Open source software, democratizing statistical analysis in data science

Authors

DOI:

https://doi.org/10.62305/alcon.v4i5.361

Keywords:

Open source; data analysis; democratization

Abstract

Open source software revolutionized statistical analysis by democratizing access to advanced tools in data science. This model allowed users from various disciplines to access, modify and adapt high-level technologies to meet their specific needs, promoting equity and innovation. Featured technologies included Apache Hadoop and Apache Sqoop, which made it easier to manage and analyze large volumes of data in distributed environments. Hadoop provided a scalable and efficient framework for massive data storage and processing, while Sqoop has enabled data transfer between relational databases and Hadoop, integrating heterogeneous sources into analytical workflows. The democratization of open source software was driven by initiatives such as the Free Software Foundation (FSF) and the concept of copyleft, which have ensured that code remains accessible and reusable for any user. This movement was based on the four fundamental freedoms of free software: the freedom to run the program for any purpose, the freedom to study how it works and adapt it to your own needs, the freedom to distribute copies to help others, and the freedom to improve the program and publish those improvements for the benefit of the community. These freedoms ensured that improvements to the software benefited the entire community, strengthening collaboration and transparency in research.

Downloads

Download data is not yet available.

References

Cifras, E. e. (2024). Ecuador en Cifras . Obtenido de Ecuador en Cifras : https://www.ecuadorencifras.gob.ec/

Danilov, H. (2024). Execution of sql-like queries in distributed heterogeneous systems based on Apache Hadoop. Boletín de la Universidad Técnica Estatal de Voronezh, 2-8.

Demchenko, Y. (2024). Big Data Algorithms, MapReduce and Hadoop ecosystem. Oklahoma: Springer.

Jiang, P. (2024). Application Status of Hadoop in Data Cloud Computing. International Journal of Computer Science and Information Technology, 279-282.

Kim, S.-Y. (2024). Optimizing hadoop data locality: performance enhancement. Scalable Computing: Practice and Experience, 4558–4575.

Kusuma, D. (2024). Perbandingan Uji Performa Impala dan Hive-Hadoop. Journal Syntax Dmiration, 4837-4845.

Picarella, L. (2024). Critical Education and Digital Media: A Binomial for the Exercising of Human Rights. Review of human rights, 86-114.

Rahmani, A. M. (2024). Scheduling of Big Data Workflows in the Hadoop Framework with Heterogeneous Computing Cluster. Arabian Journal for Science and Engineering , 1-9.

Salah, H. (2024). From Micro-benchmarks to Machine Learning: Unveiling the Efficiency and Scalability of Hadoop and Spark. iJIM journal, 46-60.

Varde, A. (2024). HaaS in Environmental Computing: Hadoop-as-a-Service for Big Data Mining in Environmental Computing Applications. ACM SIGWEB Newsletter, 1-19.

Published

2024-12-16

How to Cite

Nuñez Laje , E. A. ., Estrella de las Mercedes Baquero Tapia, Rengel Sandoval , H., & Noboa Ramírez , M. Y. (2024). Open source software, democratizing statistical analysis in data science . Scientific Journal of Educational Innovation and Current Society "ALCON". ISSN 2960-8473, 4(5), 179–188. https://doi.org/10.62305/alcon.v4i5.361

Issue

Section

Original articles