Antes del algoritmo: ética de la información y responsabilidad en los datos de entrenamiento de la inteligencia artificial generativa
DOI:
https://doi.org/10.54167/roee.v4i2.2446Palabras clave:
datos de entrenamiento, ética de la información, gobernanza de datos, inteligencia artificial generativa, procedencia de la información, responsabilidad informacionalResumen
La expansión de la inteligencia artificial generativa ha intensificado el uso de grandes volúmenes de información para el entrenamiento de modelos computacionales. Aunque buena parte de la discusión ética se ha concentrado en los resultados producidos por estos sistemas, este artículo sostiene que la responsabilidad debe analizarse desde una etapa anterior: la selección, recopilación, transformación e incorporación de información a los conjuntos de entrenamiento. Mediante una investigación documental de carácter crítico-interpretativo se examina literatura especializada sobre ética de la información, documentación de conjuntos de datos, procedencia, gobernanza, sesgos y consentimiento, además de marcos internacionales recientes sobre inteligencia artificial. A partir de este análisis se propone una cadena de responsabilidad informacional integrada por seis dimensiones: procedencia, legitimidad de uso, calidad, representación, trazabilidad y responsabilidad. Se sostiene que la accesibilidad técnica de la información no constituye, por sí misma, una justificación ética suficiente para incorporarla al entrenamiento de sistemas generativos. Se concluye que una inteligencia artificial responsable requiere extender la gobernanza hacia las condiciones informacionales que hacen posible su desarrollo.
DOI: https://doi.org/10.54167/roee.v4i2.2446
Descargas
Referencias
Bender, Emily M., y Batya Friedman. “Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science.” Transactions of the Association for Computational Linguistics 6 (2018): 587–604. https://doi.org/10.1162/tacl_a_00041.
Bender, Emily M., Timnit Gebru, Angelina McMillan-Major y Shmargaret Shmitchell. “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜.” En Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–23. New York: Association for Computing Machinery, 2021. https://doi.org/10.1145/3442188.3445922.
Birhane, Abeba, Vinay Uday Prabhu y Emmanuel Kahembwe. “Multimodal Datasets: Misogyny, Pornography, and Malignant Stereotypes.” arXiv, 2021. https://doi.org/10.48550/arXiv.2110.01963.
Bommasani, Rishi, Kevin Klyman, Sayash Kapoor, et al. “The 2024 Foundation Model Transparency Index.” Transactions on Machine Learning Research, 2024. https://openreview.net/forum?id=38cwP8xVxD.
Budapest Open Access Initiative. “Read the Budapest Open Access Initiative.” Open Society Institute, 2002. https://www.budapestopenaccessinitiative.org/read/.
Dodge, Jesse, Maarten Sap, Ana Marasović, et al. “Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.” En Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 1286–1305. Association for Computational Linguistics, 2021. https://doi.org/10.18653/v1/2021.emnlp-main.98.
European Parliament and Council of the European Union. “Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act).” Official Journal of the European Union, 2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj.
Floridi, Luciano. “Information Ethics, Its Nature and Scope.” ACM SIGCAS Computers and Society 36, núm. 3 (2006): 21–36. https://doi.org/10.1145/1195716.1195719.
Floridi, Luciano. The Ethics of Information. Oxford: Oxford University Press, 2013. https://doi.org/10.1093/acprof:oso/9780199641321.001.0001.
Gebru, Timnit, Jamie Morgenstern, Briana Vecchione, et al. “Datasheets for Datasets.” Communications of the ACM 64, núm. 12 (2021): 86–92. https://doi.org/10.1145/3458723.
Kandpal, Nikhil, Brian Lester, Colin Raffel, et al. “The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text.” arXiv, 2025. https://doi.org/10.48550/arXiv.2506.05209.
Longpre, Shayne, Robert Mahari, Anthony Chen, et al. “A Large-Scale Audit of Dataset Licensing and Attribution in AI.” Nature Machine Intelligence 6 (2024): 975–987. https://doi.org/10.1038/s42256-024-00878-8.
Longpre, Shayne, Robert Mahari, Ariel Lee, et al. “Consent in Crisis: The Rapid Decline of the AI Data Commons.” Advances in Neural Information Processing Systems 37 (2024): 108042–108087. https://doi.org/10.52202/079017-3431.
Pushkarna, Mahima, Andrew Zaldivar y Oddur Kjartansson. “Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI.” En Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 1776–1826. New York: Association for Computing Machinery, 2022. https://doi.org/10.1145/3531146.3533231.
Sambasivan, Nithya, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen K. Paritosh y Lora M. Aroyo. “‘Everyone Wants to Do the Model Work, Not the Data Work’: Data Cascades in High-Stakes AI.” En Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, artículo 39, 1–15. New York: Association for Computing Machinery, 2021. https://doi.org/10.1145/3411764.3445518.
UNESCO. Recommendation on the Ethics of Artificial Intelligence. Paris: UNESCO, 2021. https://unesdoc.unesco.org/ark:/48223/pf0000380455
Wan, Alexander, Kevin Klyman, Sayash Kapoor, et al. “The 2025 Foundation Model Transparency Index.” arXiv, 11 de diciembre de 2025. https://doi.org/10.48550/arXiv.2512.10169.
Publicado
Número
Sección
Licencia
Derechos de autor 2026 Orexis. Exploraciones Éticas

Esta obra está bajo una licencia internacional Creative Commons Atribución-NoComercial 4.0.








