The ‘Open Source’ AI Illusion—And How Researchers Can Reclaim It

The scientific community must take a stand against misleading ‘open-source’ claims and push for AI systems that truly uphold the principles of transparency and reproducibility.

1 min read
A representational image [Shahadat Rahman/Unsplash]

The term ‘open source’ has long been associated with transparency, collaboration, and free access to software. However, as Nature recently highlighted, many AI companies are misusing this label, undermining the very principles that have driven scientific progress for decades. While some AI models are branded as ‘open source,’ they often fail to meet the core criteria of openness, particularly when it comes to access to training data and model transparency.

Decades of open-source software, from R Studio for statistical computing to OpenFOAM for fluid dynamics, have accelerated scientific discovery by ensuring research reproducibility and fostering global collaboration. However, conventional open-source frameworks were built around source code—something that AI systems, which rely heavily on training data, do not easily conform to. Many foundational AI models, such as Meta’s Llama series, Microsoft’s Phi-2, and Mistral AI’s Mixtral, fall short of true openness due to restricted access to their training data. In contrast, models like OLMo, developed by the Allen Institute for AI, and community-driven initiatives such as LLM360’s CrystalCoder better align with genuine open-source principles.

One of the biggest concerns is the rising trend of ‘openwashing,’ where companies exploit the open-source label while restricting access to key components. Nature reports that some firms may be doing this to evade regulatory scrutiny under the European Union’s 2024 AI Act, which exempts open-source AI from certain regulations. This deceptive practice threatens scientific integrity, as researchers may be forced to rely on closed corporate systems with unverifiable methodologies.

To address this challenge, the Open Source Initiative (OSI) launched a global effort to define Open Source AI (OSAID 1.0), setting standards for true AI transparency. This initiative introduces the concept of ‘data information,’ requiring companies to disclose details about training data sources, characteristics, and preparation methods—even if full data release isn’t legally possible. By advocating for a transition to a more inclusive ‘data-commons’ model, OSI and organizations like Open Future aim to preserve transparency and accountability in AI development.

For scientists and researchers who rely on AI-driven tools, engaging with OSAID 1.0 is a critical first step. Public funding agencies and governments also have a role to play in ensuring AI models meet genuine open-source standards. As Nature points out, institutions such as the US National Institutes of Health already mandate open licensing for research data, and countries like Italy require open-source software for public administration. These policies can serve as a blueprint for regulating AI in a way that prioritizes openness, replicability, and scientific integrity.

The scientific community must take a stand against misleading ‘open-source’ claims and push for AI systems that truly uphold the principles of transparency and reproducibility. Without these efforts, the future of AI in research could be dictated by closed, proprietary models—threatening not just access, but the very foundation of scientific progress.

Sri Lanka Guardian

The Sri Lanka Guardian is an online web portal founded in August 2007 by a group of concerned Sri Lankan citizens including journalists, activists, academics and retired civil servants. We are independent and non-profit. Email: editor@slguardian.org

Leave a Reply

Your email address will not be published.

Latest from Blog