Teaching Everyone to Fish for Tokens
Open Models and the Linux Analogy A common comparison for today’s open model ecosystem is the development of foundational open source software projects such as Linux. The analog...
By AI Engineering Team
Open Models and the Linux Analogy
A common comparison for today’s open model ecosystem is the development of foundational open-source software projects such as Linux. The analogy is useful, but it describes only one possible path toward a self-sustaining ecosystem.
Once Linux became widely adopted, its popularity reinforced its position as a practical tool for many types of work. Fully open language models, meaning models released with a complete training recipe, data, and code, are closer to an open-source operating system. Open-weight models, which generally provide model weights and inference code, are more like particular software versions installed within projects that depend on them.
Model weights can become outdated quickly, but heavily used versions may remain relevant for years. Many companies, for example, still operate workflows based on Llama 3, even though agentic behaviors became a major focus later.
A complete open-source training recipe, such as the one associated with the Olmo models developed at Ai2 and earlier projects such as Pythia from EleutherAI, requires substantial resources. However, any company can theoretically take the recipe, modify it, and run it to produce a new set of model weights. In the strongest version of this model, the community can also contribute improvements to data and training code for future releases.
Nvidia is investing heavily in models that are close to open source. For its Nemotron models, the company releases as much legally permissible data and training code as possible. This approach supports a world in which many organizations can build models that generate tokens and intelligence is not controlled by a small number of providers. Such a world would also create substantial demand for inference across companies that may purchase Nvidia hardware and services.
The Economic Challenge of Open Model Development
Open-source AI faces a difficult future because building leading models requires enormous capital. The ability of companies to build competitive models has remained more accessible for longer than many observers expected. Still, the prevailing assumption is that model training will become too expensive and that open-source methods will fall too far behind. Under that assumption, creating a new laboratory focused on large language model training would be difficult to sustain.
Two broad outcomes are possible.
A Self-Reinforcing Open Model Ecosystem
If the open-source training approach works for Nvidia, the company could generate more demand for its chips and services than it spends developing the models. Nvidia is reportedly spending $26 billion on this effort, although it remains uncertain whether the investment will produce that result or whether the capital requirements of AI will force more companies to leave model training.
There have not yet been many signs that companies are broadly abandoning the field. Databricks and 01.ai appear to be exceptions rather than evidence of a widespread retreat.
The open-source ecosystem may become increasingly dependent on Nvidia financing in the coming years. This creates a limited period in which the profits generated by the strategy must eventually return to Nvidia, or another open model company must develop platform-like financial feedback loops based on openness.
For this approach to remain viable over decades of language model development, its economic rewards would need to approach the scale of the profits generated by the APIs of Anthropic and OpenAI. That could happen through superior performance, or because the overall AI market becomes large enough to provide extensive demand for open model training, inference, and fine-tuning.
A Different Role for Open Models
If neither financially positive path develops, open models may follow a different trajectory from leading closed models. Their strengths could center more on efficiency, modifiability, and specialization.
Under this outcome, open models would remain highly useful but would serve a long-tail ecosystem alongside closed models. Closed providers may retain the most valuable areas, including knowledge-work collaboration, drug discovery, and software engineering. Open models could instead support applications such as enterprise-specific agents running on-premises, using private data for repetitive business tasks.
Increasing Complexity in Training
One reason open-source training may struggle to expand is that training is becoming more complex and abstracted. The current open model ecosystem is supported in part by strong interest in post-training. Developers can take models such as DeepSeek V4 Flash, Inkling Small, or GLM 5.X and fine-tune them for specific agentic tasks, including through services such as Tinker.
In recent years, post-training often referred broadly to the process of modifying a base model so that it became intelligent and usable. A shift is now underway: training a base model to become a general agentic reasoner is becoming less transparent, resembling the large-scale pretraining practices that became difficult to reproduce several years ago.
This development could change the terminology that has guided the field. Instead of a simple pretraining, midtraining, and post-training sequence, the process may increasingly be described as pretraining, reasoning training, and post-training.
As fewer organizations show interest in training complete models, investment in open-source AI may also decline. One indication could be a continued reduction in the number of open model builders releasing base models, meaning versions produced before core reasoning training.
This trend is occurring alongside experiments with revenue-sharing licenses for downstream products and inference. These licenses are attempts to create sustainable financing for near-frontier open-weight models. Their success may influence whether Nvidia’s strategy of expanding demand through open models can continue.
Open Weights as an Indirect Business Strategy
Despite these economic pressures, open-weight models are likely to remain active because making intelligence broadly available can support indirect business strategies.
Companies with very large balance sheets, including Meta and other hyperscalers, can monetize AI without relying entirely on direct token sales. If Meta released the weights of its reported Muse Spark 1.2 model, for example, it could significantly reduce the potential revenue growth of competitors such as Anthropic and OpenAI, whose businesses depend on selling access to tokens.
Nvidia and Meta would be making intelligence more widely available, but their strategies would differ. Nvidia’s goal is to enable many organizations to build and operate token-generating systems, helping create a self-sustaining ecosystem. Meta’s approach is to increase the supply of accessible tokens and place pressure on companies whose revenue depends on selling them.