Rapid AI Progress Will Accelerate Engineering, Not Necessarily Lead to General Superintelligence
Engineering Is Becoming Less of a Bottleneck Leading industry researchers increasingly expect AI systems to outperform them at their jobs within a few years. One reason for this...
By Software Development Team
Engineering Is Becoming Less of a Bottleneck
Leading industry researchers increasingly expect AI systems to outperform them at their jobs within a few years. One reason for this expectation is the rapid improvement in infrastructure and engineering capabilities surrounding AI models.
AI systems are likely to become superhuman at distributed GPU engineering. That capability would make it much easier to experiment with and refine model designs, but it would not necessarily make the models fundamentally different in nature.
Ilya Sutskever may have been early in describing AI as entering an “era of research.” Research has already become easier with coding agents, which provide a flexible and engaging way to carry out technical work. This suggests the beginning of a period in which good ideas may become more valuable than strong execution in software development.
This transition is still in its early stages, and it is not the first shift in the balance between AI research and engineering. Before deep learning became dominant, AI was primarily a research field. Today, leading researchers are also evaluated by their ability to implement and scale ideas across complex infrastructure.
Over the next few years, engineering is likely to become less of a bottleneck again. Some people describe this possibility as recursive self-improvement, or RSI. A more grounded description is “parallelized, AI-assisted language modeling.” This outcome does not require accepting rapid-takeoff scenarios. Improvements to infrastructure alone could significantly change how AI work is performed.
The main point is that engineering acceleration will help spread AI throughout the economy, even if models do not develop economically valuable superhuman abilities outside areas such as mathematics and coding. Much of the near-term progress is likely to come from scaling inference-time compute with existing tools, rather than from dramatic improvements in the models’ ability to conduct research.
Optimizing Training and Inference
Many parts of the AI training and inference stack are easy to measure and verify. Training metrics include tokens per second per GPU, which represents training speed. Inference metrics include tokens per prompt, FLOPs per token, and the cost of producing an answer.
These metrics are highly optimizable. The field already has established subproblems and architecture trade-offs that can improve efficiency. Within a few years, AI agents may help optimize the process from end to end, bringing inference capabilities close to the maximum compute available on accelerators such as GPUs.
Companies have already achieved substantial improvements in inference efficiency. In recent years, it has been possible to reduce the cost of serving a model by around 10% to 30% after announcing a particular price point. As improvements accumulate across the entire stack, the effective cost of model intelligence could decline at a near-exponential rate over the coming years, potentially faster than recent trends.
The easiest efficiency gains may be captured within only a few years. Longer-term improvements are likely to come from co-designing accelerators and models. This process could produce additional orders of magnitude in efficiency beyond the flexibility of a general-purpose GPU platform.
During the period of rapid efficiency gains, GPU flexibility will remain valuable because the ability to explore different architectures appears to be an important lever for automated research. Automating pretraining research, particularly architecture design and data selection for the current class of models, could reasonably happen within two to three years.
Demand and Agentic Systems
These efficiency gains are likely to create a significant Jevons paradox for agentic models. As the cost of using agents declines, demand may increase rather than remain constant. The industry is still largely constrained by the challenge of determining how agents should be directed and delivered to users.
Meta’s Muse agent provides an early indication of this direction. More Muse-like experiences are likely to appear for different audiences and use cases. Their value will come from understanding how agents work and applying them effectively, rather than solely from pushing the frontier of model performance.
Effects on Scientific Research
A comparable phase in fields such as biology and chemistry could be defined by models identifying connections across scientific subfields. AI systems are highly capable of searching large bodies of literature and finding relationships across sparse networks. Those networks are often maintained by small communities of scientists who rarely interact directly.
It may be difficult to determine whether this will produce a new era of scientific discovery, such as cures for most cancers, or whether it will primarily accelerate the trajectory that science was already following.
Reinforcement Learning Environments
Improving the general quality of reinforcement learning environments is another large source of relatively accessible progress. New reinforcement learning data companies have reached more than $100 million or $1 billion in revenue, and there are more of them than many observers might expect.
However, the average output from this sector remains notably low quality. Many researchers believe that a significant portion of the data they purchase is poor. At the same time, leading AI laboratories continue to see a clear return on investment from buying it.
The shortcomings of many reinforcement learning environments appear to be fixable. Improving their quality could therefore provide another major avenue for advancing AI systems without requiring a fundamental change in their underlying nature.