Skip to main content
Back to Blog
AI/MLInnovationSecurity
13 September 202615 min readUpdated 22 September 2026

Debating RSI, the US-China AI Gap, and Model Safeguards

Overview Nathan Lambert speaks with Jean Stanislas “JS” Denain of Epoch AI, where Denain leads the Insights Team. Their discussion examines recursive self improvement (RSI), AI...

By AI Engineering Team

Overview

Nathan Lambert speaks with Jean-Stanislas “JS” Denain of Epoch AI, where Denain leads the Insights Team. Their discussion examines recursive self-improvement (RSI), AI research automation, robotics, the relative progress of Chinese and US AI labs, distillation, model safeguards, and how Epoch AI studies a rapidly changing industry.

Both speakers emphasize substantial uncertainty about AI’s trajectory. The conversation focuses on current evidence rather than definitive predictions.

Recursive self-improvement and AI research

OpenAI and Anthropic have published analyses of how AI systems may accelerate AI development. Denain says these publications do not yet provide strong evidence of an imminent, self-sustaining acceleration or complete automation of AI research.

One notable signal is the growing use of AI systems in model development. OpenAI reported that researcher spending on Codex had been increasing by roughly 2x per month. Denain views this as evidence that researchers are receiving significant value from the system, but not as proof that a software intelligence explosion will occur within six months.

Internal AI labs may also track metrics that are not publicly available, such as compute multipliers for pre-training teams or changes in the performance of specialized research groups. These measurements could provide early warnings of accelerating progress, although Denain does not think recent public statements necessarily demonstrate that such metrics are already “going crazy.”

The speakers connect RSI to existential risk through the scale of capabilities that future AI systems might acquire. Denain describes himself as broadly focused on capabilities: disagreements about extreme AI scenarios often depend on how capable AI systems become, how quickly they reach those levels, and how effectively their capabilities spread through the economy.

A high-capability scenario could involve AI systems automating AI research, managing factories, accelerating robotics, and supporting a rapidly expanding industrial and scientific economy. The key uncertainties concern real-world bottlenecks, including compute, physical infrastructure, strategic decision-making, and the ability to give AI systems useful control over complex projects.

What does 10x AI research productivity mean?

The speakers distinguish between the productivity of individual researchers and the overall progress of an AI company or research field.

A typical researcher may become 10 times more productive at the tasks they currently perform, such as writing code, troubleshooting experiments, or running analyses. That does not necessarily mean that every researcher improves by 10x, or that an entire organization advances 10 times faster. Compute availability, coordination, research direction, and other bottlenecks may limit the effect.

Lambert compares this to junior PhD students becoming much more productive while their advisors’ research agendas fail to progress at the same rate. Faster experimentation may also be difficult to distinguish from the arrival of much larger amounts of inference compute, which can make existing research processes more efficient without fundamentally changing scientific understanding.

OpenAI’s reported Codex usage patterns illustrate this distinction. Engineering, troubleshooting, and related tasks showed increased usage, while strategic decisions such as compute allocation and research direction appeared to show less obvious improvement. This raises two separate questions:

  1. How much faster can AI research become if strategic decision-making does not improve substantially?
  2. Can longer-horizon reinforcement learning or transfer from other domains eventually improve strategic research decisions as well?

Lambert expects major gains in well-scoped knowledge work and in the efficiency of existing research processes. He is less certain that current methods can generate environments that are an order of magnitude harder, such as discovering fundamental chemical breakthroughs or solving major open mathematical problems.

Denain agrees that some forms of progress, especially inference efficiency, are well-scoped and therefore easier to accelerate. He also suggests measuring the improvement in a company’s overall capabilities using trends such as Epoch AI’s Effective Compute Index (ECI). A large increase in the slope of that trend would be an important signal, although benchmark optimization could cause measured performance to improve faster than broader real-world capabilities.

Robotics and industrial expansion

Lambert is skeptical that robotics will scale as quickly as language models. He expects robotics to resemble the gradual deployment of self-driving cars more than the rapid adoption of large language models. In his view, a two-to-five-year timeline for widespread robotic industrial expansion appears unlikely, although much longer-term progress remains possible.

Several obstacles could slow deployment:

  • Building factories that can produce robots at scale
  • Re-industrializing regions such as the United States
  • Improving visual and action models
  • Handling the variety and dexterity required in physical environments
  • Establishing reliability in real-world operations

Denain agrees that a two-to-five-year timeline is uncertain, but questions whether factory construction is necessarily the main constraint. If robots could reliably substitute for blue-collar workers, the potential market and financial incentives would be enormous. The rapid construction of US data centers suggests that physical build-out can occur quickly when demand and investment are strong.

He remains less certain about the capabilities required for industrial robotics. Extremely dexterous, human-like robots may not be necessary for large-scale industrial expansion if factories are designed around more limited, specialized machines. Amazon’s approach of building facilities with robotics in mind illustrates how environments can be adapted to available robotic capabilities.

Both speakers distinguish between constrained, repetitive work and flexible activity in human environments. Manufacturing robots may become highly capable sooner than robots that navigate public spaces and interact with people. They also note that the relevance of robotics to catastrophic risk depends on threat modeling: the minimum capabilities required to acquire significant physical power may be lower than those required for fully human-level dexterity.

The gap between US and Chinese AI models

When asked to estimate the difference between leading US and Chinese AI models, Denain gives a rough estimate of six to eight months based on public model release dates and ECI-style comparisons. He notes that the gap could be larger when comparing the date a model was completed rather than the date it was publicly released, because US labs may delay deployment longer than Chinese labs.

Lambert suggests that Kimi K3 and GLM 5.2 appeared closer to leading US models, perhaps two to four months behind according to some measurements. Denain considers a four-to-eight-month range plausible, while a two-month estimate may reflect measurement noise or overfitting.

Epoch AI examined whether benchmark contamination or overfitting could explain much of the difference. Comparisons using private benchmarks and benchmarks published after model release did not show a statistically significant change in the estimated gap. Denain nevertheless cautions that ECI may understate the difference because some capabilities, such as broad user-facing performance and practical deployment experience, are not fully represented by benchmarks.

Distillation and other explanations for the gap

If distillation from leading US models were effectively stopped, Denain estimates that the gap might increase from roughly six months to eight or nine months over the following period. He has become more convinced that distillation is an important factor, partly because recent work, including the Stolen Thoughts paper, suggests that extracting useful reasoning information may be relatively easy.

He also points to Claude routers, which can provide data reflecting realistic usage patterns. Such data may improve general capabilities, although it does not fully explain differences on established benchmarks, since benchmark-like prompts can be generated directly.

Another possibility is that supervised fine-tuning on high-quality language trajectories provides a valuable initialization. Recent discussions of mid-training and reinforcement learning suggest that reinforcement learning may be especially effective when combined with repeated cycles of supervised fine-tuning, mid-training, and further reinforcement learning, rather than used as a single isolated stage.

Lambert remains unconvinced that distillation explains most of the gap. He argues that Chinese labs may benefit from several other factors:

  • Faster public release after training is complete
  • Strong focus on benchmark performance
  • Large pools of technical talent
  • Effective execution of incremental improvements
  • Access to ideas and techniques developed by US labs
  • Commercial availability of reinforcement-learning environments and evaluation tools

The speakers also discuss a “four-minute mile” effect. Once one organization demonstrates that a technical approach works, other organizations can pursue it with greater confidence and avoid spending as much time exploring alternatives. Lambert thinks Chinese labs may be operating on shorter development horizons, while US labs invest more in open-ended research and longer-term architectural or training innovations.

Other possible explanations include smaller differences in access to researchers, data, or labor than in access to compute. Ideas may also spread through personnel movement, public model use, data cleaning, reward modeling, and the purchase of commercially developed training environments.

What Chinese AI job postings reveal

Epoch AI analyzed 1,604 Chinese AI job postings in work by Denain and Cheryl Wu. Denain describes the results as a source of useful facts and anecdotes rather than a precise trend line.

Job listings can reveal how companies organize their operations. For example, hiring for data center roles may suggest that some companies are building their own facilities rather than relying entirely on rented compute. Listings also provide information about company strategies:

  • Z.ai, the brand associated with Zhipu, appeared to have more business-to-business sales roles than MiniMax or Moonshot.
  • MiniMax appeared more focused on international expansion than Z.ai.
  • DeepSeek showed fewer obvious signals of these particular strategies.

AI can help automate some parts of this analysis, such as classifying positions as research, engineering, research engineering, or go-to-market roles. This makes it possible to track changes in hiring patterns over time. However, Denain says that identifying the most important findings remains more dependent on human judgment because language models do not always select the most meaningful trends.

Data markets are more difficult to study. Data is strategically important, difficult to quantify, and often handled confidentially. Chinese companies appear to be bringing more data work in-house, and many are hiring interns and junior staff who may contribute to data operations. Frontier US companies have also increased their internal data efforts. Job postings can reveal individual initiatives, but they do not yet provide a reliable picture of the full market.

Model safeguards, open models, and misuse

The speakers discuss the difficulty of evaluating safeguards for frontier models. A reported incident involving three people who used public Claude models to access OpenAI systems raises questions about both model safeguards and the security practices of AI companies. Denain says it is difficult to determine how much such an incident should update beliefs about either area.

Open models create a different safety trade-off. They may be less capable than the strongest closed models, but users can run them without provider monitoring, classifiers, or other controls. Closed models may have stronger safeguards while offering more powerful capabilities, making it unclear which category creates the greater overall misuse risk.

Lambert distinguishes between several layers of protection:

  • Model-level safeguards built into the weights
  • Inference-time classifiers and monitoring
  • Restrictions on API functionality
  • Synchronous detection during use
  • Asynchronous analysis of longer trajectories and campaigns

A determined user may be able to bypass some protections, especially if they can use a less-aligned model to search for ways around the safeguards of a more-aligned model. Extracting hidden reasoning traces through API or serving weaknesses is another concern because it can expose information useful for distillation.

Denain notes that the importance of a successful attack depends partly on its duration. A short attack may be difficult to detect in real time, while a longer campaign may be easier to identify through retrospective analysis. The effectiveness of identity checks and other monitoring systems remains uncertain.

For open models, the immediate risk may come more from the absence of provider-side safeguards than from deliberate safety fine-tuning. At present, relatively few people combine the motivation, expertise, and willingness required to fine-tune a model for harmful purposes. That barrier could fall as AI research becomes more accessible.

Both speakers argue that better threat modeling is needed. The relevant question is not only whether a model can be misused, but which groups are most likely to attempt harmful actions, what resources they possess, and whether open or closed systems would be more useful to them.

How Epoch AI selects research projects

Denain says project selection at Epoch AI is primarily driven by curiosity about questions that appear important. The organization was created in part to provide a public, maintained resource on major AI trends that industry and academia might not cover adequately.

The team also looks for information that is available but underused. Job postings and data center tracking are examples of this approach. Data centers are large physical objects that can potentially be identified and monitored, while company disclosures and public filings may provide additional information as companies grow and become public.

Access varies by research area. Data centers are currently relatively observable, although they could become more strategically protected. Tracking model-training compute has become more difficult when companies do not publish detailed technical papers. Data is even harder to study because it is secretive and difficult to standardize.

What a frontier post-training process might look like

The speakers attempt to outline a possible post-training process at a frontier AI lab, while emphasizing that the description is uncertain.

One possible structure includes specialized teams working on areas such as mathematics, biology, or finance. These teams may:

  • Request and curate training data
  • Build and evaluate reinforcement-learning environments
  • Identify high-quality evaluations to use as targets
  • Tune hyperparameters for their domains
  • Produce collections of environments and recommended training settings

Those teams may then compete for inclusion in a large final reinforcement-learning run. Another possibility is that domain-specific experts are distilled into a broader model before additional training. The exact balance between multiple expert runs, distillation, model merging, mid-training, supervised fine-tuning, and a final reinforcement-learning run remains unclear.

Lambert says open-weight models show different approaches, including multi-teacher on-policy distillation and sequential reinforcement learning. Understanding how frontier labs choose environments, schedule training stages, and combine domain-specific experts could provide a better model of what is actually being scaled.

Denain’s tentative view is that frontier training may have moved away from many separate experts and checkpoints toward one large late-stage run, preceded by extensive preparation. Neither speaker claims firm evidence for this structure.

Conclusion

The discussion leaves many questions unresolved. Current evidence shows rapid progress in AI-assisted software development, benchmark performance, inference efficiency, and some forms of research automation. It does not yet establish how quickly AI systems will improve strategic research, robotics, industrial production, or open-ended scientific discovery.

The US-China capability gap may be measured in months rather than years, with estimates affected by release timing, benchmarks, distillation, data, compute, labor, and organizational choices. Safeguards also remain difficult to assess because model capabilities, provider monitoring, and real-world threat models interact in complex ways.

For both speakers, the central task is to improve measurement and identify the bottlenecks that determine how quickly AI capabilities translate into broader economic, scientific, and physical-world effects.