Skip to main content
Back to Blog
AI/MLInnovationEnterprise
2 October 202613 min readUpdated 2 October 2026

Ai2 Open-Sources AstaBrief, a Fast Model for Scientific Report Generation

Ai2 Open Sources AstaBrief, a Fast Model for Scientific Report Generation Published October 2, 2026 Language models can help researchers search literature, synthesize evidence,...

By AI Engineering Team

Ai2 Open-Sources AstaBrief, a Fast Model for Scientific Report Generation

Published October 2, 2026

Language models can help researchers search literature, synthesize evidence, and investigate complex questions. Scientific work also requires careful grounding: reports must remain faithful to the evidence, avoid expanding a study's conclusions, and provide citations that researchers can verify.

Ai2 developed AstaBrief 8B to address those requirements. The model converts a research question and retrieved literature excerpts into a cited report. It is available in Asta's Generate a report feature as Fast mode, alongside the Claude-powered Thinking mode. Ai2 is also releasing the model weights and training data so that others can study and adapt the approach.

AstaBrief was designed to reduce report-generation time while maintaining answer quality, relevance, structure, and citation grounding. Across the complete Asta pipeline, Fast mode averages 51.1 seconds per report, compared with 178.5 seconds for Thinking mode, making it approximately 3.5 times faster.

The model is based on Qwen3-8B. Its development used tens of thousands of real research queries, citation-focused filtering, preference data, and a report-generation pipeline that produces the complete report in one pass instead of writing it section by section.

Open weights also allow institutions to run AstaBrief on their own infrastructure. This can be important when research questions contain sensitive or unpublished information. Ai2 is additionally releasing an example workflow that researchers can adapt to generate reports from their own PDF collections.

Most of the training and evaluation described for AstaBrief was completed in 2025. The proprietary models used to create training data and provide comparison points therefore reflect the frontier at that time. Ai2 did not rerun the complete evaluation against current frontier models, so the results primarily describe the training and system-design choices tested during development.

Training the model

The goal was to build an open-weights model capable of long-form scientific synthesis while preserving four key qualities:

  • Answer quality
  • Relevance
  • Report structure
  • Citation grounding

Ai2 focused on post-training data, evaluation, and the surrounding report-generation system rather than training a new base model from scratch.

The organization is also studying scientific language models through NSF OMAI, a U.S. national initiative led by Ai2 to develop fully open AI infrastructure and models for scientific discovery. That work examines how scientific communities use models, how needs differ across fields, and where general-purpose systems fall short.

Ai2 considered reinforcement-learning-based methods for AstaBrief. Earlier work, including DR Tulu, showed that reinforcement learning can improve long-form report generation for open-weights models, particularly when judge models participate in training. For AstaBrief, the team instead used a simpler combination of supervised fine-tuning (SFT) and direct preference optimization (DPO).

The choice was motivated by the cost and instability of reinforcement-learning-based training. SFT and DPO offered a more manageable setup that was easier to debug and iterate on. This placed greater importance on constructing training examples that demonstrated the desired report-writing behavior.

The model was also trained to produce a complete report in one pass from a user query and relevant retrieved snippets. This bypassed the snippet summarization and clustering stages used by Claude-based Thinking mode and avoided generating answers section by section. During development, Ai2 found that this approach could improve speed without sacrificing performance.

Collecting SFT training data

The training pipeline began with real user queries submitted through the system described in Ai2's paper, “Synthesizing scientific literature with retrieval-augmented LMs,” and through ScholarQA, the framework behind Asta's report-generation feature.

Rather than relying only on synthetic prompts or benchmark tasks, the team wanted AstaBrief to learn from questions submitted by scientists. An analysis of hundreds of thousands of Asta queries found that expert researchers often provide substantial context, multiple constraints, and relationships between concepts instead of short keyword-style prompts.

User studies also found that researchers have different preferences for how AI should participate in their work. Some use models for ideation and experimentation, while others prefer more limited roles in synthesis, literature monitoring, or pattern discovery. Across those preferences, researchers consistently want clearer source traceability, more visibility into model behavior, and greater control over the context used for a response.

Ai2 filtered the collected user logs for quality, relevance, and privacy. The process removed beta-tester and bot traffic, discarded queries that were too short to be meaningful, and used an LLM-based filter to identify non-English queries, non-scientific requests, and prompts containing personal information. The remaining pool contained 90,000 research-focused queries.

For SFT, the team generated full-report targets using the multi-step ScholarQA pipeline. The pipeline retrieved relevant literature, organized the material into sections, and used a report-generation model to synthesize the evidence into a cited report.

The training data used a combination of proprietary systems:

  • Claude 3.5 Sonnet
  • Claude 3.7 Sonnet
  • o3
  • o4-mini
  • GPT-4.1

After quality filtering, the process produced 47,000 usable training examples.

Creating DPO preference pairs

DPO requires pairs of reports, with one report identified as preferable to the other. Ai2 created these pairs from a separate set of queries that was not used to generate the SFT data.

One report for each query came from the existing ScholarQA pipeline, usually backed by Claude 3.5 Sonnet or Claude 3.7 Sonnet. The competing report was generated by providing ScholarQA's retrieved literature excerpts to one of several other models:

  • o3
  • o4-mini
  • DeepSeek-V3
  • DeepSeek-R1

Two judge models, GPT-4.1 and DeepSeek-R1, compared each pair and selected a winner. The judges were aligned with human preferences at 95% agreement. Ai2 kept only pairs for which both judges agreed, reducing noise in the preference data.

The final DPO dataset contained approximately 6,000 examples. Using multiple report generators and requiring agreement between two judges provided a way to construct preference data without treating the output or judgment of a single model as ground truth.

Filtering for stronger attribution

The primary evaluation target was SQABench-CS2, a set of 200 user-written computer science research questions. Ai2 tracked four metrics during development:

  • Rubric score: Measures how much necessary content the report covers.
  • Answer precision: Measures whether each paragraph is relevant to the question.
  • Citation precision: Measures whether each citation supports the claim to which it is attached.
  • Citation recall: Measures whether the report's claims are fully supported by its citations.

The final model was also evaluated on DeepScholarBench, a 63-query benchmark for long-form research synthesis built from recent arXiv papers. Additional pairwise evaluations compared AstaBrief with reports from the Claude-powered pipeline, including an LLM-judged comparison on SQABench-CS2 and a small human study.

A report can appear polished while drifting from the question or attaching citations to claims that the underlying evidence does not support. Citation support also does not guarantee that a report preserves the strength and scope of a study's conclusions.

For example, a model might turn a finding about a particular sample into a general claim about an entire population. It might convert a result stated in the past tense into a present-tense statement that appears universally true, or transform a descriptive finding into a recommendation for clinicians, policymakers, or researchers.

These problems can broaden the apparent scope of evidence without producing an obviously false statement. A cited sentence may be related to its source while still overstating what the researchers established. The development metrics emphasized relevance, coverage, and citation grounding, but a more comprehensive evaluation of scientific report writers should also test whether they preserve the scope and strength of source claims.

Initial SFT runs improved overall content quality but remained behind the Claude-powered report-generation pipeline in answer precision and citation quality. The model was writing better reports, but it was not consistently grounded in evidence.

Ai2 therefore tested four statistical filters for identifying weaker synthetic training examples:

  • Output-to-input token ratio: Very high ratios often indicated that the report generated too much text from too little evidence.
  • Citation relevance: The team averaged the retrieval relevance scores of papers cited in each synthetic report. Lower averages suggested greater reliance on lower-ranked evidence.
  • Citation density: This measured the share of statements with at least one citation. Low-density reports often contained long stretches of unsupported text.
  • Citation diversity: This measured the share of retrieved papers cited in the answer. Low scores suggested excessive reliance on a small number of papers.

The largest improvement came from removing synthetic reports with low citation density. More aggressive filtering, combinations of filters, and learning-rate sweeps did not produce meaningful additional gains.

This result suggested that more elaborate filtering was not automatically better. A relatively simple signal, whether reports consistently cited their claims, was more useful than several complex combinations. Scientific specialization does not necessarily require adding more scientific text during pretraining. The composition and quality of post-training data, along with whether that data demonstrates behaviors such as grounding and attribution, can materially affect model performance.

The emphasis on grounded and useful output also reflects findings from Asta user research. Participants indicated that more text is not necessarily more helpful. They wanted concise synthesis and sufficient source traceability to review and verify results without processing unnecessary material.

After producing a stronger SFT checkpoint, Ai2 applied DPO training. This further improved performance, bringing AstaBrief within the range of the Claude-powered report pipeline in Asta and DR Tulu on report-generation evaluations.

Validating the approach

AstaBrief was intended to operate within Asta's agentic report-generation framework rather than necessarily function as a standalone model. The central question was whether it could retain important report qualities while enabling a faster and less expensive pipeline.

The evaluation therefore considered more than whether AstaBrief could match a stronger proprietary model on individual benchmarks. It also examined how much quality could be retained with a simpler system.

In the development evaluations, AstaBrief was competitive with the Claude-powered pipeline and DR Tulu across several measures of answer and citation quality. Qwen3-8B was evaluated only on the SQABench-CS2 test set. SQABench-CS2 contains user-written computer science research questions, while DeepScholarBench measures long-form research synthesis using different metrics, so the scores from the two benchmarks are not directly comparable.

A separate 14-question human study involved three scientific researchers. Each researcher contributed four or five questions and ranked reports from three systems for overall preference, completeness, relevance, organization, and citation accuracy. Ties were allowed.

DR-Tulu received the highest overall preference. However, two of the three researchers preferred AstaBrief to the other systems on citation-accuracy measures. This result highlighted the contribution of the SFT data-quality filters.

The reported results are best interpreted as validation of the engineering approach used during development, rather than as a current ranking of the base model against today's frontier systems. Model capabilities change quickly, while the lessons about data construction, attribution filtering, and serving are intended to be more broadly applicable.

Early use in Asta

Fast mode also produced early usage data among Asta users:

  • Of 374 users who tried Fast mode, 29.1% used it on two or more days.
  • Users generated an average of 3.67 report threads with Fast mode.
  • 23% of users who tried Fast mode continued using it and did not switch back to Thinking mode for later threads.
  • An additional 18% switched between Fast and Thinking modes depending on their goals. These users used Fast mode for approximately 40% of their threads.

Feedback was limited, so the results do not support strong conclusions. Fast mode received positive feedback at a rate similar to Thinking mode, 84.2% versus 85.2%.

Future work

Asta's report-generation feature is the first production use of AstaBrief. It provides an open-weights Fast mode alongside the existing Thinking mode. Institutions can deploy the model on their own hardware, including behind their own firewall, without depending on a proprietary model API for report generation.

Within Asta, this also allows Ai2 to study and improve the report-generation pipeline directly while retaining Thinking mode for more compute-intensive tasks.

Planned areas of investigation include:

  • More fine-grained preference learning
  • Stronger RAG-plus-RL approaches
  • Multi-turn and multi-tool capabilities
  • Additional scientific data sources
  • Query decomposition
  • Evaluations that assess whether a model preserves the evidentiary scope of its sources

Future evaluations may look beyond whether a claim has a supporting citation. They may also assess concision, organization, and whether a model turns sample-specific findings into broad generalizations or descriptive results into recommendations.

AstaBrief is part of Ai2's continuing work on language models for science, including ScholarQA and DR Tulu, as well as future versions of Olmo. The project's findings about training data, filtering, and evaluation provide guidance for subsequent model development.