Complete AI Training

AI news ·

Ai2 open-sources AstaBrief, a fast scientific report generation model

Ai2 released AstaBrief 8B, an open-weights model that generates cited scientific reports in an average of 51.1 seconds-3.5 times faster than the Claude-powered pipeline it runs alongside.

Share

Ai2 has released AstaBrief 8B, an open-weights language model that generates cited scientific reports from a research question and retrieved literature excerpts. The model, available today as a Fast mode in the Asta platform, cuts report generation time to an average of 51.1 seconds - roughly 3.5 times faster than the 178.5 seconds averaged by the Claude-powered Thinking mode it runs alongside. The weights and training data are open-sourced so institutions can run the model on their own infrastructure, a requirement when research questions involve sensitive or unpublished work.

Training for speed and grounding

AstaBrief was built on Qwen3-8B using supervised fine-tuning and direct preference optimization rather than more expensive reinforcement-learning-based methods. The team said this simpler recipe was chosen because "RL-based training can be unstable and expensive," and they wanted to see how far report quality could be pushed with a setup that is easier to debug and iterate on. That decision put heavy emphasis on training data quality.

The SFT data came from 90,000 filtered research queries drawn from real Asta user logs. Target reports were generated using a mix of proprietary models - Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1 - and after quality filtering, 47,000 examples remained. For the DPO stage, the team built 6,000 preference pairs where two judge models had to agree on which report was better, a requirement that "cut much of the noise that typically shows up in preference data generated at scale."

The model also generates the full report in a single pass, bypassing the snippet summarization and clustering stages the Claude-based pipeline uses. "We found it was possible to do so without sacrificing performance," the team said.

Filtering for citations that actually support claims

Early SFT runs improved content quality but lagged on answer precision and citation grounding. The team tested four statistics-based filters on the synthetic training data: output-to-input token ratio, citation relevance, citation density, and citation diversity. The strongest gains came from filtering out reports with low citation density - essentially, removing examples where large stretches of text lacked any supporting citation.

"A relatively simple signal - whether the synthetic reports consistently cited their claims - was more useful than several more complicated combinations we tried," the researchers wrote. They added that "scientific specialization isn't necessarily a matter of adding more scientific text to pretraining; the composition and quality of post-training data and whether it demonstrates behaviors like grounding and attribution can materially change how the resulting model performs."

On the SQABench-CS2 evaluation set, AstaBrief proved competitive with the Claude-powered pipeline and DR Tulu across measures of answer and citation quality. In a small 14-question human study, two of three scientific researchers preferred AstaBrief over other systems on citation accuracy metrics.

Early usage and where the model fits

Among 374 Asta users who tried Fast mode, 29.1% used it for two or more days, and users averaged 3.67 report threads. Twenty-three percent of those who tried it never switched back to Thinking mode, while another 18% switched between modes depending on their goals. Positive feedback rates were similar between Fast mode (84.2%) and Thinking mode (85.2%).

Because the model weights are open, institutions can deploy AstaBrief behind their own firewall without relying on a proprietary model API. Ai2 is also releasing an example workflow that researchers can adapt to create reports from their own PDFs. The team noted that future work will explore more fine-grained preference learning, stronger RAG-plus-RL approaches, and evaluations that test whether models preserve the evidentiary scope of their sources - catching cases where a model turns a sample-specific finding into a broad generalization.

Why this matters for researchers and writers

For anyone producing or consuming scientific literature reviews, AstaBrief represents a concrete shift toward faster, verifiable synthesis. The model's single-pass architecture and open weights mean reports arrive in under a minute, and every claim can be traced back to its source. The emphasis on citation density filtering during training also addresses a persistent problem with AI-generated research summaries: text that sounds authoritative but drifts beyond what the cited evidence actually supports. Professionals working across AI research and generative AI will find the open-source release useful as a test case for adapting small models to domain-specific demands without the cost and latency of large proprietary systems.

Share