Ai2 released AstaBrief 8B on October 2 as the model behind Fast mode in Asta's Generate a report feature. The downloadable model takes a research question and excerpts retrieved from scientific papers, then writes a cited report in one pass. For a researcher, the useful distinction is where the papers come from and what the citations actually support: AstaBrief writes from supplied evidence; it is not, by itself, a literature search service. Ai2's release and model card describe that boundary.
The big change
- What changed: Ai2 put AstaBrief into Asta's live Fast mode and released the model weights and training materials.
- Why it matters: Research groups can inspect or adapt the report generator and run the open-weights model on their own infrastructure, according to Ai2.
- What to watch: Report quality depends on the retrieved literature and on checking each consequential claim against its source. Asta's Claude-powered Thinking mode remains a separate, multistep option.
From question to report
The hosted route begins in Asta's Generate a report feature, where Fast mode uses AstaBrief. Ai2 says its report system retrieves relevant literature before passing the question and excerpts to the model. The model then writes the report in one pass, omitting the quote extraction, clustering and section-by-section drafting stages used by Thinking mode. The public ScholarQA repository describes retrieval from Semantic Scholar's passage and keyword search, followed by reranking; its lite implementation builds a prompt from reranked references and calls the configured generator. These are documented components, not a test of the live Asta service by BIG CHANGE.
For the downloadable route, the AstaBrief model card recommends its linked prompt format and an input containing the question and section references. It provides example transformers/vLLM inference code with 4,096 maximum output tokens. The card's example sets model_name to the SFT checkpoint, while the page itself describes the later DPO-tuned AstaBrief_8B; a researcher choosing a checkpoint should resolve that discrepancy rather than assume the example runs the final model as printed. Ai2 also links example ScholarQA lite code that researchers can adapt for reports from their own PDFs. The linked directory is code, not a documented turnkey PDF setup with a stated hardware minimum.
That leaves a practical sequence for a local experiment: obtain papers one is permitted to use; extract and select relevant passages with paper identifiers; format the question and references as the card recommends; run the chosen checkpoint; and inspect the resulting report beside those passages and the original papers. These are the components the sources document, not steps BIG CHANGE executed. A full local replica also needs retrieval, PDF parsing, reference handling and serving choices that the model weights alone do not supply.
Check the evidence before using the prose
Start with retrieval. Record the query and the set of papers or PDFs offered to the generator. An omitted study or irrelevant excerpt can distort the answer before AstaBrief writes a word. Then open each cited paper for material claims. Check whether the cited passage supports the precise sentence, including its population, method, date and degree of certainty. A citation can be real and related while the report widens a result from one sample to an entire field, or turns a descriptive finding into advice. Ai2 explicitly identifies this claim-scope failure in its release.
Finally, look for important uncited claims and places where the report relies repeatedly on a narrow slice of the retrieved papers. Ai2 says filtering training reports for citation density improved its development results. That finding explains why attribution matters, but citation density alone cannot establish that the cited claims are faithful. Treat the generated report as a working synthesis to revise against the papers, especially before citing it in research or making a consequential decision.
What the published evaluation establishes
Ai2 reports that, across the Asta pipeline, Fast mode averaged 51.1 seconds per report against 178.5 seconds for Thinking mode, about 3.5 times faster in that comparison. Its model card reports AstaBrief's results on the ScholarQA-CS2 computer science questions and DeepScholarBench, and the release describes a separate, small human comparison: three researchers contributed 14 questions. Those are Ai2's evaluations of its systems and test sets, not an independent demonstration that any particular report will be accurate. The release says most training and evaluation took place in 2025 and the full evaluation was not rerun against current frontier models.
The final checkpoint is public under Apache 2.0, and Ai2 lists its training datasets and prompts in an AstaBrief collection. The model card says the model is intended for research and education under Ai2's responsible-use guidance. Its training note mentions eight H100 GPUs for DPO training; that is not a published minimum for inference. The reviewed primary materials do not state a per-report price for hosted Asta, a minimum local GPU specification or a dependable local cost per report. Anyone planning deployment needs to establish those costs for their own setup.
Sources & further reading
- Ai2's October 2 release explains the live Fast mode, the one-pass design, timing and evaluation limits. Its performance claims are Ai2's own.
- AstaBrief 8B model card gives the license, intended use, prompt recommendation and inference example, including the SFT checkpoint name in that example.
- AstaBrief collection lists the published checkpoints, datasets and prompts; their presence does not establish an end-to-end local reproduction.
- ScholarQA code and documentation describe the retrieval system; the lite generator shows how reranked references become a one-pass generation prompt. We inspected documentation and code; we did not run them.
- ScholarQA-CS2 evaluation repository provides benchmark code and data, including rubric construction. It helps explain the test design; it does not independently verify AstaBrief's reported scores.



