# AstaBrief turns retrieved research excerpts into a cited report

> Ai2's open AstaBrief model writes cited reports from a research question and retrieved excerpts. Here is the documented workflow and how to check its claims.

By BIG CHANGE Editorial

Published: 2026-10-03T09:22:25.732Z
Updated: 2026-10-03T09:22:25.732Z
Canonical: https://bigchange.ai/blog/astabrief-research-report-citation-workflow

![Conceptual charcoal illustration of one researcher holding a report page while pointing to a passage in an open scientific paper on a reading stand.](https://bigchange.ai/api/media/file/astabrief-source-check-hero-v1.png)
AI-generated conceptual editorial illustration by BIG CHANGE.

Ai2 released AstaBrief 8B on October 2 as the model behind **Fast mode** in Asta's Generate a report feature. The downloadable model takes a research question and excerpts retrieved from scientific papers, then writes a cited report in one pass. For a researcher, the useful distinction is where the papers come from and what the citations actually support: AstaBrief writes from supplied evidence; it is not, by itself, a literature search service. [Ai2's release](https://huggingface.co/blog/allenai/astabrief) and [model card](https://huggingface.co/allenai/AstaBrief_8B) describe that boundary.

## The big change

- **What changed:** Ai2 put AstaBrief into Asta's live Fast mode and released the model weights and training materials.
- **Why it matters:** Research groups can inspect or adapt the report generator and run the open-weights model on their own infrastructure, according to Ai2.
- **What to watch:** Report quality depends on the retrieved literature and on checking each consequential claim against its source. Asta's Claude-powered Thinking mode remains a separate, multistep option.

## From question to report

The hosted route begins in [Asta's Generate a report feature](https://asta.allen.ai/), where Fast mode uses AstaBrief. Ai2 says its report system retrieves relevant literature before passing the question and excerpts to the model. The model then writes the report in one pass, omitting the quote extraction, clustering and section-by-section drafting stages used by Thinking mode. The [public ScholarQA repository](https://github.com/allenai/ai2-scholarqa-lib) describes retrieval from Semantic Scholar's passage and keyword search, followed by reranking; its [lite implementation](https://github.com/allenai/ai2-scholarqa-lib/blob/main/api/scholarqa/lite/scholar_qa_lite.py) builds a prompt from reranked references and calls the configured generator. These are documented components, not a test of the live Asta service by BIG CHANGE.

For the downloadable route, the [AstaBrief model card](https://huggingface.co/allenai/AstaBrief_8B) recommends its linked prompt format and an input containing the question and section references. It provides example `transformers`/`vLLM` inference code with 4,096 maximum output tokens. The card's example sets `model_name` to the **SFT checkpoint**, while the page itself describes the later DPO-tuned `AstaBrief_8B`; a researcher choosing a checkpoint should resolve that discrepancy rather than assume the example runs the final model as printed. Ai2 also links [example ScholarQA lite code](https://github.com/allenai/ai2-scholarqa-lib/tree/main/api/scholarqa/lite) that researchers can adapt for reports from their own PDFs. The linked directory is code, not a documented turnkey PDF setup with a stated hardware minimum.

That leaves a practical sequence for a local experiment: obtain papers one is permitted to use; extract and select relevant passages with paper identifiers; format the question and references as the card recommends; run the chosen checkpoint; and inspect the resulting report beside those passages and the original papers. These are the components the sources document, not steps BIG CHANGE executed. A full local replica also needs retrieval, PDF parsing, reference handling and serving choices that the model weights alone do not supply.

## Check the evidence before using the prose

Start with retrieval. Record the query and the set of papers or PDFs offered to the generator. An omitted study or irrelevant excerpt can distort the answer before AstaBrief writes a word. Then open each cited paper for material claims. Check whether the cited passage supports the precise sentence, including its population, method, date and degree of certainty. A citation can be real and related while the report widens a result from one sample to an entire field, or turns a descriptive finding into advice. Ai2 explicitly identifies this claim-scope failure in its [release](https://huggingface.co/blog/allenai/astabrief).

Finally, look for important uncited claims and places where the report relies repeatedly on a narrow slice of the retrieved papers. Ai2 says filtering training reports for citation density improved its development results. That finding explains why attribution matters, but citation density alone cannot establish that the cited claims are faithful. Treat the generated report as a working synthesis to revise against the papers, especially before citing it in research or making a consequential decision.

## What the published evaluation establishes

Ai2 reports that, across the Asta pipeline, Fast mode averaged **51.1 seconds per report** against **178.5 seconds** for Thinking mode, about 3.5 times faster in that comparison. Its model card reports AstaBrief's results on the ScholarQA-CS2 computer science questions and DeepScholarBench, and the release describes a separate, small human comparison: three researchers contributed 14 questions. Those are Ai2's evaluations of its systems and test sets, not an independent demonstration that any particular report will be accurate. The release says most training and evaluation took place in 2025 and the full evaluation was **not rerun against current frontier models**.

The [final checkpoint](https://huggingface.co/allenai/AstaBrief_8B) is public under Apache 2.0, and Ai2 lists its training datasets and prompts in an [AstaBrief collection](https://huggingface.co/collections/allenai/astabrief). The model card says the model is intended for research and education under Ai2's responsible-use guidance. Its training note mentions eight H100 GPUs for DPO training; that is **not** a published minimum for inference. The reviewed primary materials do not state a per-report price for hosted Asta, a minimum local GPU specification or a dependable local cost per report. Anyone planning deployment needs to establish those costs for their own setup.

## Sources & further reading

- [Ai2's October 2 release](https://huggingface.co/blog/allenai/astabrief) explains the live Fast mode, the one-pass design, timing and evaluation limits. Its performance claims are Ai2's own.
- [AstaBrief 8B model card](https://huggingface.co/allenai/AstaBrief_8B) gives the license, intended use, prompt recommendation and inference example, including the SFT checkpoint name in that example.
- [AstaBrief collection](https://huggingface.co/collections/allenai/astabrief) lists the published checkpoints, datasets and prompts; their presence does not establish an end-to-end local reproduction.
- [ScholarQA code and documentation](https://github.com/allenai/ai2-scholarqa-lib) describe the retrieval system; [the lite generator](https://github.com/allenai/ai2-scholarqa-lib/blob/main/api/scholarqa/lite/scholar_qa_lite.py) shows how reranked references become a one-pass generation prompt. We inspected documentation and code; we did not run them.
- [ScholarQA-CS2 evaluation repository](https://github.com/allenai/ai2-scholarqa-eval) provides benchmark code and data, including rubric construction. It helps explain the test design; it does not independently verify AstaBrief's reported scores.

## Sources

- [Open-sourcing AstaBrief, the fast report-generation model in Asta](https://huggingface.co/blog/allenai/astabrief) — Ai2 release. Documents Fast mode, one-pass generation, reported speed and benchmark caveats. Claims are vendor-reported.
- [AstaBrief-8B model card](https://huggingface.co/allenai/AstaBrief_8B) — Official checkpoint card: Apache 2.0, prompt and inference example, limits, training compute. Its code example names the SFT checkpoint.
- [AstaBrief collection](https://huggingface.co/collections/allenai/astabrief) — Official index of final/SFT checkpoints, datasets and prompts. Individual dataset terms need separate review.
- [Ai2 Scholar QA repository](https://github.com/allenai/ai2-scholarqa-lib) — Official repository explains retrieval and reranking; may not describe exact hosted Asta configuration.
- [ScholarQALite generator implementation](https://github.com/allenai/ai2-scholarqa-lib/blob/main/api/scholarqa/lite/scholar_qa_lite.py) — Official code shows reranked references into one-pass generation. Inspected, not executed.
- [Ai2 ScholarQA Evaluation repository](https://github.com/allenai/ai2-scholarqa-eval) — Ai2 benchmark code/data and rubric construction context; not an independent AstaBrief rerun.
