StandardsAboutContact
The Weights
AstaBrief-8B Teardown: Ai2's Open 8B Report Writer Posts

AstaBrief-8B Teardown: Ai2's Open 8B Report Writer Posts

Ai2 has released a Qwen3-8B fine-tune that writes a full cited research report in one pass, averaging 51.1 seconds in Fast mode. The model card's 78.2% citation recall is the figure to check before you trust it.

Use it, for cited research briefs where seconds matter. AstaBrief-8B, an Apache 2.0 Qwen3-8B fine-tune from Ai2, scores 87.0 on the ScholarQA-CS2 test set with 90.5% citation precision, and Fast mode averages 51.1 seconds per report versus 178.5 for Thinking mode. Citation recall of 78.2% means some claims arrive unsupported.

The Weights Desk · 4 min read

The bottom line: Use it, with one condition. Ai2 released AstaBrief-8B on October 2, 2026, an open 8B-parameter model that writes a full cited research report in one pass, and citation recall is the number that decides. The model card reports 90.5% citation precision but only 78.2% citation recall on the ScholarQA-CS2 test set. The speed gain is real by Ai2's account, with Fast mode averaging 51.1 seconds per report against 178.5 seconds for Thinking mode.

What shipped, and under what terms

AstaBrief-8B is a fine-tune of Qwen3-8B released under Apache 2.0, with weights and training data published on Hugging Face. Ai2's post says it is built to generate a whole report in one pass, not section by section. The model card limits the intended use to research and educational work under Ai2's Responsible Use Guidelines. It also says best results require the recommended prompt format, and that other prompts may degrade or destabilize behavior. That is a real integration constraint.

How it was trained

Training had two stages. Ai2 reports supervised fine-tuning on 47,000 examples drawn from 90,000 filtered research queries, then direct preference optimization on about 6,000 preference pairs. The model card says the pairs came from real user queries, with paired reports judged by both GPT-4.1 and DeepSeek-R1. Using two judges reduces the chance that a single grader's bias shapes the preferences. Neither source gives a per-stage ablation in the material we reviewed, so the contribution of each stage is not independently established.

The numbers that decide it

The model card reports an average score of 87.0, citation precision of 90.5% and citation recall of 78.2% on the ScholarQA-CS2 test set. Precision says that when the model cites, the citation is usually accurate. Recall says how much of what it claims is backed by a citation. A 12-point gap means some claims arrive unsupported, so a human has to check the claims that carry weight before the report is relied on.

Latency and what users did with it

Ai2 states that Fast mode averages 51.1 seconds per report, against 178.5 seconds for Thinking mode, about 3.5 times faster. Usage data is thin but useful: among 374 Asta users, 29.1% used it on two or more days, and 23% of those who tried Fast mode kept using it. A small human study had three researchers rate 14 questions, and two preferred AstaBrief on citation accuracy. That sample is too small to settle the question.

Limits of the evidence

Ai2 itself cautions that its results are best read as evidence about the training and system design choices it tested, not as a comparison with current frontier models, because the evaluation was run in 2025. The benchmark is also Ai2's own, and the post and model card name it differently (SQABench-CS2 and ScholarQA-CS2). Anyone adopting the model should rerun it on their own question set, with their own judge, before committing.

The bottom line: Use it, with a verification step

Use it for first-draft literature briefs where a human checks the cited origins, since an open 8B model with 90.5% citation precision and about 51 seconds per report is cheap to run and easy to audit. Wait if your workflow needs every claim supported, because 78.2% recall does not meet that bar. Skip it for non-research writing, since the card ties good behavior to a specific prompt format and a scientific-literature task. The deciding number is 78.2% citation recall.

What is AstaBrief-8B?
It is an 8-billion-parameter model fine-tuned from Qwen3-8B to write a cited report from a research question and excerpts of scientific literature. Ai2 released it on October 2, 2026 under Apache 2.0.
How fast is it?
Ai2 reports that Fast mode averages 51.1 seconds per report, against 178.5 seconds for Thinking mode. That is about 3.5 times faster.
How reliable are its citations?
The model card reports 90.5% citation precision and 78.2% citation recall on the ScholarQA-CS2 test set. Precision measures whether a cited source is accurate. Recall measures whether claims are supported by citations. A reviewer should still verify claims that matter.
What is the bottom line on adopting it?
Use it for first-draft cited literature briefs with a human verification step. Wait if every claim must be supported, and skip it for general writing, since the model card ties good behavior to a specific prompt format.
  1. Open-sourcing AstaBrief, the fast report-generation model in Asta — Hugging Face (Ai2)
  2. allenai/AstaBrief_8B model card — Hugging Face (Ai2)
  3. Asta — Ai2