Who checks what the AI threw out?
Two papers published on 27 August 2026 come at the same problem from opposite directions. Together they make the most-quoted number in LLM screening research much harder to use.
The first, from Lai and colleagues in Research Synthesis Methods, is an argument about arithmetic. Papers on LLM-assisted title-and-abstract screening now routinely report workload savings between 79% and 91%; the figures reported by Sciurti and colleagues for compact models are the specific target here. Lai's objection is that this metric is calculated retrospectively, from the classifier's performance on a fully labelled dataset, and that it quietly assumes something no review team should assume: that records the model excluded are never looked at again.
Anyone who has actually run a screening stage knows this is not how it works. Somebody spot-checks the excludes. A second reviewer samples them. The clinical lead asks to see everything the model threw out below a certain confidence score. Each of those choices is a real cost, and none of it appears in the published saving.
So Lai and colleagues propose an alternative: Verification-Adjusted Workload Reduction, parameterised by the proportion of model-excluded records a human still verifies. That parameter is the one worth remembering. On identical data, the same model swings from roughly 0% to 91% saving depending purely on how much of the excluded pile gets a second look.
That range is the finding. It means the workload saving is not a property of the model at all. It is a property of the model combined with a governance decision, and the governance decision is usually the one nobody has made.
The second paper, arriving at the same place from the other side
Also on 27 August, BMC Medical Research Methodology published a screening study using a five-model Delphi ensemble over 4,745 abstracts. Recall reached 98%, which is the number most teams would care about first. Work saved over sampling at 95% recall came in at 13–42%.
High recall, modest relief. The two figures sit uncomfortably together, and the gap between them is roughly the space that marketing copy occupies.
One incidental result from that study deserves more attention than it will get: the simplified prompt outperformed the elaborate ones. Prompt engineering in screening has developed a folklore of long, heavily structured instructions. At least in this dataset, the folklore lost.
What a team should be able to answer
If you are planning a review that uses LLM screening, four questions decide what your timeline actually looks like, and none of them are about the model.
What proportion of model-excluded records will a human verify? This is the parameter. If the answer is "we hadn't decided", the saving figure in your grant application is unsupported.
Is that verification sample random, or targeted at borderline confidence scores? Targeted sampling is more efficient per record checked, but it changes what the sample tells you. A random sample estimates the miss rate. A borderline-band sample estimates the miss rate in the band, which is not the same quantity and cannot be reported as if it were.
How are model–human disagreements resolved, and by whom? Any disagreement rule that defaults to the human is a rule that imports human error rates back into a process advertised on model performance.
Who signs off? A verification policy nobody owns tends to shrink under deadline pressure, which is exactly when the record of it matters.
Some arithmetic to keep the scale honest. On a search returning 14,000 records with a 4% inclusion rate at title and abstract, the model-excluded pile is somewhere near 13,000 records. A 10% verification fraction is 1,300 abstracts to screen by hand. That is a fortnight of a reviewer's time, and it is the difference between a saving you can bank and one you can only publish.
Which model you use is a methodological decision
A third paper from August is worth reading alongside these. In BMC Medical Research Methodology on 18 August, Claude Sonnet 4.5, Gemini 3.0 Thinking and ChatGPT 5.1 each reviewed 50 anonymised landmark oncology trials. Their verdicts differed significantly (p<0.001), with Claude issuing major-revision decisions in 74% of cases and rejecting 6%, and Gemini the most permissive of the three. Recognition of the trials ranged from 14% to 54% across models.
That study is about peer review, not screening. But the model-to-model variance finding transfers directly. If three frontier models disagree that sharply on a bounded judgement task, the choice of model is a methodological decision on the same footing as choosing a risk-of-bias tool, and it needs pre-specifying and reporting for the same reasons.
A preprint benchmark posted on arXiv in mid-August (2608.12741) tested GPT-5, Claude Sonnet 4, Gemini 2.5 Pro and NotebookLM across screening, extraction, analysis and synthesis on 244 documents. No system led on every task. Claude scored highest on screening accuracy at 82.8%; GPT-5 recorded the best recall at 91.8%, at a cost in specificity. Performance fell off most on interpretive analysis and synthesis across sources, which is the part of a review that a reader is actually paying for.
And a medRxiv preprint from 21 August, MetaFemina, is the useful counterexample for anyone asking whether the whole pipeline can be automated. It runs PubMed retrieval, LLM extraction and statistical synthesis end to end for nutritional exposures and gynaecological cancers. Benchmarked against two published meta-analyses, it recovered the included studies at 81.8% and 80.0% sensitivity. Those are respectable engineering numbers and they are well below any recall threshold a review would accept at screening.
Two practical consequences
First, version and date everything. Nested Knowledge upgraded the models behind its entire AI stack on 23–24 August, reporting roughly 50% faster performance with improved instruction-following and modest recall gains. Anyone who benchmarked that platform before 23 August is holding a result about software that no longer exists. Vendor model upgrades are silent, frequent, and they invalidate your pilot.
Second, write the verification policy into the protocol rather than the methods section. A verification fraction chosen after screening is a description of what happened. Chosen before, it is a design parameter, and it is the only version that supports a claim about workload.
One deadline worth knowing: the joint AI-disclosure reporting standard from STM, COPE, the International Science Council and the Global Young Academy is in its second consultation round, closing 16 October 2026. It covers disclosure thresholds, where and how disclosure appears, taxonomy, and what documentation of accountability should look like. A third round follows in late 2026 or early 2027. If your work involves AI-assisted synthesis, this is the consultation that will shape what you are required to declare.
References
Delporte, M., Tamimi, R., Mehta, S., Choi, E., Zhang, Y., & Shi, Y. (2026). MetaFemina: Development and evaluation of a large language model-assisted platform for automated meta-analysis of nutritional exposures and breast, ovarian, and uterine cancer risk [Preprint]. medRxiv. https://doi.org/10.64898/2026.08.18.26360713
Erdat, E. C., & Utkan, G. (2026). AI as gatekeeper: A cross-sectional study of AI-generated peer review of landmark oncological studies. BMC Medical Research Methodology, 26, Article 2978. https://doi.org/10.1186/s12874-026-02978-y
International Science Council. (2026). AI disclosure in research: Consultation on a global reporting standard. https://council.science/our-work/ai-disclosure-in-research/
Lai, H., Liu, J., Ge, L., & Estill, J. (2026). Reframing workload saving in LLM-assisted screening: A commentary on Sciurti et al. Research Synthesis Methods, 1-5. Advance online publication. https://doi.org/10.1017/rsm.2026.10114
Nested Knowledge. (2026). Release notes: Versions 1.115.0 and 1.115.1. https://about.nested-knowledge.com/docs/releases/
Sciurti, A., Migliara, G., Siena, L. M., Isonne, C., De Blasiis, M. R., Sinopoli, A., Iera, J., Marzuillo, C., De Vito, C., Villari, P., & Baccolini, V. (2026). Compact large language models for title and abstract screening in systematic reviews: An assessment of feasibility, accuracy, and workload reduction. Research Synthesis Methods, 17(2), 332-347. https://doi.org/10.1017/rsm.2025.10044
Shafqat, W., Patterson, M., & Liss, S. N. (2026). Knowledge Synthesis Review Framework: Task-level benchmarking of LLM-based systems for multi-source evidence synthesis (arXiv:2608.12741). arXiv. https://arxiv.org/abs/2608.12741
Tolend, M., Halabi, R., Ghaouari, K., Lau, Y. C. Y., Alda, M., Hintze, A., Mulsant, B. H., & Ortiz, A. (2026). Impact of prompt engineering and workflow modifications on the performance of a Delphi-based automated abstract screening system in a psychiatric systematic review. BMC Medical Research Methodology, 26, Article 2981. https://doi.org/10.1186/s12874-026-02981-3
