An auditable evidence compiler for large language model-assisted systematic reviews

Wait 5 sec.

Background: Large language models (LLMs) can support systematic reviews, but accurate individual outputs do not establish whether the final synthesis preserves the clinical question, accounts for statistical dependence and incorporates corrections. Objective: To develop and evaluate a framework linking LLM-assisted evidence processing to a versioned, auditable release of synthesis outputs. Methods: We used LLM agents to interpret sources and extract data. Deterministic code enforced statistical rules; investigators resolved material ambiguities and authorized release. We specified ten release properties covering evidence identities, statistical contributions and propagation of corrections. We retrospectively evaluated six integrity domains and historical failure events in one registered prognostic review, without an external comparator or held-out domain. Results: Fifty distinct root-cause events were documented, including 12 that had changed a pooled result before correction. Forty-six were resolved, and four remained disclosed limitations. The corpus comprised 454 reports, 445 studies, 441 cohort entities and 421 dependence clusters. Forty-one of 49 registered analyses were fitted, and eight retained explicit non-fitted states. All 39 source records across five principal analysis families reached a terminal source state. Two implementations within the project agreed across 1,217 numerical comparisons. All 94 file comparisons between release and publication packages were byte-identical. Two reviewers confirmed 39 principal records after seeing the same recommendations. Conclusions: This case provides a framework for inspecting synthesized evidence together with its provenance, statistical meaning and correction history. Comparative validity, generalizability and benefit in patient-centred care require independent evaluation.