A schema-enforced large language model framework produces largely reproducible decision codes in simulated sarcoma tumor boardsDownload PDF Download PDF ArticleOpen accessPublished: 11 October 2026Tekoshin Ammo1 na1,Moritz Englich1 na1,Fabio Dos Santos Adrego1,Maximilian Jacobi1,Alida Wilckens1,Philipp Kruppa1,Bastian Bonaventura1,Stefan Niemuth2,Selina M. Weiler1,Victoria Wachenfeld-Teschner1,Anja M. Boos1 &…Gerrit Freund1 Scientific Reports (2026) Cite this articleSave articleView saved research We’re sharing this article early to provide faster access to peer-reviewed, accepted research. It is citable and carries a permanent DOI. This version is subject to further edits and will be replaced automatically by the final Version of Record. All legal disclaimers apply.AbstractLarge language models require reproducible outputs for systematic evaluation in tumor boards. We developed a sarcoma tumor-board simulator using Claude Opus 4.5 at temperature 0.0, schema-enforced JSON output, and prompt-specified semantic and conditional-logic rules. The framework was evaluated on 51 de-identified cases across 154 runs. Reproducibility was defined as agreement across all runs of a case per decision code. In the 41 cases not used during development, reproducibility was 94.2% (cluster-bootstrap 95% CI 91.9–96.1); in the 10 development cases, 93.3% (95% CI 88.6–97.6). Pooled, it was 94.0% (95% CI 92.0–95.8); 17 cases agreed on all 21 codes. Estimates are model-specific. A post hoc ablation on 15 cases compared the framework with and without these rules: 94.6% versus 82.5% (difference 12.1 points; 95% CI 7.6–17.1). The improvement was driven mainly by exact wording agreement in three free-text fields (+ 60.0 points), whereas the 18 categorical and numeric codes differed by 4.1 points (95% CI − 0.7 to 9.3). Without any schema, a decision was stated for only 37.6% of instances versus 100% under schema enforcement; seven codes were never addressed. Single-reviewer screening detected no fabricated numbers, references, or measurements. Clinical correctness was not assessed. Consistent recommendations may still be incorrect.AcknowledgementsThe authors thank the institutional sarcoma multidisciplinary tumor board for prior case discussion underpinning the case set used in this methodology study.FundingThe authors declare that no external funding was received for the research, authorship, or publication of this article.Author informationAuthor notesTekoshin Ammo and Moritz Englich contributed equally to this work.Authors and AffiliationsDepartment of Plastic, Reconstructive and Aesthetic Surgery — Hand Surgery and Burn Center, University Hospital Schleswig-Holstein, Campus Lübeck, Campus Lübeck, GermanyTekoshin Ammo, Moritz Englich, Fabio Dos Santos Adrego, Maximilian Jacobi, Alida Wilckens, Philipp Kruppa, Bastian Bonaventura, Selina M. Weiler, Victoria Wachenfeld-Teschner, Anja M. Boos & Gerrit FreundDepartment of Anaesthesiology and Intensive Care Medicine, University Hospital Schleswig-Holstein, Campus Lübeck, GermanyStefan NiemuthAuthorsTekoshin AmmoView author publicationsSearch author on:PubMed Google ScholarMoritz EnglichView author publicationsSearch author on:PubMed Google ScholarFabio Dos Santos AdregoView author publicationsSearch author on:PubMed Google ScholarMaximilian JacobiView author publicationsSearch author on:PubMed Google ScholarAlida WilckensView author publicationsSearch author on:PubMed Google ScholarPhilipp KruppaView author publicationsSearch author on:PubMed Google ScholarBastian BonaventuraView author publicationsSearch author on:PubMed Google ScholarStefan NiemuthView author publicationsSearch author on:PubMed Google ScholarSelina M. WeilerView author publicationsSearch author on:PubMed Google ScholarVictoria Wachenfeld-TeschnerView author publicationsSearch author on:PubMed Google ScholarAnja M. BoosView author publicationsSearch author on:PubMed Google ScholarGerrit FreundView author publicationsSearch author on:PubMed Google ScholarCorresponding authorCorrespondence to Tekoshin Ammo.Ethics declarationsCompeting interestsThe authors declare no competing interests.Ethics statementThis retrospective methodological pilot study used fully de-identified case texts and was carried out in accordance with relevant guidelines and regulations, including the principles of the Declaration of Helsinki. The study was reviewed by the Ethics Committee of the University of Lübeck (Ethikkommission der Universität zu Lübeck), University Hospital Schleswig-Holstein (UKSH), Campus Lübeck, Germany (reference number / Aktenzeichen 2025 − 520). The need to obtain ethical approval was waived by the Ethics Committee of the University of Lübeck, given the retrospective design and the fully de-identified nature of the data. The requirement for informed consent was also waived by the Ethics Committee of the University of Lübeck for the same reason. No patients were treated, contacted, or recruited as part of this methodological study; all analyses were performed on de-identified textual records of cases that had previously been presented to the institutional sarcoma multidisciplinary tumor board as part of routine clinical care.Additional informationPublisher’s noteSpringer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.Supplementary InformationGlossaryJSON (JavaScript Object Notation)a lightweight, human-readable data-interchange format used here for the structured tumor-board output.Pydantic / Pydantic validationa Python data-validation library (version 2.13.3 in this study) used here to check declared field types and the allowed values of enumerated fields at run time. The schema used in this study contains no custom validators, so no cross-field or conditional rules were enforced programmatically; the conditional rules described in the Methods were specified in the system prompt only. The library was used in its default (lax) mode, which coerces compatible types rather than rejecting them.Tool-use APIan interface provided by the LLM vendor that allows the caller to define tools and their input schemas. Whether conformity is guaranteed during generation depends on the configured API mode; application-level validation checks returned data against the schema.Zero-shot promptinga prompting strategy in which the model is asked to perform a task without being shown completed input/output examples in the prompt.Few-shot semantic anchoringa prompting strategy in which a small number of brief examples are provided to clarify the intended semantics of specific fields, without giving full output exemplars.Reproducibility (in this study)the property that identical inputs yield identical outputs; quantified here as the proportion of decision codes whose values are identical across all available runs of the same case (three for 50 cases and four for case 11).Hallucination (LLM)the generation by an LLM of content not derivable from the input — for example, fabricated guideline citations, invented prognostic statistics, or unauthorized output sections (12).Rights and permissionsOpen Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.Reprints and permissionsAbout this article