Deep learning framework for automated cervical vertebral maturation staging using CBCT scansDownload PDF Download PDF ArticleOpen accessPublished: 23 September 2026Omid Halimi Milani1,Amanda Nikho2,Lauren Mills2,Marouane Tliba1,Veerasathpurush Allareddy2,Rashid Ansari1,Ahmet Enis Cetin1,Ahmet Atakan1 &…Mohammed H. Elnagar ORCID: orcid.org/0000-0002-2332-92502,3,4 Scientific Reports (2026) Cite this articleSave articleView saved research We’re sharing this article early to provide faster access to peer-reviewed, accepted research. It is citable and carries a permanent DOI. This version is subject to further edits and will be replaced automatically by the final Version of Record. All legal disclaimers apply.AbstractAutomated skeletal maturity assessment is essential in orthodontic diagnosis and treatment planning, guiding the optimal timing of interventions such as myofunctional appliances and orthognathic surgery. This study aims to develop and evaluate deep learning algorithms for the assessment and classification of cervical vertebral maturation (CVM) stages using sagittal views derived from cone beam computed tomography (CBCT) scans. The sample consisted of 364 CBCT scans of orthodontic patients from private practices in the midwestern United States. The CVM stages were classified by two orthodontists and an oral and maxillofacial radiologist. Distinct experimental settings evaluated image-based classification, prompt-guided anatomical localization using the Segment Anything Model (SAM), demographic conditioning with age and sex information, knowledge distillation from a teacher network pretrained on a skeletal maturity dataset, and tri-modal input fusion. Because data availability differed across settings, these components were evaluated on different eligible subsets rather than as a unified sequential ablation. In the full 364-sample six-class baseline setting, ResNet34 achieved 66.76% accuracy. In the separate 96-sample, stage-aligned five-class knowledge-distillation setting, ResNet50 achieved the highest observed accuracy of 88.54% (5-fold cross-validation average), compared with 87.10% for its non-distilled counterpart in the stage-aligned five-class knowledge-distillation setting using a paired subset of 96 scans. The distillation effect varied across architectures, and the 96-sample subset was small and class-imbalanced. The findings support the feasibility of deep learning-based CVM staging from CBCT-derived sagittal views across distinct experimental settings. The results do not establish that anatomical guidance, demographic conditioning, and knowledge distillation jointly improve classification because these components were not evaluated through a unified ablation on a common cohort. Within the separate five-class knowledge-distillation setting, performance changes were backbone-dependent, with only a modest numerical improvement observed for ResNet50. The findings should therefore be considered preliminary and internally validated. Larger, balanced, and externally validated cohorts are required before broader clinical applicability can be established.FundingThis study was supported by the American Association of Orthodontists Foundation (AAOF) through funding awarded to M.H.E. The funder had no role in the study design, data collection and analysis, decision to publish, or preparation of the manuscript.Author informationAuthors and AffiliationsDepartment of Electrical and Computer Engineering, University of Illinois Chicago, Chicago, Illinois, USAOmid Halimi Milani, Marouane Tliba, Rashid Ansari, Ahmet Enis Cetin & Ahmet AtakanDepartment of Orthodontics, College of Dentistry, University of Illinois Chicago, Chicago, Illinois, USAAmanda Nikho, Lauren Mills, Veerasathpurush Allareddy & Mohammed H. ElnagarDepartment of Orthodontics, Faculty of Dentistry, Tanta University, Tanta, EgyptMohammed H. ElnagarDepartment of Orthodontics (M/C 841), College of Dentistry, University of Illinois Chicago, 801 S. Paulina Street, RM 131, 60612-7211, Chicago, IL, USAMohammed H. ElnagarAuthorsOmid Halimi MilaniView author publicationsSearch author on:PubMed Google ScholarAmanda NikhoView author publicationsSearch author on:PubMed Google ScholarLauren MillsView author publicationsSearch author on:PubMed Google ScholarMarouane TlibaView author publicationsSearch author on:PubMed Google ScholarVeerasathpurush AllareddyView author publicationsSearch author on:PubMed Google ScholarRashid AnsariView author publicationsSearch author on:PubMed Google ScholarAhmet Enis CetinView author publicationsSearch author on:PubMed Google ScholarAhmet AtakanView author publicationsSearch author on:PubMed Google ScholarMohammed H. ElnagarView author publicationsSearch author on:PubMed Google ScholarCorresponding authorCorrespondence to Mohammed H. Elnagar.Ethics declarationsCompeting interestsThe authors declare no competing interests.Ethics approval and informed consentThis research was a retrospective study utilizing archived data. All data used in the study were deidentified to ensure confidentiality. The University of Illinois Chicago’s Office for the Protection of Research Subjects (OPRS) Institutional Review Board (IRB) granted an exemption from the requirement for informed consent and provided IRB-exempt status to the study under Study ID STUDY2022-1048. All methods were performed in accordance with the relevant guidelines and regulations of the University of Illinois Chicago’s OPRS and IRB.Additional informationPublisher's noteSpringer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.AppendixAppendixFigures 10 and 11 provide the row-normalized confusion matrices for different backbone architectures evaluated on the CVM dataset. The first set of matrices corresponds to the baseline models trained solely on the original images, while the second set illustrates the performance when age and sex positional encodings are incorporated alongside masked inputs. These visualizations highlight the distribution of correct and incorrect predictions across CVM stages for each architecture.Fig. 10Full size imageConfusion matrices (row-normalized percentages) of the base models on the original CVM dataset. Each panel corresponds to a different backbone architecture.Fig. 11Full size imageConfusion matrices (row-normalized %) for models with combined age and sex positional encoding (Age+Sex PE) on the original CVM dataset along with the masked images. Each panel corresponds to a different backbone architecture.Rights and permissionsOpen Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.Reprints and permissionsAbout this article