CellExLink: End-to-end cell-type recognition and normalization in biomedical text

Wait 5 sec.

by Alimire Nabijiang, Leili ShahriyariCell types are described in biomedical literature using diverse names, abbreviations, and phenotype phrases, which complicates their recognition and normalization. We developed CellExLink, an end-to-end pipeline that identifies cell-type mentions and normalizes them to Cell Ontology (CL) identifiers. The recognizer was fine-tuned and evaluated on five heterogeneous biomedical corpora spanning full-length articles, article excerpts, figure captions, abstracts, and anatomical text passages. These resources include fine-grained phenotype-defined populations, heterogeneous cell populations, and abbreviated mentions. Across the five corpora, CellExLink achieved macro-average exact- and relaxed-span F1 scores of 0.766 and 0.855, respectively. For CL identifier normalization on gold-standard mention spans, F1 scores ranged from 0.690 to 0.874. In strict end-to-end evaluation, which required both an exact mention span and the correct CL identifier, F1 scores ranged from 0.552 on a figure-caption corpus to 0.725 on a corpus of full-text article excerpts. CellExLink outperformed the evaluated off-the-shelf systems in cell mention recognition, CL identifier normalization, and end-to-end extraction. By converting unannotated biomedical text into cell-type spans linked to standardized CL identifiers, CellExLink provides a practical foundation for downstream applications, including literature curation, relation extraction, and knowledge graph construction.