NEWS AND VIEWS07 October 2026Most LLMs cannot reliably evaluate text on the level of individual letters. A technique called byteification retrofits existing models to enable it.ByZhao Zhang0 &Yingfei Xiong1Zhao ZhangZhao Zhang is in the Key Lab of High Confidence Software Technologies, School of Computer Science, Peking University, Beijing 100871, China.View author publicationsSearch author on: PubMed Google ScholarYingfei XiongYingfei Xiong is in the Key Lab of High Confidence Software Technologies, School of Computer Science, Peking University, Beijing 100871, China.View author publicationsSearch author on: PubMed Google ScholarSave articleView saved researchStrawberry contains the letter ‘r’ three times, but when asked, many large language models (LLMs) answer that the letter appears twice. This happens because most LLMs encode words as ‘tokens’ that represent sequences of letters. LLMs that operate in this way can achieve excellent performance, but they cannot access the individual characters in each word, which are encoded as binary sequences called bytes. Writing in Nature, Minixhofer et al.1 now report an approach called byteification that retrofits token-based LLMs to operate at the byte level. The authors show that byteified models can achieve competitive performance while retaining the ability to read individual characters.Access options Access through your institutionAccess Nature and 54 other Nature Portfolio journalsGet Nature+, our best-value online-access subscription27,99 € / 30 dayscancel any timeLearn moreSubscribe to this journalReceive 52 print issues and online access199,00 € per yearonly 3,83 € per issueLearn moreRent or buy this articlePrices vary by article typefrom$1.95to$39.95Learn morePrices may be subject to local taxes which are calculated during checkoutdoi: https://doi.org/10.1038/d41586-026-03059-2ReferencesMinixhofer, B. et al. Nature https://doi.org/10.1038/s41586-026-11111-4 (2026).Article Google Scholar Sennrich, R., Haddow, B. & Birch, A. in Proc. 54th Ann. Meet. Assoc. Comput. Linguist. 1715–1725 (2016).Cosma, A., Ruseti, S., Radoi, E. & Dascalu, M. in Proc. 2025 Conf. Empir. Meth. Nat. Lang. Proc. 28252–28263 (2025).Pagnoni, A. et al. Proc. 63rd Ann. Meet. Assoc. Comput. Linguist. 9238–9258 (2025).Download referencesCompeting InterestsThe authors declare no competing interests.Related Articles Read the paper: Retrofitting language models to operate over bytes Expert-level test is a head-scratcher for AI Algorithm that gets ‘under the hood’ of AI models could effectively steer their responsesSee all News & ViewsSubjectsMachine learningLatest on:Machine learningJobs Open-Rank Faculty Position in Theoretical BiologyThe Department of Biological Sciences seeks an outstanding individual for an open-rank faculty position in Theoretical Biology.Nashville, TennesseeVanderbilt University Department of Biological SciencesProteomics SpecialistBuild the science that shapes the future of human health. Application closing date: 08 November 2026 Join a place where ambitious science thrives...Milan (IT)Human TechnopoleTenure-Track Professor/Associate Professor/Assistant ProfessorThe University of Hong Kong We are seeking a motivated Tenure-Track Professor/Associate Professor/Assistant Professor to join our Department of Pha...Hong Kong (HK)The University of Hong Kong