A joint mixture Tobit method with latent microbial abundance improves detection of microbiome–disease associations in zero-inflated data

Wait 5 sec.

by Junrong Deng, Dong Chen, Siyuan Shen, Yuyang Zhou, Hanyu Cui, Yu Qiu, Qiming Li, Yue-Qing HuZero inflation remains a major challenge in microbiome differential abundance analysis, often resulting in inflated type I error rates or reduced statistical power. Although numerous methods have been proposed to address excess zeros, many existing approaches do not explicitly distinguish between distinct zero-generating mechanisms, and censoring-based modeling perspectives remain relatively underexplored in microbiome data analysis. To advance methodological development in this area, we introduce a novel censoring-based modeling perspective for microbiome differential abundance analysis. Specifically, we develop a joint mixture Tobit (joint mTobit) method for zero-inflated microbiome data. Building on a mixture Tobit formulation, the model incorporates a point-mass component to represent structural zeros, while modeling latent microbial abundance via a Tobit regression component. The proposed framework jointly links disease status, latent true microbial abundance, and relevant covariates within a unified probabilistic model, allowing appropriate adjustment for confounding factors and facilitating more reliable statistical inference. By explicitly modeling the latent abundance underlying observed counts, the joint mTobit method improves estimation stability and enhances detection power under zero inflation. Extensive simulation studies demonstrate that the joint mTobit method achieves effective type I error control while maintaining high statistical power and stable coefficient estimation across a wide range of settings. Application to a real-world colorectal cancer and adenoma microbiome dataset further illustrates its ability to identify biologically meaningful differentially abundant taxa. Overall, this work develops a joint mTobit modeling framework for zero-inflated microbiome data, enabling inference on latent microbial abundance and providing a censoring-based statistical framework for investigating disease–microbiome associations in differential abundance analysis.