Cell-type-specific eQTLs underlie the genetic architecture of complex traits

Wait 5 sec.

MainThe challenge of connecting genetic effects on complex traits with their underlying gene regulatory mechanisms has motivated population-scale studies of gene expression using RNA-sequencing (RNA-seq) in bulk tissues, such as GTEx2. Although these studies have identified thousands of expression quantitative trait loci (eQTLs) that contribute to variation in transcript abundance, at present, known eQTLs only explain about 10–30% of genetic effects on complex traits5,6. This suggests that gene regulatory effects on complex traits are likely to depend on contexts that are not evident in bulk tissues, such as cell types, environments or developmental stages. Here, we sought to characterize the role of eQTL cell-type-specificity on complex traits.Previous studies have focused on eQTLs that reach statistical significance in bulk tissues, a strategy that, due to limited power, identifies only a small fraction of suspected eQTLs. Furthermore, this strategy biases towards atypically large eQTL effects that are shared across all cell types within a tissue and are depleted in the genetic architecture of complex traits and relevant gene features3. Recently, eQTL studies using single-cell RNA-seq (scRNA-seq) have improved power to identify cell-type-specific eQTLs that are likely to contribute to gene regulatory effects on complex traits7,8,9,10,11. However, these studies also focus on statistically significant eQTLs, leaving the overall role of eQTL cell-type specificity unclear.We propose an alternative approach based on partitioning variation in scRNA-seq data. Rather than identifying individual eQTLs, our model unbiasedly quantifies the overall contributions of cell-type-shared and cell-type-specific eQTLs. Using this approach, we established the importance of cell-type-specific eQTLs in the genetic architecture of complex traits and partly explained why known eQTLs are depleted in functionally relevant gene features.CIGMA quantifies cell-type-specific eQTLsWe set out to characterize cell-type-specific eQTLs by combining recent population-scale scRNA-seq datasets, which greatly improve cell-type resolution over bulk tissues, with statistical genetic models that unbiasedly quantify genetic effects, similar to the genomic-relatedness-based restricted maximum-likelihood (GREML) model of complex trait heritability12.Our approach, cell-type-informed genetic mixed-model analysis (CIGMA), is shown in Fig. 1 (Methods). For a given gene, CIGMA quantifies the variance explained by cell-type-shared eQTLs, \({\sigma }_{{\rm{g}}}^{2}\); the variance explained by eQTLs specific to cell-type c, vc; and the overall eQTL specificity, which we define as \(\bar{v}/({\sigma }_{{\rm{g}}}^{2}+\bar{v})\), where \(\bar{v}\) averages vc across cell types. When all vc = 0, eQTLs are identical across cell types; otherwise, eQTLs are partly cell-type-specific, and we call the gene a cell-type-specific eGene (cs-eGene)9.Fig. 1: Study overview.Full size imagea, We used population-scale scRNA-seq data and paired genotypes as input. Typically, CIGMA is fit to common cis variants, but it can fit multiple genotype matrices, for example, cis and trans. b,c, CIGMA fits cell-type-shared and cell-type-specific eQTLs (αl and γlc) (b), and then outputs the variance explained by each group (\({\sigma }_{{\rm{g}}}^{2}\) and vc) and eQTL specificity \((\bar{v}/({\sigma }_{{\rm{g}}}^{2}+\bar{v}))\) (c). d, We used CIGMA outputs to characterize how cell-type-specific eQTLs relate to key genomic features, such as conserved genes and complex trait heritability.CIGMA has several key features for robust inference in scRNA-seq data (Methods and Supplementary Note 1). First, CIGMA accounts for cell-to-cell variation (δ) within each individual and cell type, including experimental noise. Second, CIGMA can jointly fit multiple random effects, such as cis and trans genomic regions or experimental batches. Third, our primary analyses assume that cell-type-specific eQTLs are independent across cell types for parsimony, but CIGMA can also fit a more realistic model allowing arbitrary eQTL correlations across cell types. Finally, CIGMA parameters are unbiasedly estimated with the method of moments and nonparametrically tested with jackknife.To assess the ability of CIGMA to quantify the variance explained by cell-type-specific eQTLs, we performed simulations based on peripheral blood mononuclear cells (PBMCs) from OneK1K, which has the most individuals of available scRNA-seq datasets10. After quality control, we included 10,288 genes, 928 individuals, 7 cell types and 1,190 cells per individual on average (Methods and Supplementary Fig. 1). We applied CIGMA to real cis SNPs (within 500 kilobases (kb) of the gene body) for each gene.We established that CIGMA is well calibrated in the absence of eQTLs by permuting genotypes across individuals (Fig. 2a). Next, we asked whether CIGMA is biased by cell-type-shared eQTLs by permuting cells across cell types. As expected, CIGMA detected the cell-type-shared eQTLs yet also provided calibrated null estimates for cell-type-specific eQTLs (Fig. 2a, Extended Data Fig. 1a and Supplementary Fig. 2). These results held across genes passing our quality control regardless of their total levels of expression or variance (Extended Data Fig. 1a). Interestingly, permuting subsets of cell types shows that eQTL specificity is reduced less by mixing cell types within the same lineage (Extended Data Fig. 1b). We found that the Wald test of CIGMA for significant cs-eGenes was calibrated when permuting individuals but slightly inflated when permuting cells, although this was negligible compared with real data signals (Supplementary Fig. 2 and Supplementary Table 1).Fig. 2: CIGMA yields unbiased and powerful estimates of cell-type-shared and cell-type-specific eQTLs in simulations.Full size imagea, Permutations of real OneK1K data across 10,288 genes. Genotype permutation breaks the connection between genotypes and expression, eliminating all eQTLs. Cell permutation across cell types within individuals renders cell types meaningless while preserving cell-type-shared eQTLs. About 10% of points lie outside the range (−0.2, 0.3) and are not shown. b, Pseudobulk simulations with 1,000 replicates show that CIGMA is unbiased and grows more precise with more individuals. c, CIGMA compared with simpler methods while varying the number of cells relative to OneK1K (which has approximately 1,190 cells per individual) across 1,000 replicates. ‘CIGMA without δ’ ignores cell-to-cell variation by assuming δ = 0 in equation (1) (Methods). ‘GCTA’ is used to fit the standard GREML model, which we apply either to pseudobulk per cell type (CTP) or to overall pseudobulk which combines cells across all cell types (OP). ‘GxEMM’ extends the GREML model to model GxE heritability, which we apply to OP using cell-type proportion as an ‘environment’. ‘BOLT-REML’ extends GREML to multiple traits, which we apply to each cell type as a ‘trait’. The OTD framework applies GREML after partitioning expression into cell-type-shared and cell-type-specific components. The white dots in a represent the medians. The boxes in b represent the first, second (median) and third quartiles, whereas the whiskers extend to values within 1.5 times the interquartile range from the first and third quartiles. Data points beyond the whiskers are outliers. The error bars in c represent the first, second (median) and third quartiles of estimates.Source dataWe next simulated data varying the number of individuals, N, and found that estimates were unbiased for all N and grew more precise as N grew (Fig. 2b and Methods). Power to detect cs-eGenes grew with N, rising from about 40% with N = 1,000 to about 100% with N = 2,000 (Supplementary Fig. 3). Power and precision also grew with the number of cells, the number of cell types and eQTL cell-type specificity (Extended Data Fig. 2, Supplementary Fig. 3 and Supplementary Table 2). We also found that CIGMA is robust to simulations that violate its assumptions on how eQTL effects are distributed across variants or cell types in theory and simulations (Extended Data Fig. 3, Supplementary Figs. 4 and 5 and Supplementary Notes 1.4 and 1.6) and to real genotype data (Supplementary Fig. 6).We compared CIGMA with simpler methods to quantify eQTLs. First, we tested a version of CIGMA that ignores cell-to-cell variation (Methods), which resulted in about 90% deflated estimates of eQTL variance explained with realistic data, although this bias vanished in the limit of infinite cells (Fig. 2c). Second, we tested GREML12, a common approach to quantify genetic effects on complex traits that neither distinguishes shared and specific eQTLs nor models cell-to-cell variation. With infinite cells, GREML unbiasedly estimated total eQTL variance when applied separately to each cell type (cell-type pseudobulk, CTP). However, when applied to a proxy for bulk RNA-seq that sums over cell types (overall pseudobulk, OP), GREML was biased in different directions depending on the number of cells; the mixture preserves cell-type-shared effects but partly cancels cell-type-specific effects, changing the relative strength of genetic and nongenetic variance13 (Methods). Third, we applied GREML to cell-type-shared and cell-type-specific components of expression by applying orthogonal tissue decomposition (OTD), which was downward biased with realistic data and remained biased with infinite cells14 (Methods). Fourth, we tested a multi-trait model, treating the expression of each cell type as a trait (BOLT-REML15), which is similar to CIGMA without cell-to-cell variation and gave similar results. Finally, we applied a GxE model to OP expression treating cell-type proportion as an ‘environment’ (GxEMM16), which also gave similar estimates on average, although they were much noisier because GxEMM uses less informative data (OP instead of CTP). GxEMM and BOLT-REML yielded biased specificity estimates due to unmodelled cell-to-cell variation (Supplementary Fig. 7). Overall, existing methods for complex traits cannot quantify cell-type-specific eQTLs in realistic scRNA-seq data because of cell-type-shared eQTLs and/or cell-to-cell variation.Cell-type-specific eQTLs in OneK1KWe analysed 928 individuals and the seven most common cell types from the OneK1K scRNA-seq dataset10. We applied CIGMA to common cis SNPs for each of 10,288 genes, defined as SNPs within 500 kb of the gene body with minor allele frequency (MAF) greater than 5% (ref. 14). We detected cell-type-specific eQTLs for 193 genes (P  0.95; Supplementary Fig. 19).Analysis of scRNA-seq data from CLUES and ImmVarWe applied CIGMA to the scRNA-seq dataset from the CLUES and the ImmVar9. After quality control in the original study, this dataset spans 1.2 million PBMCs from 162 patients with SLE and 99 healthy controls. Cells were clustered into 11 cell types: CD14+ classical monocytes (cM), CD16+ non-classical monocytes (ncM), conventional dendritic cells (cDC), plasmacytoid dendritic cells (pDC), CD4+ T cells (CD4), CD8+ T cells (CD8), NKs, B cells (B), plasmablasts (PB), proliferating T and NKs (Prolif), and progenitor cells (Progen). For our analysis, we excluded individuals without genotype data and focused on the three largest subgroups analysed in ref. 9: 70 controls of European ancestry, 70 patients of European ancestry and 75 patients of Asian ancestry.For CIGMA, we performed two types of analyses: (1) separate CIGMA analyses for each subgroup followed by meta-analysis; and (2) a mega-analysis jointly analysing all three subgroups in a single CIGMA analysis. In the first approach, we conducted quality controls independently for each subgroup using the same procedure as in OneK1K. After quality control, we retained 11,424, 10,842 and 10,786 genes in 70 European controls, 65 European patients and 74 Asian patients, respectively, involving the seven largest cell types: B, NK, CD4, CD8, cDC, cM and ncM. In CIGMA, we corrected for sex, age, cell processing batch and 10 PCs of OP expression as fixed effects and sequencing batch as a random effect. For age, similar to OneK1K, we categorized the cohort into 5-year intervals from 20 years to 70 years, with an additional group for individuals older than 70 years. Moreover, we corrected for three, four and four genotype PCs as fixed effects in European controls, European patients and Asian patients, respectively, based on elbows in the eigenvalue scree plots (Supplementary Fig. 41). Then, we conducted a meta-analysis on 10,553 genes common across all three subgroups. Cell-type-shared and cell-type-specific estimates from CIGMA were combined using inverse-variance weighting, with precision matrices estimated by jackknife resampling. For heritability and specificity, which have noisier precision estimates, we used sample-size weighting. Meta-analysed cell-type-specific genetic and residual interindividual effects were tested using Wald tests as in individual runs of CIGMA, using meta-analysed precision estimates across subgroups. In the mega-analysis, we analysed the same set of 10,553 genes across all 209 individuals. Apart from covariates used in the subgroup-based analyses, the mega-analysis adjusted for five genotype PCs, health state (case–control), cohort (CLUES–ImmVar) and ancestry.We use Pearson correlation across genes to measure replication across cohorts. Owing to estimation error, these correlations will be below 1 even when the true underlying parameters are identical across cohorts. To model this null, we draw CIGMA pseudo-estimates from Gaussian distributions, in which the mean for each gene is its weighted average across cohorts and the standard deviation for the estimate of each cohort is given by its real data standard error. We simulate 200 datasets per gene to calculate empirical P-values.Gene feature analysisTo investigate attributes related to cs-eGenes, we evaluated three gene-level features: LOEUF, enhancer count and connectedness in gene co-expression networks. LOEUF scores quantify the tolerance of a gene to loss-of-function mutations, serving as an approximate measure of selection strength acting on the gene. Genes with higher LOEUF scores are more tolerant to loss-of-function mutations and less conserved. We obtained LOEUF scores from the Genome Aggregation Database (gnomAD) v.2.1 (ref. 21). Enhancer counts reflect the regulatory complexity of a gene. We used counts from ref. 25, which were derived from enhancer–gene links identified through chromatin states and the association between histone modifications and gene expression levels62. Gene connectedness, as defined in refs. 3,63, was assessed by ranking genes based on their number of neighbours in co-expression networks constructed in ref. 64. Genes with higher connectedness are likelier to have regulatory effects on more genes.In OneK1K, we analysed 7,042 genes with positive total genetic variances (Fig. 4a). We confirmed these results using a complete set of 10,035 genes, which included genes with negative total genetic variances, and a refined subset of 6,578 genes, which excluded 464 genes with specificity s.e. exceeding 100 (Supplementary Fig. 23). In CLUES and ImmVar, we analysed 2,325 genes (Fig. 5d) after excluding genes with specificity s.e. exceeding 100 in any subgroup or total genetic variances below zero in meta-analysis. We also evaluated 3,075 genes defined solely by the latter criterion, which gave qualitatively similar results (Supplementary Fig. 34).LD score regressionWe used LD score regression (LDSC)65 to investigate the impact of cs-eGenes on complex diseases by adapting its approach for cell-type-specific heritability enrichment30. We included the top 200 genes with the most significant cell-type-specific genetic effects as defined by P-value. We defined the genomic annotation per gene as in the cis windows for CIGMA analyses (within 500 kb of the gene body). To validate our findings, we repeated the analysis using the top 100 and 300 genes and alternative window sizes of 300 kb and 700 kb.We tested three additional gene sets for comparison: (1) shared eGenes identified by CIGMA, with the most significant cell-type-shared genetic variance; (2) additively heritable genes identified by GCTA, with the most significant OP heritability; and (3) DEGs identified by CIGMA, with the largest variance in mean expression levels (μ). The DEG set was chosen based on variance instead of significance levels because most of the genes exhibited highly significant differential gene expression across cell types57. We chose to threshold eGenes into discrete sets, but larger datasets will enable modelling eQTL specificity as a continuous annotation.Our analysis included seven autoimmune diseases: ulcerative colitis, rheumatoid arthritis, primary biliary cirrhosis, multiple sclerosis, SLE, Crohn’s disease and celiac disease. As negative controls, which are less relevant to immune cells, we included height, coronary artery disease and schizophrenia. GWAS summary statistics for these diseases and traits were obtained from https://alkesgroup.broadinstitute.org/sumstats_formatted.In each analysis, LDSCs were computed using genotype data from the 1,000 Genomes Phase 3 European populations66, restricting the analysis to SNPs in HapMap 3 and using a window size of 1 centimorgan. We removed the major histocompatibility complex (MHC) region because of its unusual LD and genetic architecture. Apart from the baseline LDSC model v.1.2 (refs. 65,66), we conditioned on an annotation defined by all genes analysed in CIGMA so that our results are not merely biases from genes included in our study.To control for potential confounding due to gene length and expression level, we compared LDSC P-values from each of the four tested gene sets against matched control genes to achieve empirical P-values. Control genes were selected by ranking all genes by gene length and mean expression level (average OP across individuals) and identifying, for each target gene, a matched gene that was (1) within ±500 gene ranks for both metrics and (2) located >500 kb away from all target genes. LDSC heritability enrichment analyses were then conducted on these random matched genes, replicated 999 times.Abstract mediation modelWe applied the abstract mediation model (AMM)6 to estimate the heritability mediated by the same gene sets (cs-eGenes, sh-eGenes, GCTA eGenes and DEGs) and complex traits as in LDSC analyses. We estimated SNP × gene rank matrix for all genes passing quality control in our OneK1K analysis, excluding the MHC region. For each SNP, we considered the closest 50 genes and binned them as suggested in ref. 6 (closest, 2nd closest, 3rd–5th, 6th–10th, 11th–20th, 21st–30th, 31st–40th and 41st–50th). We then estimated the fraction of heritability mediated by each bin and used these fractions to test enrichment in gene sets. Apart from individual diseases or traits, we also meta-analysed blood-related diseases (ulcerative colitis, rheumatoid arthritis, primary biliary cirrhosis, multiple sclerosis, Crohn’s disease, celiac disease and SLE) and less-blood-related diseases (height, coronary artery disease and schizophrenia).Reporting summaryFurther information on research design is available in the Nature Portfolio Reporting Summary linked to this article.Data availabilityOneK1K single-cell gene expression and genotype data are publicly available on the Gene Expression Omnibus (GEO) under accession number GSE196830. For CLUES and ImmVar, single-cell gene expression data are available via GEO at GSE174188 and genotype data are available at dbGap under accession number phs002812.v1.p1. Summary statistics for the main CIGMA analyses in OneK1K and CLUES are provided in Supplementary Tables 3, 4 and 13. Additional publicly available data used in this study include LOEUF scores from the Genome Aggregation Database (gnomAD; https://storage.googleapis.com/gcp-public-data--gnomad/release/2.1.1/constraint/gnomad.v2.1.1.lof_metrics.by_gene.txt.bgz); enhancer counts (https://ars.els-cdn.com/content/image/1-s2.0-S0002929720300124-mmc2.xlsx); gene connectedness (https://zenodo.org/records/6618073)63; GWAS summary statistics (https://alkesgroup.broadinstitute.org/sumstats_formatted); LDSC baseline annotations and 1000 Genomes Phase 3 European genotype data (https://zenodo.org/records/10515792)66; and candidate cis-regulatory elements (cCREs; https://decoder-genetics.wustl.edu/catlasv1/humanenhancer/data/cCREs). Source data are provided with this paper.Code availabilityThe CIGMA Python package and all code used for analyses are available on GitHub https://github.com/Minhui-Chen/CIGMA and Zenodo (https://doi.org/10.5281/zenodo.19424343)67.References Maurano, M. T. et al. Systematic localization of common disease-associated variation in regulatory DNA. Science 337, 1190–1195 (2012).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar GTEx Consortium. The GTEx consortium atlas of genetic regulatory effects across human tissues. Science 369, 1318–1330 (2020).Article  Google Scholar Mostafavi, H., Spence, J. P., Naqvi, S. & Pritchard, J. K. Systematic differences in discovery of genetic effects on gene expression and complex traits. Nat. Genet. 55, 1866–1875 (2023).Article  CAS  PubMed  PubMed Central  Google Scholar Connally, N. J. et al. The missing link between genetic association and regulatory function. eLife 11, e74970 (2022).Article  CAS  PubMed  PubMed Central  Google Scholar Yao, D. W., O’Connor, L. J., Price, A. L. & Gusev, A. Quantifying genetic effects on disease mediated by assayed gene expression levels. Nat. Genet. 52, 626–633 (2020).Article  CAS  PubMed  PubMed Central  Google Scholar Weiner, D. J., Gazal, S., Robinson, E. B. & O’Connor, L. J. Partitioning gene-mediated disease heritability without eQTLs. Am. J. Hum. Genet. 109, 405–416 (2022).Article  CAS  PubMed  PubMed Central  Google Scholar Emani, P. S. et al. Single-cell genomics and regulatory networks for 388 human brains. Science 384, eadi5199 (2024).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Nathan, A. et al. Single-cell eQTL models reveal dynamic T cell state dependence of disease loci. Nature 606, 120–128 (2022).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Perez, R. K. et al. Single-cell RNA-seq reveals cell type–specific molecular and genetic associations to lupus. Science 376, eabf1970 (2022).Article  CAS  PubMed  PubMed Central  Google Scholar Yazar, S. et al. Single-cell eQTL mapping identifies cell type–specific genetic control of autoimmune disease. Science 376, eabf3041 (2022).Article  CAS  PubMed  Google Scholar Krockenberger, L. et al. FastGxC: fast and powerful context-specific eQTL mapping in bulk and single-cell data. Cell Genomics https://doi.org/10.1016/j.xgen.2026.101250 (2025).Yang, J. et al. Common SNPs explain a large proportion of heritability for human height. Nat. Genet. 42, 565–569 (2010).Article  CAS  PubMed  PubMed Central  Google Scholar Price, A. L. et al. Single-tissue and cross-tissue heritability of gene expression via identity-by-descent in related or unrelated individuals. PLoS Genet. 7, e1001317 (2011).Article  CAS  PubMed  PubMed Central  Google Scholar Wheeler, H. E. et al. Survey of the heritability and sparse architecture of gene expression traits across human tissues. PLoS Genet. 12, e1006423 (2016).Article  PubMed  PubMed Central  Google Scholar Loh, P.-R. et al. Contrasting genetic architectures of schizophrenia and other complex diseases using fast variance-components analysis. Nat. Genet. 47, 1385–1392 (2015).Article  CAS  PubMed  PubMed Central  Google Scholar Dahl, A. et al. A robust method uncovers significant context-specific heritability in diverse complex traits. Am. J. Hum. Genet. 106, 71–91 (2020).Article  CAS  PubMed  PubMed Central  Google Scholar Onengut-Gumuscu, S. et al. Fine mapping of type 1 diabetes susceptibility loci and evidence for colocalization of causal variants with lymphoid gene enhancers. Nat. Genet. 47, 381–386 (2015).Article  CAS  PubMed  PubMed Central  Google Scholar Oelen, R. et al. Single-cell RNA-sequencing of peripheral blood mononuclear cells reveals widespread, context-specific gene expression regulation upon pathogenic exposure. Nat. Commun. 13, 3267 (2022).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Strober, B. J. et al. Dynamic genetic regulation of gene expression during cellular differentiation. Science 364, 1287–1290 (2019).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Lin, W. et al. Disease-associated loci share properties with response eQTLs under common environmental exposures. Preprint at bioRxiv https://doi.org/10.1101/2025.04.30.651602 (2025).Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434–443 (2020).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar GTEx Consortium. Genetic effects on gene expression across human tissues. Nature 550, 204–213 (2017).Article  Google Scholar Grundberg, E. et al. Mapping cis- and trans-regulatory effects across multiple tissues in twins. Nat. Genet. 44, 1084–1089 (2012).Article  CAS  PubMed  PubMed Central  Google Scholar Spence, J. P. et al. Specificity, length and luck drive gene rankings in association studies. Nature 649, 918–925 (2026).Article  ADS  CAS  PubMed  Google Scholar Wang, X. & Goldstein, D. B. Enhancer domains predict gene pathogenicity and inform gene discovery in complex disease. Am. J. Hum. Genet. 106, 215–233 (2020).Article  CAS  PubMed  PubMed Central  Google Scholar Xu, S. et al. Using clusterProfiler to characterize multiomics data. Nat. Protoc. 19, 3292–3320 (2024).Article  CAS  PubMed  Google Scholar Ochoa, D. et al. The next-generation Open Targets Platform: reimagined, redesigned, rebuilt. Nucleic Acids Res. 51, D1353–D1359 (2023).Article  PubMed  PubMed Central  Google Scholar Karczewski, K. J. et al. Systematic single-variant and gene-based association testing of thousands of phenotypes in 394,841 UK Biobank exomes. Cell Genom. 2, 100168 (2022).Article  CAS  PubMed  PubMed Central  Google Scholar Zhang, K. et al. A single-cell atlas of chromatin accessibility in the human genome. Cell 184, 5985–6001 (2021).Article  CAS  PubMed  PubMed Central  Google Scholar Finucane, H. K. et al. Heritability enrichment of specifically expressed genes identifies disease-relevant tissues and cell types. Nat. Genet. 50, 621–629 (2018).Article  CAS  PubMed  PubMed Central  Google Scholar Yang, Y. et al. Investigating the shared genetic architecture between multiple sclerosis and inflammatory bowel diseases. Nat. Commun. 12, 5641 (2021).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Blanco, P. et al. Increase in activated CD8+ T lymphocytes expressing perforin and granzyme B correlates with disease activity in patients with systemic lupus erythematosus. Arthritis Rheum. 52, 201–211 (2005).Article  CAS  PubMed  Google Scholar Buang, N. et al. Type I interferons affect the metabolic fitness of CD8+ T cells from patients with systemic lupus erythematosus. Nat. Commun. 12, 1980 (2021).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar McKinney, E. F., Lee, J. C., Jayne, D. R. W., Lyons, P. A. & Smith, K. G. C. T-cell exhaustion, co-stimulation and clinical outcome in autoimmunity and infection. Nature 523, 612–616 (2015).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Khunsriraksakul, C. et al. Multi-ancestry and multi-trait genome-wide association meta-analyses inform clinical risk prediction for systemic lupus erythematosus. Nat. Commun. 14, 668 (2023).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Farh, K. K.-H. et al. Genetic and epigenetic fine mapping of causal autoimmune disease variants. Nature 518, 337–343 (2015).Article  ADS  CAS  PubMed  Google Scholar Manolio, T. A. et al. Finding the missing heritability of complex diseases. Nature 461, 747–753 (2009).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Yengo, L. et al. A saturated map of common genetic variants associated with human height. Nature 610, 704–712 (2022).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Giambartolomei, C. et al. Bayesian test for colocalisation between pairs of genetic association studies using summary statistics. PLoS Genet. 10, e1004383 (2014).Article  PubMed  PubMed Central  Google Scholar Hormozdiari, F. et al. Colocalization of GWAS and eQTL signals detects target genes. Am. J. Hum. Genet. 99, 1245–1260 (2016).Article  CAS  PubMed  PubMed Central  Google Scholar Mitchel, J. et al. A single-cell genetic colocalization test improves power and resolves disease-mediating cell types. Preprint at bioRxiv https://doi.org/10.1101/2025.10.10.681685 (2025).Zhang, Z. E., Kim, A., Suboc, N., Mancuso, N. & Gazal, S. Efficient count-based models improve power and robustness for large-scale single-cell eQTL mapping. Preprint at medRxiv https://doi.org/10.1101/2025.01.18.25320755 (2025).Zhou, W. et al. Efficient and accurate mixed model association tool for single-cell eQTL analysis. Preprint at medRxiv https://doi.org/10.1101/2024.05.15.24307317 (2024).Engelmann, J. P., Palma, A., Tomczak, J. M., Theis, F. J. & Casale, F. P. Mixed models with multiple instance learning. Preprint at arxiv.org/abs/2311.02455 (2024).Urbut, S. M., Wang, G., Carbonetto, P. & Stephens, M. Flexible statistical methods for estimating and testing effects in genomic studies with multiple conditions. Nat. Genet. 51, 187–195 (2019).Article  CAS  PubMed  Google Scholar Gamazon, E. R. et al. Using an atlas of gene regulation across 44 human tissues to inform complex disease- and trait-associated variation. Nat. Genet. 50, 956–967 (2018).Article  CAS  PubMed  PubMed Central  Google Scholar Hormozdiari, F. et al. Leveraging molecular quantitative trait loci to understand the genetic architecture of diseases and complex traits. Nat. Genet. 50, 1041–1047 (2018).Article  CAS  PubMed  PubMed Central  Google Scholar Dobbyn, A. et al. Landscape of conditional eQTL in dorsolateral prefrontal cortex and co-localization with schizophrenia GWAS. Am. J. Hum. Genet. 102, 1169–1184 (2018).Article  CAS  PubMed  PubMed Central  Google Scholar Zeng, B. et al. Multi-ancestry eQTL meta-analysis of human brain identifies candidate causal variants for brain-related traits. Nat. Genet. 54, 161–169 (2022).Article  CAS  PubMed  PubMed Central  Google Scholar Brotman, S. M. et al. Adipose tissue eQTL meta-analysis highlights the contribution of allelic heterogeneity to gene expression regulation and cardiometabolic traits. Nat. Genet. 57, 180–192 (2025).Article  CAS  PubMed  PubMed Central  Google Scholar Natri, H. M. et al. Cell-type-specific and disease-associated expression quantitative trait loci in the human lung. Nat. Genet. 56, 595–604 (2024).Article  CAS  PubMed  PubMed Central  Google Scholar Soskic, B. et al. Immune disease risk variants regulate gene expression dynamics during CD4+ T cell activation. Nat. Genet. 54, 817–826 (2022).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Popp, J. M. et al. Cell type and dynamic state govern genetic regulation of gene expression in heterogeneous differentiating cultures. Cell Genom. 4, 100701 (2024).Article  CAS  PubMed  PubMed Central  Google Scholar Qi, G. et al. Transcriptome-wide association studies at cell-state level using single-cell eQTL data. Cell Genom. 6, 101060 (2026).Article  CAS  PubMed  Google Scholar Ahlmann-Eltze, C. & Huber, W. Comparison of transformations for single-cell RNA-seq data. Nat. Methods 20, 665–672 (2023).Article  CAS  PubMed  PubMed Central  Google Scholar Liu, X., Li, Y. I. & Pritchard, J. K. Trans effects on gene expression can drive omnigenic inheritance. Cell 177, 1022–1034 (2019).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Chen, M. & Dahl, A. A robust model for cell type-specific interindividual variation in single-cell RNA sequencing data. Nat. Commun. 15, 5229 (2024).Article  ADS  CAS  PubMed  PubMed Central  Google Scholar Steinsaltz, D., Dahl, A. & Wachter, K. W. Statistical properties of simple random-effects models for genetic heritability. Electron. J. Stat. 12, 321–358 (2018).Article  MathSciNet  PubMed  PubMed Central  Google Scholar Chen, J. et al. A quantitative framework for characterizing the evolutionary history of mammalian gene expression. Genome Res. 29, 53–63 (2019).Article  CAS  PubMed  Google Scholar Chang, C. C. et al. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience 4, 7 (2015).Article  PubMed  PubMed Central  Google Scholar Steinsaltz, D., Dahl, A. & Wachter, K. W. On negative heritability and negative estimates of heritability. Genetics 215, 343–357 (2020).Article  PubMed  PubMed Central  Google Scholar Liu, Y., Sarkar, A., Kheradpour, P., Ernst, J. & Kellis, M. Evidence of reduced recombination rate in human regulatory domains. Genome Biol. 18, 193 (2017).Article  PubMed  PubMed Central  Google Scholar Mostafavi, H. Supplementary Data for ‘Systematic differences in discovery of genetic effects on gene expression and complex traits’. Zenodo. https://doi.org/10.5281/zenodo.6618073 (2023).Saha, A. et al. Co-expression networks reveal the tissue-specific regulation of transcription and splicing. Genome Res. 27, 1843–1858 (2017).Article  CAS  PubMed  PubMed Central  Google Scholar Finucane, H. K. et al. Partitioning heritability by functional annotation using genome-wide association summary statistics. Nat. Genet. 47, 1228–1235 (2015).Article  CAS  PubMed  PubMed Central  Google Scholar Gazal, S. S-LDSC reference files. Zenodo. https://doi.org/10.5281/zenodo.10515792 (2024).Chen, M. et al. Cell type-specific eQTLs underlie the genetic architecture of complex traits. Zenodo. https://doi.org/10.5281/zenodo.19424343 (2026).Wright, F. A. et al. Heritability and genomics of gene expression in peripheral blood. Nat. Genet. 46, 430–437 (2014).Article  CAS  PubMed  PubMed Central  Google Scholar Gusev, A. et al. Integrative approaches for large-scale transcriptome-wide association studies. Nat. Genet. 48, 245–252 (2016).Article  CAS  PubMed  PubMed Central  Google Scholar Lloyd-Jones, L. R. et al. The genetic architecture of gene expression in peripheral blood. Am. J. Hum. Genet. 100, 228–237 (2017).Article  CAS  PubMed  PubMed Central  Google Scholar Liu, X. et al. Functional architectures of local and distal regulation of gene expression in multiple human tissues. Am. J. Hum. Genet. 100, 605–616 (2017).Article  CAS  PubMed  PubMed Central  Google Scholar Kachuri, L. et al. Gene expression in African Americans, Puerto Ricans and Mexican Americans reveals ancestry-specific patterns of genetic architecture. Nat. Genet. 55, 952–963 (2023).Article  CAS  PubMed  PubMed Central  Google Scholar Saitou, M., Dahl, A., Wang, Q. & Liu, X. Allele frequency impacts the cross-ancestry portability of gene expression prediction in lymphoblastoid cell lines. Am. J. Hum. Genet. 111, 2814–2825 (2024).Article  CAS  PubMed  PubMed Central  Google Scholar Download referencesAcknowledgementsWe thank the Center for Research Informatics for providing the computing resources. The Center for Research Informatics is funded by the Biological Sciences Division at the University of Chicago with additional funding provided by the Institute for Translational Medicine, CTSA grant no. 2U54TR002389-06 from the National Institutes of Health. This work was funded by the National Institute of General Medical Sciences of the National Institutes of Health (R35GM150822 to A.D.).Author informationAuthors and AffiliationsSection of Genetic Medicine, University of Chicago, Chicago, IL, USAMinhui Chen  (陈敏惠), Xinpei Wang  (王鑫培), Sebastian Pott, Xuanyao Liu  (刘轩尧) & Andy DahlVirginia Institute for Psychiatric and Behavioral Genetics, Virginia Commonwealth University, Richmond, VA, USAMinhui Chen  (陈敏惠)Department of Psychiatry, Virginia Commonwealth University, Richmond, VA, USAMinhui Chen  (陈敏惠)Bioinformatics Interdepartmental Graduate Program, University of Californaia, Los Angeles, Los Angeles, CA, USALena KrockenbergerTranslational Genomics, Garvan Institute of Medical Research, Darlinghurst, New South Wales, AustraliaRika Tyebally & Joseph E. PowellUNSW Cellular Genomics Futures Institute, University of New South Wales, Sydney, New South Wales, AustraliaRika Tyebally & Joseph E. PowellDepartment of Human Genetics, University of Chicago, Chicago, IL, USAJeremy J. BergDepartment of Human Genetics, David Geffen School of Medicine, University of California, Los Angeles, Los Angeles, CA, USAJonathan FlintDepartment of Pathology and Laboratory Medicine, Department of Computational Medicine, Department of Biostatistics, University of California, Los Angeles, Los Angeles, CA, USABrunilda BalliuAuthorsMinhui Chen  (陈敏惠)View author publicationsSearch author on:PubMed Google ScholarXinpei Wang  (王鑫培)View author publicationsSearch author on:PubMed Google ScholarLena KrockenbergerView author publicationsSearch author on:PubMed Google ScholarRika TyeballyView author publicationsSearch author on:PubMed Google ScholarJeremy J. BergView author publicationsSearch author on:PubMed Google ScholarSebastian PottView author publicationsSearch author on:PubMed Google ScholarJonathan FlintView author publicationsSearch author on:PubMed Google ScholarJoseph E. PowellView author publicationsSearch author on:PubMed Google ScholarBrunilda BalliuView author publicationsSearch author on:PubMed Google ScholarXuanyao Liu  (刘轩尧)View author publicationsSearch author on:PubMed Google ScholarAndy DahlView author publicationsSearch author on:PubMed Google ScholarContributionsM.C. and A.D. developed statistical methodology, conducted analyses and wrote the paper. X.W. developed statistical methodology and conducted analyses. L.K. and R.T. assisted in data analyses. J.J.B., S.P., J.F., J.E.P., B.B. and X.L. provided suggestions and feedback to the manuscript.Corresponding authorCorrespondence to Andy Dahl.Ethics declarationsCompeting interestsThe authors declare no competing interests.Peer reviewPeer review informationNature thanks Shamil Sunyaev and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available.Additional informationPublisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.Extended data figures and tablesExtended Data Fig. 1 Permuting OneK1K data validates CIGMA robustness to expression levels, covarying cell types, and the count nature of scRNA-seq data.(a) The results in Fig. 2a are stratified across genes based on their expression level (left) or their variance in expression (right). (b) The cell permutations (as in Fig. 2a) are restricted to subsets of cell types; The x-axis indicates the number of permuted cell types: 0 represents the real data analysis in Fig. 3a (unpermuted), and ‘All’ represents permuting cells across all cell types as in Fig. 2a. Solid lines indicate permutation across cell types within the T-cell lineage (2 means CD4 ET and CD4 NC; 4 means CD4 ET, CD4 NC, CD8 ET, and CD8 NC). Dashed lines indicate permutation across T- and B-cell lineages (2 means CD4 NC and B IN; 4 means CD4 ET, CD4 NC, B Mem, and B IN). Error bars indicate 95% confidence intervals. In (a) and (b), points show transcriptome-wide medians for genes with nonnegative combined genetic and residual interindividual variances. (c) (Left) The number of cells per individual and cell type was limited to a specified maximum; otherwise, a random subset of cells was selected. ‘All’ indicates the real OneK1K analysis, in which all cells were retained, giving an average of 170 cells per individual and cell type. The plot shows the median heritability for the 3,994 genes that have positive combined genetic and residual interindividual variances in all scenarios. (Right) A specified proportion of reads was randomly sampled for each cell. The plot shows the median heritability for the 4,616 genes that have positive combined genetic and residual interindividual variances in all scenarios. Error bars indicate 95% confidence intervals for the medians.Source dataExtended Data Fig. 2 CIGMA estimates of cell-type-shared and cell-type-specific heritabilities in simulations.The simulations vary seven parameters: (a) the number of individuals, (b) the number of cell types, (c) cell type proportions, (d) cell numbers per individual (Supplementary Note 1), (e) the estimation error of cell-to-cell variation, simulated by drawn from Beta(noise level, 1) distribution (Supplementary Note 1), (f) the cell-type specificity, and (g) the ratio of specific genetic variance in the first cell type relative to other cell types, with all other cell types having equal specific genetic variance. (h) REML-based CIGMA estimates are deflated by estimation error in cell-to-cell variation (using the same simulations as in e). Gray regions represent the parameter values used in the baseline simulations. Red points are the true heritabilities. Box plots show the distribution of estimated heritabilities from 1,000 replicate simulations. Each box plot shows the median, first quartile, and third quartile, with whiskers extending up to 1.5 times the interquartile range. Data points beyond the whiskers are outliers.Source dataExtended Data Fig. 3 CIGMA is robust to varying genetic relationships across cell types in simulations.Simulations follow the Free model, where all cell types have equal eQTL covariance, except for a single pair of cell types whose eQTL correlation varies from −1 to 1 (x-axis). Panels show that CIGMA’s Free model nonetheless gives unbiased estimates (y-axis) for shared and specific genetic variances (a), heritabilities (b, c), and specificity (d). Red points are the true values. Each box plot shows the median, first quartile, and third quartile across 1,000 replicate simulations, with whiskers extending up to 1.5 times the interquartile range.Source dataExtended Data Fig. 4 Transcriptome-wide heritability estimates varying quality control parameters.Error bars represent 95% confidence intervals. Median, Mean, and Ratio refer to three approaches to aggregate estimates across genes, where Ratio refers to the ratio of transcriptome-wide means. Top panel: When filtering genes based on the standard errors of shared and specific heritability (h2) at 0.1, 0.5, 1, and 10, the number of remaining genes is 4,930, 8,427, 9,039, and 9,814, respectively. Bottom panel: When filtering genes based on the total genetic and residual interindividual variance at 0.01, 0,1, 0.2, and 0.3, the number of remaining genes is 8,992, 8,090, 6,296, and 4,298, respectively. Figure 3a in the Main text uses genes with positive total genetic and residual interindividual variances. Supplementary Table 5 provides published heritability estimates for comparison13,14,68,69,70,71,72,73.Source dataExtended Data Fig. 5 Robust enrichment of gene features in eQTL specificity after correcting for gene length, expression level, and mean expression differentiation across cell types in OneK1K.Expression level and mean expression differentiation across cell types were measured as described in Fig. S22. The same set of 7,042 genes was used as in Fig. 4a, except 106 genes with extreme specificity (>10 or