AbstractLarge-scale randomized experiments involving millions of observations have transformed social science across academic and applied domains. These methods are now widely used in fields ranging from economics and political science to public health, education and digital product development. Drawing on the experience of industry practitioners and academic collaborators, this Perspective outlines six areas where the expanding scale and scope of experimental studies have introduced new challenges. We describe open problems related to organizational incentives and experimental governance; privacy, fairness and ethics; long-term impact estimation; time-adaptive experimental studies; heterogeneous treatment effects; and generative artificial intelligence. For each area, we summarize recent methodological advances, highlight ongoing limitations and suggest directions for future work. Our goal is to surface practical challenges that merit greater attention from the research community and to encourage sustained collaboration between academics and practitioners in shaping the future of experimental studies.This is a preview of subscription content, access via your institutionAccess options Access through your institutionAccess Nature and 54 other Nature Portfolio journalsGet Nature+, our best-value online-access subscription27,99 € / 30 dayscancel any timeLearn moreSubscribe to this journalReceive 12 digital issues and online access to articles118,99 € per yearonly 9,92 € per issueLearn moreBuy this articlePurchase on SpringerLinkInstant access to the full article PDF.39,95 €Prices may be subject to local taxes which are calculated during checkoutReferencesDuflo, E. & Banerjee, A. Poor Economics Vol. 619 (PublicAffairs, 2011).Rajkumar, K., Saint-Jacques, G., Bojinov, I., Brynjolfsson, E. & Aral, S. A causal test of the strength of weak ties. Science 377, 1304–1310 (2022).Article PubMed CAS Google Scholar Horton, J. J. Price floors and employer preferences: evidence from a minimum wage experiment. Am. Econ. Rev. 115, 117–146 https://doi.org/10.1257/aer.20170637 (2025).Bojinov, I. & Gupta, S. Online experimentation: benefits, operational and methodological challenges, and scaling guide. Harv. Data Sci. Rev. https://doi.org/10.1162/99608f92.a579756e (2022).Fisher, R. A. The Design of Experiments (Oliver and Boyd, 1935).Neyman, J. On the application of probability theory to agricultural experiments: essay on principles. Stat. Sci. 5, 465–480 (1990).Google Scholar Rothamsted Experimental Station Harpenden. Rothamsted Experimental Station Report for 1925-1926, with the Supplement to the Guide to the Experimental Plots (Rothamsted Experimental Station, 1926).Kohavi, R., Longbotham, R., Sommerfield, D. & Henne, R. M. Controlled experiments on the web: survey and practical guide. Data Min. Knowl. Discov. 18, 140–181 (2009).Article Google Scholar Polonioli, A. et al. The ethics of online controlled experiments (A/B testing). Minds Mach. 33, 667–693 (2023).Article Google Scholar Larsen, N. et al. Statistical challenges in online controlled experiments: a review of A/B testing methodology. Am. Stat. 78, 135–149 https://doi.org/10.1080/00031305.2023.2257237 (2024).Gupta, S. et al. Top challenges from the first practical online controlled experiments summit. ACM SIGKDD Explor. Newsl. 21, 20–35 (2019).Article Google Scholar Ioannidis, J. P. A. Why most published research findings are false. PLoS Med. 2, e124 (2005).Article PubMed PubMed Central Google Scholar Head, M. L., Holman, L., Lanfear, R., Kahn, A. T. & Jennions, M. D. The extent and consequences of p-hacking in science. PLoS Biol. 13, e1002106 (2015).Article PubMed PubMed Central Google Scholar Lee, J. D., Sun, D. L., Sun, Y. & Taylor, J. E. Exact post-selection inference, with application to the lasso. Ann. Statist. 44, 907–927 https://doi.org/10.1214/15-AOS1371 (2016).Deng, A., Li, Y., Lu, J. & Ramamurthy, V. On post-selection inference in A/B testing. In Proc. 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining 2743–2752 (Association for Computing Machinery (ACM), 2021).Johari, R., Koomen, P., Pekelis, L. & Walsh, D. Peeking at A/B tests: why it matters, and what to do about it. In Proc. 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 1517–1525 (Association for Computing Machinery (ACM), 2017).Johari, R., Li, H., Liskovich, I. & Weintraub, G. Y. Experimental design in two-sided platforms: an analysis of bias. Manage. Sci. 68, 7069–7089 (2022).Article Google Scholar Ham, D. W., Lindon, M., Tingley, M. & Bojinov, I. Design-based confidence sequences: a general approach to risk mitigation in panel experiments. Manag. Sci. https://doi.org/10.1287/mnsc.2024.08235 (2026).Hayek, F. A. The use of knowledge in society. Am. Econ. Rev. 35, 519–530 (1945).Google Scholar Walsh, D. How to Design and Analyze Online A/B Tests within Decentralized Organizations (Stanford Univ., 2019).Hall, T. A. & Hasan, S. Organizational decision-making and the returns to experimentation. J. Organ. Des. 11, 129–144 (2022).Google Scholar Min, D. Screening for experiments. Games Econ. Behav. 142, 73–100 (2023).Article Google Scholar Bates, S., Jordan, M. I., Sklar, M. & Soloff, J. A. Principal-agent hypothesis testing. J. Am. Stat. Assoc. https://doi.org/10.1080/01621459.2026.2724031 (2026).Azevedo, E. M., Deng, A., Montiel Olea, J. L., Rao, J. & Weyl, E. G. A/B testing with fat tails. J. Polit. Econ. 128, 4614–000 (2020).Article Google Scholar Colson, E., Sibley, D. & Spiegel, D. The sobering truth about the impact of your business ideas. O’Reilly Radar/Business https://www.oreilly.com/radar/the-sobering-truth-about-the-impact-of-your-business-ideas/ (26 October 2021).Fernandes, M., Walls, L., Munson, S., Hullman, J. & Kay, M. Uncertainty displays using quantile dotplots or CDFs improve transit decision-making. In Proc. 2018 CHI Conference on Human Factors in Computing Systems, paper no 144, 1–12 (ACM, 2018).Yu, E. C., Sprenger, A. M., Thomas, R. P. & Dougherty, M. R. When decision heuristics and science collide. Psychon. Bull. Rev. 21, 268–282 (2014).Article PubMed Google Scholar Ruginski, I. T. et al. Non-expert interpretations of hurricane forecast uncertainty visualizations. Spat. Cogn. Comput. 16, 154–172 (2016).Article Google Scholar Azevedo, E. M., Deng, A., Montiel Olea, J. L. & Weyl, E.G. Empirical Bayes estimation of treatment effects with many A/B tests: an overview. In AEA Papers and Proceedings Vol. 109, 43–47 (American Economic Association, 2019).MD National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research, Bethesda. The Belmont Report: Ethical Principles and Guidelines for the Protection of Human Subjects of Research (Superintendent of Documents, 1978).Dwork, C. et al. The algorithmic foundations of differential privacy. Found. Trends Theor. Comput. Sci. 9, 211–407 (2014).Article Google Scholar Chen, W.-N., Cormode, G., Bharadwaj, A., Romov, P. & Ozgur, A. Federated experiment design under distributed differential privacy. In International Conference on Artificial Intelligence and Statistics 2458–2466 (PMLR, 2024).Movahedi, M. et al. Privacy-preserving randomized controlled trials: a protocol for industry scale deployment. In Proc. 2021 on Cloud Computing Security Workshop 59–69 (Association for Computing Machinery (ACM), 2021).Shi, J., Wang, D., Tesei, G. & Norgeot, B. Generating high-fidelity privacy-conscious synthetic patient data for causal effect estimation with multiple treatments. Front. Artif. Intell. 5, 918813 (2022).Article PubMed PubMed Central Google Scholar Farronato, C., MacCormack, A. & Mehta, S. Innovation at Uber: the launch of Express POOL. Harv. Bus. Sch. Case 620, 062 (2018).Google Scholar Liu, Q. et al. Can synthetic data be fair and private? A comparative study of synthetic data generation and fairness algorithms. In Proc. 15th International Learning Analytics and Knowledge Conference 591–600 (Association for Computing Machinery (ACM), 2025).Amad, H. et al. Improving the generation and evaluation of synthetic data for downstream medical causal inference. Adv. Neural Inform. Proc. Syst. 38, 30435–30478 (2026).Google Scholar Kim, D., Xu, Y. & Lin, T. A technical exploration of causal inference with hybrid LLM synthetic data. Preprint at https://doi.org/10.48550/arXiv.2511.00318 (2025).Liou, K., Zheng, W. & Anand, S. Privacy-preserving methods for repeated measures designs. In Companion Proc. Web Conference 2022 105–109 (Association for Computing Machinery (ACM), 2022).Cummings, R. & Desai, D. The role of differential privacy in GDPR compliance. In Proc. Workshop on Responsible Recommendation (FATREC’18) (Oct 2018) https://facctrec.github.io/fatrec2018/program/fatrec2018-cummings.pdf (Association for Computing Machinery (ACM), 2018).Malgieri, G. The concept of fairness in the GDPR: a linguistic and contextual interpretation. In Proc. 2020 Conference on Fairness, Accountability, and Transparency 154–166 (Association for Computing Machinery (ACM), 2020).Chen, R., Fang, F., Norton, T., McDonald, A. M. & Sadeh, N. Fighting the fog: evaluating the clarity of privacy disclosures in the age of CCPA. In Proc. 20th Workshop on Workshop on Privacy in the Electronic Society 73–102 (Association for Computing Machinery (ACM), 2021).Wachter, S., Mittelstadt, B. & Russell, C. Why fairness cannot be automated: bridging the gap between EU non-discrimination law and AI. Comput. Law Secur. Rev. 41, 105567 (2021).Article Google Scholar Chen, J., Kallus, N., Mao, X., Svacha, G. & Udell, M. Fairness under unawareness: assessing disparity when protected class is unobserved. In Proc. Conference on Fairness, Accountability, and Transparency 339–348 (Association for Computing Machinery (ACM), 2019).Kallus, N., Mao, X. & Zhou, A. Assessing algorithmic fairness with unobserved protected class using data combination. Manage. Sci. 68, 1959–1981 (2022).Article Google Scholar Taylor, L., Floridi, L. & van der Sloot, B. (eds) Group Privacy: New Challenges of Data Technologies (Springer, 2017).Zwitter, A. Big data ethics. Big Data Soc. 1, 2053951714559253 (2014).Article Google Scholar Sandvik, K. B. The humanitarian cyberspace: shrinking space or an expanding frontier? Third World Q. 37, 17–32 (2016).Article Google Scholar Saint-Jacques, G., Sepehri, A., Li, N. & Perisic, I. Fairness through experimentation: inequality in A/B testing as an approach to responsible design. Preprint at https://doi.org/10.48550/arXiv.2002.05819 (2020).Kallus, N. Treatment effect risk: bounds and inference. Manage. Sci. 69, 4579–4590 (2023).Article Google Scholar Kallus, N. What’s the harm? Sharp bounds on the fraction negatively affected by treatment. Adv. Neural Inf. Process. Syst. 35, 15996–16009 (2022).Article Google Scholar Jagielski, M. et al. Differentially private fair learning. In International Conference on Machine Learning 3000–3008 (PMLR, 2019).Calvi, A., Malgieri, G. & Kotzinos, D. The unfair side of privacy enhancing technologies: addressing the trade-offs between pets and fairness. In Proc. 2024 ACM Conference on Fairness, Accountability, and Transparency 2047–2059 (Association for Computing Machinery (ACM), 2024).Holtz, D., Lobel, F., Ruben Lobel, R., Inessa Liskovich, I. & Aral, S. Reducing interference bias in online marketplace experiments using cluster randomization: evidence from a pricing meta-experiment on Airbnb. Manag. Sci. 71, 390–406 https://doi.org/10.1287/mnsc.2020.01157 (2024).Masoero, L. et al. Multiple randomization designs: estimation and inference with interference. J. R. Stat. Soc. B Stat. Methodol. 88, 958–977 https://doi.org/10.1093/jrsssb/qkaf073 (2026).Burke, R., Sonboli, N. & Ordonez-Gauger, A. Balanced neighborhoods for multi-sided fairness in recommendation. In Conference on Fairness, Accountability and Transparency 202–214 (PMLR, 2018).Chaudhari, H. A., Lin, S. & Linda, O. A general framework for fairness in multistakeholder recommendations. Preprint at https://doi.org/10.48550/arXiv.2009.02423 (2020).Fleming, T. R. & DeMets, D. L. Surrogate end points in clinical trials: are we being misled? Ann. Intern. Med. 125, 605–613 (1996).Article PubMed CAS Google Scholar Echt, D. S. et al. Mortality and morbidity in patients receiving encainide, flecainide, or placebo: the cardiac arrhythmia suppression trial. N. Engl. J. Med. 324, 781–788 (1991).Article PubMed CAS Google Scholar Moore, T. J. Deadly medicine: why tens of thousands of patients died in America’s worst drug disaster. Stat. Med. 16, 2507–2510 (1997).Google Scholar Hotz, V. J., Imbens, G. W. & Klerman, J. A. Evaluating the differential effects of alternative welfare-to-work training components: a reanalysis of the California gain program. J. Labor Econ. 24, 521–566 (2006).Article Google Scholar Dmitriev, P., Frasca, B., Gupta, S., Kohavi, R. & Vaz, G. Pitfalls of long-term online controlled experiments. In 2016 IEEE International Conference on Big Data 1367–1376 https://doi.org/10.1109/BigData.2016.7840744 (IEEE, 2016).Zantedeschi, D., Feit, E. M. & Bradlow, E. T. Measuring multichannel advertising response. Manage. Sci. 63, 2706–2728 (2017).Article Google Scholar Lee, M. R. & Shen, M. Winner’s curse: bias estimation for total effects of features in online controlled experiments. In Proc. 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 491–499 (Association for Computing Machinery (ACM), 2018).Andrews, I., Kitagawa, T. & McCloskey, A. Inference on Winners (National Bureau of Economic Research, 2019).Masini, R. & Medeiros, M. C. Counterfactual analysis with artificial controls: inference, high dimensions, and nonstationarity. J. Am. Stat. Assoc. 116, 1773–1788 (2021).Article CAS Google Scholar Prentice, R. L. Surrogate endpoints in clinical trials: definition and operational criteria. Stat. Med. 8, 431–440 (1989).Article PubMed CAS Google Scholar Athey, S., Chetty, R., Imbens, G. W. & Kang, H. The Surrogate Index: Combining Short-Term Proxies to Estimate Long-Term Treatment Effects More Rapidly and Precisely (National Bureau of Economic Research, 2019).Zhang, V., Zhao, M., Le, A. & Kallus, N. Evaluating the surrogate index as a decision-making tool using 200 A/B tests at Netflix. Preprint at https://doi.org/10.48550/arXiv.2311.11922 (2023).Duan, W., Ba, S. & Zhang, C. Online experimentation with surrogate metrics: guidelines and a case study. In Proc. 14th ACM International Conference on Web Search and Data Mining 193–201 (Association for Computing Machinery (ACM), 2021).Zito, A., Greaves, D., Soriano, J. & Richardson, L. Pareto optimal proxy metrics. Appl. Stochastic Models Bus. Ind. 41, e70003 https://doi.org/10.1002/asmb.70003 (2025).Tripuraneni, N., Richardson, L., D’Amour, A., Soriano, J. & Yadlowsky, S. Choosing a proxy metric from past experiments. In Proc. 30th ACM SIGKDD Conf. on Knowledge Discovery and Data Mining, 5803–5812 (Association for Computing Machinery (ACM), 2024).Bibaut, A., Kallus, N., Ejdemyr, S. & Zhao, M. Long-term causal inference with imperfect surrogates using many weak experiments, proxies, and cross-fold moments. Preprint at https://doi.org/10.48550/arXiv.2311.04657 (2023).Bibaut, A., Kallus, N. & Lal, A. Nonparametric jackknife instrumental variable estimation and confounding robust surrogate indices. Preprint at https://doi.org/10.48550/arXiv.2406.14140 (2024).Bibaut, A., Chou, W., Ejdemyr, S. & Kallus, N. Learning the covariance of treatment effects across many weak experiments. In Proc. 30th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 153–162 (Association for Computing Machinery (ACM), 2024).Peysakhovich, A. & Eckles, D. Learning causal effects from many randomized experiments using regularized instrumental variables. In Proc. 2018 World Wide Web Conference 699–707 (Association for Computing Machinery (ACM), 2018).Tran, A., Bibaut, A. & Kallus, N. Inferring the long-term causal effects of long-term treatments from short-term experiments. In International Conference on Machine Learning, article no. 1984, 48565–48577 (Association for Computing Machinery (ACM), 2024).Kallus, N. & Uehara, M. Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning. Oper. Res. 70, 3282–3302 (2022).Article Google Scholar Kallus, N. & Uehara, M. Double reinforcement learning for efficient off-policy evaluation in Markov decision processes. J. Mach. Learn. Res. 21, 6742–6804 (2020).Google Scholar Lee, D. S. Training, wages, and sample selection: estimating sharp bounds on treatment effects. Rev. Econ. Stud. 76, 1071–1102 (2009).Article Google Scholar Semenova, V. Generalized Lee bounds. Preprint at https://doi.org/10.48550/arXiv.2008.12720 (2020).Dorn, J., Guo, K. & Kallus, N. Doubly-valid/doubly-sharp sensitivity analysis for causal inference with unmeasured confounding. J. Am. Stat. Assoc. 120, 331–342 https://doi.org/10.1080/01621459.2024.2335588 (2025).Wald, A. Sequential tests of statistical hypotheses. Ann. Math. Stat. 16, 117–186 (1945).Article Google Scholar Jennison, C. & Turnbull, B. W. Group Sequential Methods with Applications to Clinical Trials (CRC, 1999).Johari, R., Koomen, P., Pekelis, L. & Walsh, D. Always valid inference: continuous monitoring of A/B tests. Oper. Res. 70, 1806–1821 (2022).Article Google Scholar Bibaut, A., Kallus, N. & Lindon, M. Near-optimal non-parametric sequential tests and confidence sequences with possibly dependent observations. Preprint at https://doi.org/10.48550/arXiv.2212.14411 (2022).Lindon, M. & Malek, A. Anytime-valid inference for multinomial count data. Adv. Neural Inform. Process. Syst. 35, 2817–2831 (2022).Article Google Scholar Howard, S. R. & Ramdas, A. Sequential estimation of quantiles with applications to A/B testing and best-arm identification. Bernoulli 28, 1704–1728 (2022).Article Google Scholar Lindon, M. et al. Anytime-valid linear models and regression adjusted causal inference in randomized experiments. Preprint at https://doi.org/10.48550/arXiv.2210.08589 (2022).Cho, B., Gan, K. & Kallus, N. Peeking with PEAK: sequential, nonparametric composite hypothesis tests for means of multiple data streams. In International Conference on Machine Learning, article no. 337, 8487–8509 (Association for Computing Machinery (ACM), 2024).Lindon, M., Sanden, C. & Shirikian, V. Rapid regression detection in software deployments through sequential testing. In Proc. 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining 3336–3346 https://doi.org/10.1145/3534678.3539099 (Association for Computing Machinery (ACM), 2022).Maharaj, A. et al. Anytime-valid confidence sequences in an enterprise A/B testing platform. In Companion Proc. ACM Web Conference 2023 396–400 (Association for Computing Machinery (ACM), 2023).Schmit, S. & Miller, E. Sequential confidence intervals for relative lift with regression adjustments. svenschmit.com https://svenschmit.com/assets/pdf/code_2022_ci.pdf (2022).ter Schure, J. A. et al. Bacillus Calmette–Guérin vaccine to reduce COVID-19 infections and hospitalisations in healthcare workers—a living systematic review and prospective all-in meta-analysis of individual participant data from randomised controlled trials. Preprint at medRxiv https://doi.org/10.1101/2022.12.15.22283474 (2022).Silva, I. R., Kulldorff, M. & Katherine Yih, W. Optimal alpha spending for sequential analysis with binomial data. J. R. Stat. Soc. B 82, 1141–1164 (2020).Article Google Scholar Bojinov, I. & Shephard, N. Time series experiments and causal estimands: exact randomization tests and trading. J. Am. Stat. Assoc. 114, 1665–1682 https://doi.org/10.1080/01621459.2018.1527225 (2019).Article CAS Google Scholar Bojinov, I., Simchi-Levi, D. & Zhao, J. Design and analysis of switchback experiments. Manage. Sci. 69, 3759–3777 (2023).Article Google Scholar Hu, Y. & Wager, S. Switchback experiments under geometric mixing. Preprint at https://doi.org/10.48550/arXiv.2209.00197 (2022).Jamieson, K. & Nowak, R. Best-arm identification algorithms for multi-armed bandits in the fixed confidence setting. In 2014 48th Annual Conf. on Information Sciences and Systems (CISS) https://doi.org/10.1109/CISS.2014.6814096 (IEEE, 2014).Russo, D. Simple Bayesian algorithms for best-arm identification. Oper. Res. 68, 1625–1647 https://doi.org/10.1287/opre.2019.1911 (2020).Article Google Scholar Kasy, M. & Sautmann, A. Adaptive treatment assignment in experiments for policy choice. Econometrica 89, 113–132 https://doi.org/10.3982/ECTA17527 (2021).Article Google Scholar Li, Y., Mao, J. & Bojinov, I. Balancing risk and reward: a batched-bandit strategy for automated phased release. Adv. Neural Inf. Process. Syst. 36, 76181–76201 (2024).Google Scholar Villar, S. S., Bowden, J. & Wason, J. Multi-armed bandit models for the optimal design of clinical trials: benefits and challenges. Stat. Sci. 30, 199–215 https://doi.org/10.1214/14-STS504 (2015).Article PubMed PubMed Central Google Scholar Yao, J., Brunskill, E., Pan, W., Murphy, S. & Doshi-Velez, F. Power constrained bandits. In Machine Learning for Healthcare Conference 209–259 (PMLR, 2021).Khachatryan, G., Kostyuk, V. & Rounds, N. A Community of Bandits: How OfferFit Automates Experimentation for Lifecycle Marketers https://offerfit.ai/content/white-paper-pdf/a-community-of-bandits-download (Braze, 2023).Kallus, N. & Udell, M. Dynamic assortment personalization in high dimensions. Oper. Res. 68, 1020–1037 (2020).Article Google Scholar Bibaut, A. & Kallus, N. 2025. Demystifying inference after adaptive experiments. Ann. Rev. Stat. Appl. 12, 407–423 https://doi.org/10.1146/annurev-statistics-040522-015431 (2025).Gibbs, I. & EmCandès, E. J. Conformal inference for online prediction with arbitrary distribution shifts. J. Mach. Learn. Res. 25, 1–36 (2024).Google Scholar Liu, Y., Van Roy, B. & Xu, K. Nonstationary bandit learning via predictive sampling. In Proc. 26th International Conference on Artificial Intelligence and Statistics, Vol. 206 (eds Ruiz, F. et al.) 6215–6244 https://proceedings.mlr.press/v206/liu23e.html (PMLR, 2023).Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S. & Athey, S. Confidence intervals for policy evaluation in adaptive experiments. Proc. Natl Acad. Sci. USA 118, e2014602118 (2021).Article PubMed PubMed Central CAS Google Scholar Bibaut, A., Dimakopoulou, M., Kallus, N., Chambaz, A. & van Der Laan, M. Post-contextual-bandit inference. Adv. Neural Inf. Process. Syst. 34, 28548–28559 (2021).PubMed PubMed Central Google Scholar Liang, B. & Bojinov, I. An experimental design for anytime-valid causal inference on multi-armed bandits. Preprint at https://doi.org/10.48550/arXiv.2311.05794 (2023).Grbovic, M. & Cheng, H. Real-time personalization using embeddings for search ranking at Airbnb. In Proc. 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 311–320 (Association for Computing Machinery (ACM), 2018).Wu, H. et al. Interpretable personalized experimentation. In Proc. 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 4173–4183 (Association for Computing Machinery (ACM), 2022).Athey, S., Keleher, N. & Spiess, J. Machine learning who to nudge: causal vs predictive targeting in a field experiment on student financial aid renewal. J. Econometrics 249, 105945 https://doi.org/10.1016/j.jeconom.2024.105945 (2025).Gelman, A. & Hill, J. Data Analysis Using Regression and Multilevel/Hierarchical Models (Cambridge Univ. Press, 2006).Künzel, S. R., Sekhon, J. S., Bickel, P. J. & Yu, B. Metalearners for estimating heterogeneous treatment effects using machine learning. Proc. Natl Acad. Sci. USA 116, 4156–4165 (2019).Article PubMed PubMed Central Google Scholar Wager, S. & Athey, S. Estimation and inference of heterogeneous treatment effects using random forests. J. Am. Stat. Assoc. 113, 1228–1242 (2018).Article CAS Google Scholar Chernozhukov, V., Demirer, M., Duflo, E. & Fernandez-Val, I. Generic Machine Learning Inference on Heterogeneous Treatment Effects in Randomized Experiments, with an Application to Immunization in India (no. w24678) (National Bureau of Economic Research, 2018).Athey, S. & Imbens, G. Recursive partitioning for heterogeneous causal effects. Proc. Natl Acad. Sci. USA 113, 7353–7360 (2016).Article PubMed PubMed Central CAS Google Scholar Dwivedi, R. et al. Stable discovery of interpretable subgroups via calibration in causal studies. Int. Stat. Rev. 88, S135–S178 (2020).Article Google Scholar Alaa, A. & Van Der Schaar, M. Validating causal inference models via influence functions. In International Conference on Machine Learning 191–201 (PMLR, 2019).Kallus, N. & McInerney, J. The implicit delta method. Adv. Neural Inf. Process. Syst. 35, 37471–37483 (2022).Article Google Scholar Romano, J. P. & Wolf, M. Exact and approximate stepdown methods for multiple hypothesis testing. J. Am. Stat. Assoc. 100, 94–108 (2005).Article CAS Google Scholar Benjamini, Y. & Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J. R. Stat. Soc. B 57, 289–300 (1995).Article Google Scholar Sculley, D. et al. Hidden technical debt in machine learning systems. In Adv. Neural Inf. Process. Syst. https://proceedings.neurips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html (2015).Bojinov, I., Holtz, D., Johari, R., Schmit, S. & Tingley, M. Want your company to get better at experimentation? Learn fast by democratizing testing. Harv. Bus. Rev. 104, 96–103 (2025).Google Scholar Vaccaro, M. et al. Advancing AI negotiations: a large-scale autonomous negotiation competition. Proc. Natl Acad. Sci. USA 123, e2521774123 (2026).Article PubMed PubMed Central CAS Google Scholar Korinek, A. Generative AI for economic research: use cases and implications for economists. J. Econ. Lit. 61, 1281–1317 (2023).Article Google Scholar Bojinov, I. Keep your AI projects on track. Harv. Bus. Rev. 101, 53–59 (2023).Google Scholar Horton, J. J., Filippas, A. & Manning, B. S. Large Language Models as Simulated Economic Agents: What Can We Learn From Homo Silicus? (no. w31122) (National Bureau of Economic Research, 2023).Argyle, L. P. et al. Out of one, many: using language models to simulate human samples. Polit. Anal. 31, 337–351 (2023).Article Google Scholar Manning, B. S., Zhu, K. & Horton, J. J. Automated Social Science: Language Models as Scientist and Subjects (National Bureau of Economic Research, 2024).Manning, B. S. & Horton, J. J. General Social Agents (National Bureau of Economic Research, 2026).Aher, G. V., Arriaga, R. I. & Kalai, A. T. Using large language models to simulate multiple humans and replicate human subject studies. In International Conference on Machine Learning 337–371 (PMLR, 2023).Park, J. S. et al. Generative agents: interactive simulacra of human behavior. In Proc. 36th Annual ACM Symposium on User Interface Software and Technology, article no. 2, 1–22 (Association for Computing Machinery (ACM), 2023).Messeri, L. & Crockett, M. J. Artificial intelligence and illusions of understanding in scientific research. Nature 627, 49–58 (2024).Article PubMed CAS Google Scholar Santurkar, S. et al. Whose opinions do language models reflect? In International Conference on Machine Learning 29971–30004 (PMLR, 2023).Navigli, R., Conia, S. & Ross, B. Biases in large language models: origins, inventory, and discussion. ACM J. Data Inf. Qual. 15, 1–21 (2023).Article Google Scholar Grossmann, I. et al. AI and the transformation of social science research. Science 380, 1108–1109 (2023).Article PubMed CAS Google Scholar Binz, M. & Schulz, E. Using cognitive psychology to understand GPT-3. Proc. Natl Acad. Sci. USA 120, e2218523120 (2023).Article PubMed PubMed Central CAS Google Scholar Download referencesAcknowledgementsWe thank M. Velarde, V. Hannan, S. Liu, W. Chou and T.-P. Lee for their contributions to the Causal Inference & Digital Experimentation Roundtable at Netflix in April 2023. This work arose out of conversations that occurred at this event. The views expressed in this paper are solely those of the authors and do not necessarily represent the views of their employers.Author informationAuthors and AffiliationsDecision, Risk and Operations, Columbia Business School, New York, NY, USADavid HoltzMIT Initiative on the Digital Economy, Cambridge, MA, USADavid HoltzTechnology & Operations Management, Harvard Business School, Boston, MA, USAIavor BojinovManagement Science & Engineering, Stanford University, Stanford, CA, USARamesh JohariNetflix, Los Gatos, CA, USANathan Kallus, Sathya Anand & Matthew WardropOperations Research and Information Engineering & Cornell Tech, Cornell University, Ithaca, NY, USANathan KallusOpenAI, San Francisco, CA, USAApoorva Lal & Sven SchmitIndependent researcher, San Francisco, CA, USAKyle Carlson & Konrad MiziolekStripe, San Francisco, CA, USABrent CohnMETR, Berkeley, CA, USATom CunninghamIndependent researcher, Bellevue, WA, USAAlex DengUber, San Francisco, CA, USAMaria Dimakopoulou & Jialiang MaoEconomics, Wharton School of the University of Pennsylvania, Philadelphia, PA, USAAmit GandhiBraze, Cambridge, MA, USAVictor KostyukMarketing, Harvard Business School, Boston, MA, USAMadhav KumarIndependent researcher, Daly City, CA, USAShyue-Ming LohMicrosoft, Redmond, WA, USAWidad MachmouchiAmazon, Seattle, WA, USAJames McQueen & Dominique Perrault-JoncasMeta, Menlo Park, CA, USAJohn Meakin & Vladimir PetrovicIndependent researcher, Stonington, CT, USAJames SorensonDoorDash, San Francisco, CA, USAMichael ZhaoRoblox, San Mateo, CA, USAWenjing ZhengIndependent researcher, Boulder, CO, USAMartin TingleyAuthorsDavid HoltzView author publicationsSearch author on:PubMed Google ScholarIavor BojinovView author publicationsSearch author on:PubMed Google ScholarRamesh JohariView author publicationsSearch author on:PubMed Google ScholarNathan KallusView author publicationsSearch author on:PubMed Google ScholarApoorva LalView author publicationsSearch author on:PubMed Google ScholarSathya AnandView author publicationsSearch author on:PubMed Google ScholarKyle CarlsonView author publicationsSearch author on:PubMed Google ScholarBrent CohnView author publicationsSearch author on:PubMed Google ScholarTom CunninghamView author publicationsSearch author on:PubMed Google ScholarAlex DengView author publicationsSearch author on:PubMed Google ScholarMaria DimakopoulouView author publicationsSearch author on:PubMed Google ScholarAmit GandhiView author publicationsSearch author on:PubMed Google ScholarVictor KostyukView author publicationsSearch author on:PubMed Google ScholarMadhav KumarView author publicationsSearch author on:PubMed Google ScholarShyue-Ming LohView author publicationsSearch author on:PubMed Google ScholarWidad MachmouchiView author publicationsSearch author on:PubMed Google ScholarJialiang MaoView author publicationsSearch author on:PubMed Google ScholarJames McQueenView author publicationsSearch author on:PubMed Google ScholarJohn MeakinView author publicationsSearch author on:PubMed Google ScholarKonrad MiziolekView author publicationsSearch author on:PubMed Google ScholarDominique Perrault-JoncasView author publicationsSearch author on:PubMed Google ScholarVladimir PetrovicView author publicationsSearch author on:PubMed Google ScholarSven SchmitView author publicationsSearch author on:PubMed Google ScholarJames SorensonView author publicationsSearch author on:PubMed Google ScholarMatthew WardropView author publicationsSearch author on:PubMed Google ScholarMichael ZhaoView author publicationsSearch author on:PubMed Google ScholarWenjing ZhengView author publicationsSearch author on:PubMed Google ScholarMartin TingleyView author publicationsSearch author on:PubMed Google ScholarContributionsAll authors contributed equally to this work.Corresponding authorsCorrespondence to David Holtz or Martin Tingley.Ethics declarationsCompeting interestsD.H. is currently a paid part-time visiting researcher at OpenAI. D.H. has a significant financial interest and/or is an angel investor in Interviewing.io, Specific Labs, Datoric, Trapeze, Pelica, Ralo, Tandem, Crowdvolt, Hyphen, Alan, Workstream, HomeRoom, Coinshift and SecurityPal. D.H. has been a co-organizer of the Conference on Digital Experimentation at MIT since 2022; during this period, the conference has received sponsorship from Netflix, Booking.com, Amazon, Itaù, Eppo and Statsig. D.H.’s research is or has been funded by Microsoft, the Boston University Center for Antiracist Research, the Russell Sage Foundation, the Alfred P. Sloan Foundation and the Summer Institute in Computational Social Science. D.H. is an inventor of patents and patent applications assigned to Airbnb and Microsoft (Patents No. 10,614,385 B2 and No. US 10,644, 855 B2; Patent Application No. US 2022/0036265 A1). I.B. declares no competing interests. R.J. is co-director of the Stanford Causal Science Center (SC2). SC2 receives event sponsorship from Amazon and maintains a collaborative postdoctoral research programme with Airbnb. R.J. is also supported by the National Science Foundation, the UK AI Security Institute, the Helmsley Charitable Trust and the Stanford Human-Centered AI Institute. R.J. also holds equity in Apple and Amazon. N.K. is currently employed by Netflix. Additionally, N.K. holds equity in Netflix, Google and Amazon. A.L. is currently employed by OpenAI and was employed by Netflix during the period in which this research was conducted. A.L. holds equity in OpenAI. S.A. is currently employed by Netflix and was employed by Netflix during the period in which this research was conducted. Additionally, S.A. holds equity in Netflix and Amazon. K.C. was employed by Stripe during the period in which this research was conducted and holds equity in Stripe. B.C. holds equity in Stripe and is employed by Stripe. T.C. holds equity in OpenAI. A.D. is currently employed by Meta and was employed by Airbnb during the period in which this research was conducted. Additionally, A.D. holds equity in Meta, Microsoft, Netflix and Amazon. M.D. is or was affiliated with the following institutions: she is currently director of engineering at Uber. She was previously director of ML engineering at Spotify. Earlier, she held positions at Netflix and Google. A.G. declares no competing interests. V.K. is currently employed by Braze and was employed by OfferFit at the time this research was conducted, which was subsequently acquired by Braze. V.K. holds equity in Braze. M.K. declares no competing. S.-M.L. is currently employed by Meta and was previously employed by Instacart during the period in which this research was conducted. Additionally, S.-M.L. holds equity in Meta and Atlassian. W.M. is currently employed by Microsoft and was employed by Microsoft during the period in which this research was conducted. Additionally, W.M. holds equity in Microsoft, Netflix and Meta. J. Mao is currently employed by and holds equity in Uber. J. McQueen is currently employed by Amazon and was employed by Amazon during the period in which this research was conducted. Additionally, J. McQueen holds equity in Microsoft, Netflix, Alphabet Inc. (Google), Amazon and Meta. J. Meakin is currently employed by Meta and was employed by Meta during the period in which this research was conducted. Additionally, J. Meakin holds equity in Meta. K.M. is currently employed by and holds equity in Instacart. D.P.-J. is currently employed by Amazon and was employed by Amazon during the period in which this research was conducted. Additionally, D.P.-J. holds equity in Microsoft, Netflix, Alphabet Inc. (Google), Amazon and Meta. V.P. is currently employed by Meta and was employed by Meta during the period in which this research was conducted. Additionally, V.P. holds equity in Meta. S.S. is currently employed at OpenAI and was employed at Eppo during the period in which this research was conducted. S.S. holds equity in OpenAI, Netflix and Microsoft. J.S. is currently employed by LinkedIn and was employed by LinkedIn during the period in which this research was conducted. Additionally, J.S. holds equity in Microsoft and Alphabet Inc. (Google). M.W. is currently employed by and holds equity in Netflix. M.Z. is currently employed by DoorDash and was previously employed by Netflix during the period in which this research was conducted. Additionally, M.Z. holds equity in DoorDash, Roblox, Netflix, Alphabet Inc. (Google), Amazon, Microsoft, Uber and Meta. W.Z. is currently employed by Roblox and was previously employed by Netflix during the period in which this research was conducted. Additionally, W.Z. holds equity in Roblox, Netflix, Alphabet Inc. (Google), Amazon and Meta. M.T. is currently employed by Microsoft and was previously employed by Netflix during the period in which this research was conducted. Additionally, M.T. holds equity in Microsoft, Netflix, Alphabet Inc. (Google), Amazon and Meta.Peer reviewPeer review informationNature Human Behaviour thanks Stefano Balietti and the other, anonymous, reviewer(s) for their contribution to the peer review of this work.Additional informationPublisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.Rights and permissionsSpringer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.Reprints and permissionsAbout this article