Anderson, J. R. & Lebiere, C. The Atomic Components of Thought (Lawrence Erlbaum Associates, 1998).Swartout, W. R. et al. Toward virtual humans. AI Mag. 27, 96–108 (2006).Google Scholar Simon, H. A. The Sciences of the Artificial 3rd edn (MIT Press, 1996).Schelling, T. C. Models of segregation. Am. Econ. Rev. 59, 488–493 (1969).Google Scholar Epstein, J. M. & Axtell, R. L. Growing Artificial Societies: Social Science from the Bottom Up https://doi.org/10.7551/mitpress/3374.001.0001 (MIT Press, 1996).Brown, T. et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 33, 1877–1901 (2020).Bommasani, R. et al. On the opportunities and risks of foundation models. Preprint at https://doi.org/10.48550/arXiv.2108.07258 (2022).Gao, Y., Lee, D., Burtch, G. & Fazelpour, S. Take caution in using LLMs as human surrogates. Proc. Natl Acad. Sci. USA 122, e2501660122 (2025).Article Google Scholar Horton, J. J., Filippas, A. & Manning, B. S. Large language models as simulated economic agents: what can we learn from homo silicus? Preprint at https://doi.org/10.48550/arXiv.2301.07543 (2026).Argyle, L. P. et al. Out of one, many: using language models to simulate human samples. Polit. Anal. 31, 337–351 (2023).Article Google Scholar Park, J. S. et al. Generative agents: interactive simulacra of human behavior. In Proc. 36th Annual ACM Symposium on User Interface Software and Technology 1–22 (Association for Computing Machinery, 2023).Gao, C. et al. Large language models empowered agent-based modeling and simulation: a survey and perspectives. Humanit. Soc. Sci. Commun. 11, 1259, https://doi.org/10.1057/s41599-024-03611-3 (2024).Article Google Scholar Dillion, D., Tandon, N., Gu, Y. & Gray, K. Can AI language models replace human participants? Trends Cogn. Sci. 27, 597–600 (2023).Article Google Scholar Wang, L. et al. A survey on large language model based autonomous agents. Front. Comput. Sci. 18, 186345 (2024).Article Google Scholar Tversky, A. Features of similarity. Psychol. Rev. 84, 327–352 (1977).Article Google Scholar Bates, J. The role of emotion in believable agents. Commun. ACM 37, 122–125 (1994).Article Google Scholar Loyall, A. B. Believable Agents: Building Interactive Personalities. PhD thesis, Carnegie Mellon Univ., Pittsburgh, PA (1997).Russell, S. & Norvig, P. Artificial Intelligence: A Modern Approach 3rd edn (Prentice Hall Press, 2009).Wooldridge, M. An Introduction to MultiAgent Systems 2nd edn (John Wiley & Sons, 2009).Rahwan, I. et al. Machine behaviour. Nature 568, 477–486 (2019).Article Google Scholar Argyle, L. P. et al. Arti-‘fickle’ intelligence: using LLMs as a tool for inference in the political and social sciences. Nat. Comput. Sci. 5, 737–744 https://doi.org/10.1038/s43588-025-00843-4 (2025).Article Google Scholar Laird, J. E. The Soar Cognitive Architecture (MIT Press, 2012).Shanahan, M., McDonell, K. & Reynolds, L. Role play with large language models. Nature 623, 493–498 https://doi.org/10.1038/s41586-023-06647-8 (2023).Article Google Scholar Vaswani, A. et al. Attention is all you need. In Proc. 31st International Conference on Neural Information Processing Systems 6000–6010 (Curran Associates, 2017).Rae, J. W. et al. Scaling language models: methods, analysis and insights from training Gopher. Preprint at https://doi.org/10.48550/arXiv.2112.11446 (2022).Polino, A., Pascanu, R. & Alistarh, D. Model compression via distillation and quantization. In Proc. 6th International Conference on Learning Representations (ICLR) (2018).Frantar, E. & Alistarh, D. SparseGPT: massive language models can be accurately pruned in one-shot. In International Conference on Machine Learning 10323–10337 (PMLR, 2023).Frantar, E., Ashkboos, S., Hoefler, T. & Alistarh, D. OPTQ: accurate post-training quantization for generative pre-trained transformers. In The Eleventh International Conference on Learning Representations https://openreview.net/forum?id=tcbBPnfwxS (2023).Ouyang, L. et al. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems Vol. 35 (eds Koyejo, S. et al.) 27730–27744 https://proceedings.neurips.cc/paper_files/paper/2022/file/b1efde53be364a73914f58805a001731-Paper-Conference.pdf (Curran Associates, 2022).Binz, M. et al. A foundation model to predict and capture human cognition. Nature 644, 1002–1009 https://doi.org/10.1038/s41586-025-09215-4 (2025).Article Google Scholar Binz, M. et al. Post-training makes large language models less human-like. Preprint at https://doi.org/10.48550/arXiv.2605.07632 (2026).Park, J. S. et al. Social simulacra: creating populated prototypes for social computing systems. In Proc. 35th Annual ACM Symposium on User Interface Software and Technology 1–18 (ACM, 2022).Zhang, Z. et al. A survey on the memory mechanism of large language model-based agents. ACM Trans. Inf. Syst. 43, 1–47 (2025).Google Scholar Schick, T. et al. Toolformer: language models can teach themselves to use tools. Adv. Neural Inf. Process. Syst. 36, 68539–68551 (2023).Article Google Scholar Yao, S. et al. Tree of thoughts: deliberate problem solving with large language models. Adv. Neural Inf. Process. Syst. 36, 11809–11822 (2023).Article Google Scholar Huang, X. et al. Understanding the planning of LLM agents: a survey. Preprint at https://doi.org/10.48550/arXiv.2402.02716 (2024).Gibson, J. J. The Ecological Approach to Visual Perception: Classic Edition 1st edn https://doi.org/10.4324/9781315740218 (Psychology Press, 2014).Franklin, S. & Graesser, A. Is it an agent, or just a program?: A taxonomy for autonomous agents. In Intelligent Agents III Agent Theories, Architectures, and Languages (eds Müller, J. P. et al.) 21–35 (Springer, 1997).Shanahan, M. Talking about large language models. Commun. ACM 67, 68–79 https://doi.org/10.1145/3624724 (2024).Article Google Scholar Cronbach, L. J. & Meehl, P. E. Construct validity in psychological tests. Psychol. Bull. 52, 281–302 (1955).Article Google Scholar Jeon, M. S. et al. Simulating conversations on social media with generative agent-based models. EPJ Data Sci. 14, 79 https://doi.org/10.1140/epjds/s13688-025-00593-3 (2025).Article Google Scholar Cirulli, D., Cimini, G. & Palermo, G. How large language models play humans in online conversations: a simulated study of the 2016 US politics on Reddit. Preprint at https://doi.org/10.48550/arXiv.2506.21620 (2025).Shao, Y., Li, L., Dai, J. & Qiu, X. Character-LLM: a trainable agent for role-playing. In Proc. 2023 Conference on Empirical Methods in Natural Language Processing (eds Bouamor, H. et al.) 13153–13187 https://aclanthology.org/2023.emnlp-main.814 (Association for Computational Linguistics, 2023).Nowak, K. & Fox, J. Avatars and computer-mediated communication: a review of the definitions, uses, and effects of digital representations. Rev. Commun. Res. 6, 30–53 (2018).Article Google Scholar Cassell, J. Embodied conversational interface agents. Commun. ACM 43, 70–78 (2000).Article Google Scholar Mateas, M. An Oz-centric review of interactive drama and believable agents. In Artificial Intelligence Today: Recent Trends and Developments (eds Wooldridge, M. J. & Veloso, M.) Lecture Notes in Artificial Intelligence Vol. 1600, 297–328 (Springer, 1999).Turing, A. M. Computing machinery and intelligence. Mind 59, 433–460 (1950).Article MathSciNet Google Scholar Tversky, A. & Kahneman, D. Judgment under uncertainty: heuristics and biases. Science 185, 1124–1131 https://doi.org/10.1126/science.185.4157.1124 (1974).Article Google Scholar Jones, C. R. & Bergen, B. K. Large language models pass a standard three-party Turing test. Proc. Natl Acad. Sci. USA 123, e2524472123 (2026).Article MathSciNet Google Scholar Kotek, H., Dockum, R. & Sun, D. Gender bias and stereotypes in large language models. In Proc. ACM Collective Intelligence Conference 12–24 (ACM, 2023).Tao, Y., Viberg, O., Baker, R. S. & Kizilcec, R. F. Cultural bias and cultural alignment of large language models. PNAS Nexus 3, pgae346 https://doi.org/10.1093/pnasnexus/pgae346 (2024).Article Google Scholar Narayanan Venkit, P., Gautam, S., Panchanadikar, R., Huang, T.-H. & Wilson, S. Nationality bias in text generation. In Proc. 17th Conference of the European Chapter of the Association for Computational Linguistics (eds Vlachos, A. & Augenstein, I.) 116–122 https://aclanthology.org/2023.eacl-main.9 (Association for Computational Linguistics, 2023).Gilbert, N. & Troitzsch, K. Simulation for the Social Scientist (McGraw-Hill Education, 2005).Gilbert, N. Emergence in social simulation. In Artificial Societies: The Computer Simulation of Social Life (eds Gilbert, N. & Conte, R.) 144–156 (UCL Press, 1995).Epstein, J. M. Generative Social Science: Studies in Agent-Based Computational Modeling (Princeton Univ. Press, 2006).Shadish, W. R., Cook, T. D. & Campbell, D. T. Experimental and Quasi-Experimental Designs for Generalized Causal Inference 2nd edn (Houghton Mifflin, 2002).Acemoglu, D. & Autor, D. Skills, tasks and technologies: Implications for employment and earnings. In Handbook of Labor Economics (eds Ashenfelter, O. & Card, D.) Vol. 4B, 1043–1171 (Elsevier, 2011).Chen, E. K., Belkin, M., Bergen, L. & Danks, D. Does AI already have human-level intelligence? The evidence is clear. Nature 650, 36–40 (2026).Article Google Scholar Brynjolfsson, E., Mitchell, T. & Rock, D. What can machines learn and what does it mean for occupations and the economy? In AEA Papers and Proceedings Vol. 108, 43–47 (American Economic Association, 2018).Howison, J., Wiggins, A. & Crowston, K. Validity issues in the use of social network analysis with digital trace data. J. Assoc. Inf. Syst. 12, 767–797 https://doi.org/10.17705/1jais.00282 (2011).Article Google Scholar Loukas, L., Stogiannidis, I., Diamantopoulos, O., Malakasiotis, P. & Vassos, S. Making LLMs worth every penny: resource-limited text classification in banking. In Proc. Fourth ACM International Conference on AI in Finance 392–400 (ACM, 2023).Törnberg, P. Large language models outperform expert coders and supervised classifiers at annotating political social media messages. Soc. Sci. Comput. Rev. 43, 1181–1195 https://doi.org/10.1177/08944393241286471 (2025).Article Google Scholar Eloundou, T., Manning, S., Mishkin, P. & Rock, D. GPTs are GPTs: labor market impact potential of LLMs. Science 384, 1306–1308 (2024).Article Google Scholar Chen, W. et al. AgentVerse: facilitating multi-agent collaboration and exploring emergent behaviors. In The Twelfth International Conference on Learning Representations https://openreview.net/forum?id=EHg5GDnyq1 (2024).Lewis, P. et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Adv. Neural Inf. Process. Syst. 33, 9459–9474 (2020).Google Scholar Madaan, A. et al. Self-refine: iterative refinement with self-feedback. Adv. Neural Inf. Process. Syst. 36, 46534–46594 (2023).Article Google Scholar Woolley, A. W., Chabris, C. F., Pentland, A., Hashmi, N. & Malone, T. W. Evidence for a collective intelligence factor in the performance of human groups. Science 330, 686–688 (2010).Article Google Scholar Du, Y., Li, S., Torralba, A., Tenenbaum, J. B. & Mordatch, I. Improving factuality and reasoning in language models through multiagent debate. In Proc. 41st International Conference on Machine Learning Vol. 235, 11733–11763 (PMLR, 2024).Liang, T. et al. Encouraging divergent thinking in large language models through multi-agent debate. In Proc. 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y. et al.) 17889–17904 https://aclanthology.org/2024.emnlp-main.992/ (Association for Computational Linguistics, 2024).Chen, J., Saha, S. & Bansal, M. ReConcile: round-table conference improves reasoning via consensus among diverse LLMs. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Ku, L.-W. et al.) 7066–7085 (Association for Computational Linguistics, 2024).Wang, Q., Wang, Z., Su, Y., Tong, H. & Song, Y. Rethinking the bounds of LLM reasoning: are multi-agent discussions the key? In Proc. 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Ku, L.-W. et al.) 6106–6131 https://aclanthology.org/2024.acl-long.331/ (Association for Computational Linguistics, 2024).Zhang, Y. et al. Chain of agents: large language models collaborating on long-context tasks. Adv. Neural Inf. Process. Syst. 37, 132208–132237 (2024).Article Google Scholar Hong, S. et al. MetaGPT: meta programming for a multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations https://openreview.net/forum?id=VtmBAGCN7o (2024).Shao, E. et al. SciSciGPT: advancing human–AI collaboration in the science of science. Nat. Comput. Sci. 6, 301–315 https://doi.org/10.1038/s43588-025-00906-6 (2026).Article Google Scholar Liu, X. et al. AgentBench: evaluating LLMs as agents. In The Twelfth International Conference on Learning Representations https://openreview.net/forum?id=zAdUBOaCTQ (2024).Brynjolfsson, E. The Turing trap: the promise & peril of human-like artificial intelligence. Daedalus 151, 272–287 (2022).Article Google Scholar Autor, D. H. Why are there still so many jobs? The history and future of workplace automation. J. Econ. Persp. 29, 3–30 (2015).Article Google Scholar Dong, M., Conway, J. R., Bonnefon, J.-F., Shariff, A. & Rahwan, I. Fears about artificial intelligence across 20 countries and six domains of application. Am. Psychol. 81, 53–67 (2026).Article Google Scholar Brynjolfsson, E., Li, D. & Raymond, L. Generative AI at work. Q. J. Econ. 140, 889–942 https://doi.org/10.1093/qje/qjae044 (2025).Article Google Scholar Noy, S. & Zhang, W. Experimental evidence on the productivity effects of generative artificial intelligence. Science 381, 187–192 https://doi.org/10.1126/science.adh2586 (2023).Article Google Scholar Dell’Acqua, F. et al. Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organ. Sci. 37, 403–423 (2026).Article Google Scholar Brown, E. & Cairns, P. A grounded investigation of game immersion. In CHI ’04 Extended Abstracts on Human Factors in Computing Systems 1297–1300 https://doi.org/10.1145/985921.986048 (Association for Computing Machinery, 2004).Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G. & King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 17, 195 https://doi.org/10.1186/s12916-019-1426-2 (2019).Article Google Scholar Gottweis, J. et al. Accelerating scientific discovery with Co-Scientist. Nature 655, 487–496 https://doi.org/10.1038/s41586-026-10644-y (2026).Article Google Scholar Bronfenbrenner, U. Toward an experimental ecology of human development. Am. Psychol. 32, 513–531 (1977).Article Google Scholar Bonnefon, J.-F., Shariff, A. & Rahwan, I. The social dilemma of autonomous vehicles. Science 352, 1573–1576 (2016).Article Google Scholar Awad, E. et al. The Moral Machine experiment. Nature 563, 59–64 (2018).Article Google Scholar Ameisen, E. et al. Circuit tracing: revealing computational graphs in language models. Transformer Circuits Thread https://transformer-circuits.pub/2025/attribution-graphs/methods.html (2025).Chen, R., Arditi, A., Sleight, H., Evans, O. & Lindsey, J. Persona vectors: monitoring and controlling character traits in language models. Preprint at https://doi.org/10.48550/arXiv.2507.21509 (2025).Wei, J. et al. Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inf. Process. Syst. 35, 24824–24837 (2022).Article Google Scholar Serapio-García, G. et al. A psychometric framework for evaluating and shaping personality traits in large language models. Nat. Mach. Intell. 7, 1954–1968 https://doi.org/10.1038/s42256-025-01115-6 (2025).Article Google Scholar Chen, Y., Liu, T. X., Shan, Y. & Zhong, S. The emergence of economic rationality of GPT. Proc. Natl Acad. Sci. USA 120, e2316205120 (2023).Article Google Scholar Strachan, J., Albergo, D., Borghini, G. et al. Testing theory of mind in large language models and humans. Nat. Hum. Behav. 8, 1285–1295 https://doi.org/10.1038/s41562-024-01882-z (2024).Article Google Scholar Tsvetkova, M., Yasseri, T., Pescetelli, N. & Werner, T. A new sociology of humans and machines. Nat. Hum. Behav. 8, 1864–1876 https://doi.org/10.1038/s41562-024-02001-8 (2024).Article Google Scholar Hagendorff, T. et al. Machine psychology. Preprint at https://doi.org/10.48550/arXiv.2303.13988 (2024).Brinkmann, L. et al. Machine culture. Nat. Hum. Behav. 7, 1855–1868 (2023).Article Google Scholar Grisold, T., Berente, N. & Seidel, S. Guardrails for human-AI ecologies: norm-based coordination and design for predictability. Manage. Inf. Syst. Q. 49, 1239–1266 https://doi.org/10.25300/MISQ/2025/18058 (2025).Article Google Scholar Cheung, V., Maier, M. & Lieder, F. Large language models show amplified cognitive biases in moral decision-making. Proc. Natl Acad. Sci. USA 122, e2412015122 https://doi.org/10.1073/pnas.2412015122 (2025).Article Google Scholar Hagendorff, T. Deception abilities emerged in large language models. Proc. Natl Acad. Sci. USA 121, e2317967121 https://doi.org/10.1073/pnas.2317967121 (2024).Article MathSciNet Google Scholar Röttger, P. et al. Political compass or spinning arrow? Towards more meaningful evaluations for values and opinions in large language models. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Ku, L.-W. et al.) 15295–15311 https://aclanthology.org/2024.acl-long.816/ (Association for Computational Linguistics, 2024).Peng, R. D. Reproducible research in computational science. Science 334, 1226–1227 (2011).Article Google Scholar Feuerriegel, S. et al. A reporting checklist for large language models in behavioural science. Nat. Hum. Behav. 10, 1182–1186 https://doi.org/10.1038/s41562-026-02492-7 (2026).Article Google Scholar Grossmann, I. et al. AI and the transformation of social science research. Science 380, 1108–1109 (2023).Article Google Scholar Ziems, C. et al. Can large language models transform computational social science? Comput. Linguist. 50, 237–291 https://doi.org/10.1162/coli_a_00502 (2024).Article Google Scholar Xu, R. et al. AI for social science and social science of AI: a survey. Inf. Process. Manage. 61, 103665 (2024).Article Google Scholar Bail, C. A. Can generative AI improve social science? Proc. Natl Acad. Sci. USA 121, e2314021121 (2024).Article Google Scholar Shrestha, P. et al. Beyond WEIRD: can synthetic survey participants substitute for humans in global policy research? Behav. Sci. Policy 10, 26–45 (2024).Article Google Scholar Groves, R. M. Survey Errors and Survey Costs (John Wiley & Sons, 1989).Brand, J., Israeli, A. & Ngwe, D. Using LLMs for Market Research Working Paper (Harvard Business School Marketing Unit, 2023).Goli, A. & Singh, A. Frontiers: can large language models capture human preferences? Market. Sci. 43, 709–722 (2024).Article Google Scholar Arora, N., Chakraborty, I. & Nishimura, Y. AI–human hybrids for marketing research: leveraging large language models (LLMs) as collaborators. J. Market. 89, 43–70 https://doi.org/10.1177/00222429241276529 (2025).Article Google Scholar Larooij, M. & Törnberg, P. Validation is the central challenge for generative social simulation: a critical review of LLMs in agent-based modeling. Artif. Intell. Rev. 59, 15 (2025).Article Google Scholar Pataranutaporn, P., Powdthavee, N., Archiwaranguprok, C. & Maes, P. Simulating human well-being with large language models: systematic validation and misestimation across 64,000 individuals from 64 countries. Proc. Natl Acad. Sci. USA 122, e2519394122 https://doi.org/10.1073/pnas.2519394122 (2025).Article Google Scholar Cui, Z., Li, N. & Zhou, H. A large-scale replication of scenario-based experiments in psychology and management using large language models. Nat. Comput. Sci. 5, 627–634 https://doi.org/10.1038/s43588-025-00840-7 (2025).Article Google Scholar Kolluri, A., Wu, S., Park, J. S. & Bernstein, M. S. Finetuning LLMs for human behavior prediction in social science experiments. In Proc. 2025 Conference on Empirical Methods in Natural Language Processing (eds Christodoulopoulos, C. et al.) 30096–30111 (Association for Computational Linguistics, 2025).Park, J. S. et al. LLM agents grounded in self-reports enable general-purpose simulation of individuals. Preprint at https://doi.org/10.48550/arXiv.2411.10109 (2026).Ferraz de Arruda, H., Gracia Lázaro, C., Aleta, A. & Moreno, Y. Collective cooperation without individual fidelity in LLM agents. Preprint at https://doi.org/10.48550/arXiv.2606.30454 (2026).Cao, Y. et al. Specializing large language models to simulate survey response distributions for global populations. In Proc. 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (eds Chiruzzo, L. et al.) 3141–3154 https://aclanthology.org/2025.naacl-long.162/ (Association for Computational Linguistics, 2025).Box, G. E. P. Robustness in the strategy of scientific model building. In Robustness in Statistics (eds Launer, R. L. & Wilkinson, G. N.) 201–236 https://www.sciencedirect.com/science/article/pii/B9780124381506500182 (Academic Press, 1979).Rilla, R., Werner, T., Yakura, H., Rahwan, I. & Nussberger, A.-M. Recognising and mitigating LLM pollution in online behavioural research. Nat. Commun. 17, 5578 https://doi.org/10.1038/s41467-026-74621-9 (2026).Article Google Scholar Laird, J. E. & van Lent, M. Human-level AI’s killer application: interactive computer games. AI Mag. 22, 15–25 (2001).Google Scholar Campbell, M., Hoane, A. J. & Hsu, F.-h Deep Blue. Artif. Intell. 134, 57–83 (2002).Article Google Scholar Silver, D. et al. Mastering the game of Go with deep neural networks and tree search. Nature 529, 484–489 (2016).Article Google Scholar Isaacson, W. Einstein: His Life and Universe (Simon & Schuster, 2007).