References¶
Every work the registry cites, with its DOI or arXiv identifier.
Abdulhai, Cheng, Clay, Althoff, Levine, Jaques (2025). Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning. arXiv:2511.00222
Aher, Arriaga, Kalai (2023). Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. Proceedings of the 40th International Conference on Machine Learning. arXiv:2208.10264
Ahmad, Takemoto (2025). Large-Scale Moral Machine Experiment on Large Language Models. PLOS ONE. doi:10.1371/journal.pone.0322776
Ahn, Mao, Lee (2026). Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human Surrogates. arXiv:2608.29455
Ahnert, Haensch, Plank, Strohmaier (2026). Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2026.acl-long.1927
Ali, Silvey (1966). A General Class of Coefficients of Divergence of One Distribution from Another. Journal of the Royal Statistical Society: Series B (Methodological). doi:10.1111/j.2517-6161.1966.tb00626.x
AlKhamissi, ElNokrashy, AlKhamissi, Diab (2024). Investigating Cultural Alignment of Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2024.acl-long.671
An, Park, Shin (2026). When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing. Proceedings of the 35th ACM International Conference on Information and Knowledge Management. arXiv:2608.22438
Anderson, Darling (1952). Asymptotic Theory of Certain ``Goodness of Fit'' Criteria Based on Stochastic Processes. The Annals of Mathematical Statistics. doi:10.1214/aoms/1177729437
Anderson (1962). On the Distribution of the Two-Sample Cramér-von Mises Criterion. The Annals of Mathematical Statistics. doi:10.1214/aoms/1177704477
Angelopoulos, Bates, Fannjiang, Jordan, Zrnic (2023). Prediction-Powered Inference. Science. doi:10.1126/science.adi6000
Angelopoulos, Duchi, Zrnic (2023). PPI++: Efficient Prediction-Powered Inference. arXiv:2311.01453
Argyle, Busby, Fulda, Gubler, Rytting, Wingate (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis. doi:10.1017/pan.2023.2
Ashokkumar, Hewitt, Ghezae, Willer (2026). Large Language Models Can Predict the Results of Social Science Experiments. Nature. doi:10.1038/s41586-026-10742-x
Atari, Xue, Park, Blasi, Henrich (2023). Which Humans?. PsyArXiv preprint. doi:10.31234/osf.io/5b26t
Bae (2026). Marginal Alignment Does Not Guarantee Joint-Distribution Fidelity: An Official-Reference Audit of Nemotron-Personas-Korea with Cross-Locale Replication. arXiv:2606.12433
Beck, Schuff, Lauscher, Gurevych (2024). Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2024.eacl-long.159
Bhattacharyya (1946). On a Measure of Divergence between Two Multinomial Populations. Sankhy\=a: The Indian Journal of Statistics. link
Binz, Akata, Bethge, Brändle, Callaway, Coda-Forno, others (2025). A Foundation Model to Predict and Capture Human Cognition. Nature. doi:10.1038/s41586-025-09215-4
Bisbee, Clinton, Dorff, Kenkel, Larson (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis. doi:10.1017/pan.2024.5
Boelaert, Coavoux, Ollion, Petev, Präg (2025). Machine Bias. How Do Generative Language Models Answer Opinion Polls?. Sociological Methods & Research. doi:10.1177/00491241251330582
Bonneel, Rabin, Peyré, Pfister (2015). Sliced and Radon Wasserstein Barycenters of Measures. Journal of Mathematical Imaging and Vision. doi:10.1007/s10851-014-0506-3
Bordalo, Coffman, Gennaioli, Shleifer (2016). Stereotypes. The Quarterly Journal of Economics. doi:10.1093/qje/qjw029
Brand, Israeli, Ngwe (2023). Using LLMs for Market Research. doi:10.2139/ssrn.4395751
Bray, Curtis (1957). An Ordination of the Upland Forest Communities of Southern Wisconsin. Ecological Monographs. doi:10.2307/1942268
Brier (1950). Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review. doi:10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2
Broska, Howes, van Loon (2025). The Mixed Subjects Design: Treating Large Language Models as Potentially Informative Observations. Sociological Methods & Research. doi:10.1177/00491241251326865
Brown (1910). Some Experimental Results in the Correlation of Mental Abilities. British Journal of Psychology. doi:10.1111/j.2044-8295.1910.tb00207.x
Cao, Liu, Arora, Augenstein, Röttger, Hershcovich (2025). Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). doi:10.18653/v1/2025.naacl-long.162
Ceron, Falk, Barić, Nikolaev, Padó (2024). Beyond Prompt Brittleness: Evaluating the Reliability and Consistency of Political Worldviews in LLMs. Transactions of the Association for Computational Linguistics. doi:10.1162/tacl_a_00710
Chen, Zhu, Zheng (2026). When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses. arXiv:2607.26348
Cheng, Piccardi, Yang (2023). CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. doi:10.18653/v1/2023.emnlp-main.669
Cheng, Durmus, Jurafsky (2023). Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2023.acl-long.84
Chiu, Jiang, Lin, Park, Li, Ravi, Bhatia, Antoniak, Tsvetkov, Shwartz, Choi (2025). CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2025.acl-long.1247
Choi, Kim, Pugalenthi, Chen, Huang (2026). Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data. arXiv:2606.28963
Choulakian, Lockhart, Stephens (1994). Cramér-von Mises Statistics for Discrete Distributions. Canadian Journal of Statistics. doi:10.2307/3315828
Clogg, Petkova, Haritou (1995). Statistical Methods for Comparing Regression Coefficients Between Models. American Journal of Sociology. doi:10.1086/230638
Conover (1972). A Kolmogorov Goodness-of-Fit Test for Discontinuous Distributions. Journal of the American Statistical Association. doi:10.1080/01621459.1972.10481254
Cramér (1928). On the Composition of Elementary Errors. Skandinavisk Aktuarietidskrift. doi:10.1080/03461238.1928.10416862
Cramér (1946). Mathematical Methods of Statistics. Princeton University Press. doi:10.1515/9781400883868
Cronbach (1951). Coefficient Alpha and the Internal Structure of Tests. Psychometrika. doi:10.1007/BF02310555
Cui, Li, Zhou (2025). A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management Using Large Language Models. Nature Computational Science. doi:10.1038/s43588-025-00840-7
da Silva, Lukk, Sutani, Moturu, Yang, Silver, Ratto, Silva (2026). Population Fidelity: Evaluating Population Representativeness in LLMs. arXiv:2609.36253
Daryani, Bogen, Daepp (2026). Accurate in Space, Unreliable in Time: How LLMs Represent National Cultural Change. arXiv:2609.01902
DataViking-Tech (2026). SynthBench: Open Benchmark for Synthetic Survey Respondent Quality. link
Dominguez-Olmedo, Hardt, Mendler-Dünner (2024). Questioning the Survey Responses of Large Language Models. Advances in Neural Information Processing Systems. doi:10.52202/079017-1458
Doudkin (2026). Marginal Fidelity Does Not Establish User Simulation in Demographic Synthetic Survey Panels: Response Contracts, Support Collapse and Conditioning Failure. arXiv:2609.07305
Durmus, Nguyen, Liao, Schiefer, Askell, Bakhtin, Chen, Hatfield-Dodds, Hernandez, Joseph, Lovitt, McCandlish, Sikder, Tamkin, Thamkul, Kaplan, Clark, Ganguli (2024). Towards Measuring the Representation of Subjective Global Opinions in Language Models. First Conference on Language Modeling (COLM). arXiv:2306.16388
Dwork, Hardt, Pitassi, Reingold, Zemel (2012). Fairness through Awareness. Proceedings of the 3rd Innovations in Theoretical Computer Science Conference. doi:10.1145/2090236.2090255
Efron (1979). Bootstrap Methods: Another Look at the Jackknife. The Annals of Statistics. doi:10.1214/aos/1176344552
Efron, Tibshirani (1993). An Introduction to the Bootstrap. Chapman & Hall. doi:10.1007/978-1-4899-4541-9
Egami, Hinck, Stewart, Wei (2023). Using Imperfect Surrogates for Downstream Inference: Design-Based Supervised Learning for Social Science Applications of Large Language Models. Advances in Neural Information Processing Systems. doi:10.52202/075280-3000
Endres, Schindelin (2003). A New Metric for Probability Distributions. IEEE Transactions on Information Theory. doi:10.1109/TIT.2003.813506
Epstein (1969). A Scoring System for Probability Forecasts of Ranked Categories. Journal of Applied Meteorology. doi:10.1175/1520-0450(1969)008<0985:ASSFPF>2.0.CO;2
Feng, Sorensen, Liu, Fisher, Park, Choi, Tsvetkov (2024). Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. doi:10.18653/v1/2024.emnlp-main.240
Fleiss (1971). Measuring Nominal Scale Agreement among Many Raters. Psychological Bulletin. doi:10.1037/h0031619
Fornell, Larcker (1981). Evaluating Structural Equation Models with Unobservable Variables and Measurement Error. Journal of Marketing Research. doi:10.1177/002224378101800104
Freeman (1965). Elementary Applied Statistics. Wiley. No DOI
Friedman, Dieng (2023). The Vendi Score: A Diversity Evaluation Metric for Machine Learning. Transactions on Machine Learning Research. arXiv:2210.02410
Gao, Han, Liang (2026). How Well Do LLMs Predict Human Behavior? A Measure of their Pretrained Knowledge. arXiv:2601.12343
Gelman, Carlin (2014). Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors. Perspectives on Psychological Science. doi:10.1177/1745691614551642
Gibbs, Su (2002). On Choosing and Bounding Probability Metrics. International Statistical Review. doi:10.1111/j.1751-5823.2002.tb00178.x
Gneiting, Raftery (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association. doi:10.1198/016214506000001437
Good (1952). Rational Decisions. Journal of the Royal Statistical Society: Series B. doi:10.1111/j.2517-6161.1952.tb00104.x
Gower (1966). Some Distance Properties of Latent Root and Vector Methods Used in Multivariate Analysis. Biometrika. doi:10.1093/biomet/53.3-4.325
Gower (1975). Generalized Procrustes Analysis. Psychometrika. doi:10.1007/BF02291478
Gretton, Borgwardt, Rasch, Schölkopf, Smola (2012). A Kernel Two-Sample Test. Journal of Machine Learning Research. link
Gui, Toubia (2023). The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective. arXiv:2312.15524
Guo, Pleiss, Sun, Weinberger (2017). On Calibration of Modern Neural Networks. Proceedings of the 34th International Conference on Machine Learning. arXiv:1706.04599
Hardt, Price, Srebro (2016). Equality of Opportunity in Supervised Learning. Advances in Neural Information Processing Systems. arXiv:1610.02413
Hayes, Krippendorff (2007). Answering the Call for a Standard Reliability Measure for Coding Data. Communication Methods and Measures. doi:10.1080/19312450709336664
Hellinger (1909). Neue Begründung der Theorie quadratischer Formen von unendlichvielen Veränderlichen. Journal für die reine und angewandte Mathematik. doi:10.1515/crll.1909.136.210
Hill (1973). Diversity and Evenness: A Unifying Notation and Its Consequences. Ecology. doi:10.2307/1934352
Holm (1979). A Simple Sequentially Rejective Multiple Test Procedure. Scandinavian Journal of Statistics. link
Hong, Caldas, Leqi (2025). Hypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings. arXiv:2506.14997
Horton, Filippas, Manning (2023). Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?. doi:10.3386/w31122
Hu, Collier (2024). Quantifying the Persona Effect in LLM Simulations. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2024.acl-long.554
Hu, Lian, Xiao, Xiong, Lei, Wang, Ding, Xiao, Yuan, Xie (2025). Population-Aligned Persona Generation for LLM-based Social Simulation. arXiv:2509.10127
Hu, Baumann, Lupo, Collier, Hovy, Röttger (2026). SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors. International Conference on Learning Representations (ICLR). arXiv:2510.17516
Huang, Wu, Wang (2025). Uncertainty Quantification for LLM-Based Survey Simulations. Proceedings of the 42nd International Conference on Machine Learning. arXiv:2502.17773
Hullman, Broska, Sun, Shaw (2026). This Human Study Did Not Involve Human Subjects: Validating LLM Simulations as Behavioral Evidence. arXiv:2602.15785
Hung, Midha, Wu, Zhang (2026). Multi-dimensional Bias in Modeling Multi-dimensional Preferences: Evaluating the Ability of Synthetic Agents to Replace Human Participants in Conjoint Experiments. arXiv:2609.04243
Inman, Bradley (1989). The Overlapping Coefficient as a Measure of Agreement between Probability Distributions and Point Estimation of the Overlap of Two Normal Densities. Communications in Statistics - Theory and Methods. doi:10.1080/03610928908830127
Iyengar, Lin, Wang (2025). Model-Free Assessment of Simulator Fidelity via Quantile Curves. arXiv:2512.05024
Jackson (1995). PROTEST: A PROcrustean Randomization TEST of Community Environment Concordance. Écoscience. doi:10.1080/11956860.1995.11682297
Jang, Lee, Kim (2026). Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe. arXiv:2607.25292
Jeoung, Ge, Wang, Diesner (2025). Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes. International Conference on Learning Representations (ICLR). arXiv:2501.14294
Jia, Chen, Sharma, Diaz-Rodriguez (2026). When Can Digital Personas Reliably Approximate Human Survey Findings?. arXiv:2605.10659
Kailath (1967). The Divergence and Bhattacharyya Distance Measures in Signal Selection. IEEE Transactions on Communication Technology. doi:10.1109/TCOM.1967.1089532
Kantorovitch (1958). On the Translocation of Masses. Management Science. doi:10.1287/mnsc.5.1.1
Kendall (1938). A New Measure of Rank Correlation. Biometrika. doi:10.1093/biomet/30.1-2.81
Kendall (1945). The Treatment of Ties in Ranking Problems. Biometrika. doi:10.1093/biomet/33.3.239
Khan, Casper, Hadfield-Menell (2025). Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs. Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. doi:10.1145/3715275.3732147
Kim, Lee (2023). AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction. arXiv:2305.09620
Kirk, Whitefield, Röttger, Bean, Margatina, Ciro, Mosquera, Bartolo, Williams, He, Vidgen, Hale (2024). The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models. Advances in Neural Information Processing Systems (Datasets and Benchmarks Track). arXiv:2404.16019
Kish (1965). Survey Sampling. Wiley. No DOI
Klugman, Rodríguez, Choi (2011). The HDI 2010: New Controversies, Old Critiques. The Journal of Economic Inequality. doi:10.1007/s10888-011-9178-z
Knight (1966). A Computer Method for Calculating Kendall's Tau with Ungrouped Data. Journal of the American Statistical Association. doi:10.1080/01621459.1966.10480879
Kolluri, Wu, Park, Bernstein (2025). Finetuning LLMs for Human Behavior Prediction in Social Science Experiments. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. doi:10.18653/v1/2025.emnlp-main.1530
Kolmogorov (1933). Sulla determinazione empirica di una legge di distribuzione. Giornale dell'Istituto Italiano degli Attuari. No DOI
Kovač, Portelas, Sawayama, Dominey, Oudeyer (2024). Stick to Your Role! Stability of Personal Values Expressed in Large Language Models. PLOS ONE. doi:10.1371/journal.pone.0309114
Krippendorff (1970). Estimating the Reliability, Systematic Error and Random Error of Interval Data. Educational and Psychological Measurement. doi:10.1177/001316447003000105
Krsteski, Russo, Chang, West, Gligorić (2026). Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2026.acl-long.498
Kullback, Leibler (1951). On Information and Sufficiency. The Annals of Mathematical Statistics. doi:10.1214/aoms/1177729694
Lam, Voo, Ma (2026). The Conditional Superiority of Fast Silicon Sampling. arXiv:2608.14079
Lee, Wang (2026). Can Open-Weight Large Language Models (LLMs) Simulate Human Survey Populations? A Cross-Instrument Calibration Study. arXiv:2609.32638
Legendre, Fortin (2010). Comparison of the Mantel Test and Alternative Approaches for Detecting Complex Multivariate Relationships in the Spatial Analysis of Genetic Data. Molecular Ecology Resources. doi:10.1111/j.1755-0998.2010.02866.x
Levin, Peres, Wilmer (2009). Markov Chains and Mixing Times. American Mathematical Society. doi:10.1090/mbk/058
Li, Li, Qiu (2025). ChatGPT is not A Man but Das Man: Representativeness and Structural Consistency of Silicon Samples Generated by Large Language Models. arXiv:2507.02919
Li, Ji (2026). Statistical Realism Is Not Evidence That LLMs Can Estimate Treatment Effects in Social Science Experiments. arXiv:2604.02458
Liang, Zeger (1986). Longitudinal Data Analysis Using Generalized Linear Models. Biometrika. doi:10.1093/biomet/73.1.13
Libovick\'y (2026). On the Credibility of Evaluating LLMs using Survey Questions. Proceedings of the First Workshop on Multilingual Multicultural Evaluation. doi:10.18653/v1/2026.mme-main.2
Lin (1989). A Concordance Correlation Coefficient to Evaluate Reproducibility. Biometrics. doi:10.2307/2532051
Lin (1991). Divergence Measures Based on the Shannon Entropy. IEEE Transactions on Information Theory. doi:10.1109/18.61115
Lin, Yun, Matarić, Canny, Gretton, D'Amour (2026). The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study. arXiv:2605.20767
Liu, Diab, Fried (2024). Evaluating Large Language Model Biases in Persona-Steered Generation. Findings of the Association for Computational Linguistics: ACL 2024. doi:10.18653/v1/2024.findings-acl.586
Liu, Bhandari, Pardos (2025). Leveraging LLM Respondents for Item Evaluation: A Psychometric Analysis. British Journal of Educational Technology. doi:10.1111/bjet.13570
Lorenzo-Seva, ten Berge (2006). Tucker's Congruence Coefficient as a Meaningful Index of Factor Similarity. Methodology. doi:10.1027/1614-2241.2.2.57
Lukauskas, Sarkauskait.e (2026). Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents. arXiv:2608.14606
Maier, Aslak, Fiaschi, Rismal, Fletcher, Luhmann, Dow, Pappas, Wiecki (2025). LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings. arXiv:2510.08338
Mantel (1967). The Detection of Disease Clustering and a Generalized Regression Approach. Cancer Research. No DOI; PubMed 6018555
Masoud, Liu, Ferianc, Treleaven, Rodrigues (2025). Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions. Proceedings of the 31st International Conference on Computational Linguistics. link
McGraw, Wong (1996). Forming Inferences about Some Intraclass Correlation Coefficients. Psychological Methods. doi:10.1037/1082-989X.1.1.30
Mei, Xie, Yuan, Jackson (2024). A Turing Test of Whether AI Chatbots Are Behaviorally Similar to Humans. Proceedings of the National Academy of Sciences. doi:10.1073/pnas.2313925121
Meister, Guestrin, Hashimoto (2025). Benchmarking Distributional Alignment of Large Language Models. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). doi:10.18653/v1/2025.naacl-long.2
Miehling, Desmond, Natesan Ramamurthy, Daly, Varshney, Farchi, Dognin, Rios, Bouneffouf, Liu, Sattigeri (2025). Evaluating the Prompt Steerability of Large Language Models. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). doi:10.18653/v1/2025.naacl-long.400
Mincer, Zarnowitz (1969). The Evaluation of Economic Forecasts. Economic Forecasts and Expectations: Analysis of Forecasting Behavior and Performance. link
Monge (1781). Mémoire sur la théorie des déblais et des remblais. Histoire de l'Académie Royale des Sciences de Paris. No DOI; printed 1784
Moon, Abdulhai, Kang, Suh, Soedarmadji, Behar, Chan (2024). Virtual Personas for Language Models via an Anthology of Backstories. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. doi:10.18653/v1/2024.emnlp-main.1110
Moore, Deshpande, Yang (2024). Are Large Language Models Consistent over Value-laden Questions?. Findings of the Association for Computational Linguistics: EMNLP 2024. doi:10.18653/v1/2024.findings-emnlp.891
Mukherjee, Adilazuarda, Sitaram, Bali, Aji, Choudhury (2024). Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. doi:10.18653/v1/2024.emnlp-main.884
Murphy (1971). A Note on the Ranked Probability Score. Journal of Applied Meteorology. doi:10.1175/1520-0450(1971)010<0155:ANOTRP>2.0.CO;2
Naeini, Cooper, Hauskrecht (2015). Obtaining Well Calibrated Probabilities Using Bayesian Binning. Proceedings of the AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v29i1.9602
Nakagawa, Schielzeth (2013). A General and Simple Method for Obtaining $R^2$ from Generalized Linear Mixed-Effects Models. Methods in Ecology and Evolution. doi:10.1111/j.2041-210x.2012.00261.x
Neumann, De-Arteaga, Fazelpour (2026). Should You Use LLMs to Simulate Opinions? Quality Checks for Early-Stage Deliberation. Proceedings of the AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v40i46.41254
Nguyen, Ahmad (2026). Measurement Validity in LLM Cultural Alignment. arXiv:2608.29266
Ozkan (2026). Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling. arXiv:2607.18310
Pagan, Törnberg, Bail, Hannák, Barrie (2025). Computational Turing Test Reveals Systematic Differences Between Human and AI Language. arXiv:2511.04195
Park, Zou, Kamphorst, Egan, Shaw, Hill, Cai, Morris, Liang, Willer, Bernstein (2024). LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals. arXiv:2411.10109
Pearson (1895). Note on Regression and Inheritance in the Case of Two Parents. Proceedings of the Royal Society of London. doi:10.1098/rspl.1895.0041
Pearson (1900). On the Criterion that a Given System of Deviations from the Probable in the Case of a Correlated System of Variables is Such that it can be Reasonably Supposed to have Arisen from Random Sampling. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science. doi:10.1080/14786440009463897
Pearson (1900). Mathematical Contributions to the Theory of Evolution. VII. On the Correlation of Characters not Quantitatively Measurable. Philosophical Transactions of the Royal Society of London. Series A. doi:10.1098/rsta.1900.0022
Peng, Gui, Brucks, Merlau, Fan, Ben Sliman, Johnson, Althenayyan, Bellezza, Donati, Fong, Friedman, Guevara, Hussein, Jerath, Kogut, Kumar, Lane, Li, Morwitz, Netzer, Perkowski, Toubia (2025). Digital Twins as Funhouse Mirrors: Five Key Distortions. arXiv:2509.19088
Peres-Neto, Jackson (2001). How Well Do Multivariate Data Sets Match? The Advantages of a Procrustean Superimposition Approach over the Mantel Test. Oecologia. doi:10.1007/s004420100720
Persson, Schultzberg, Ankargren (2026). Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference. arXiv:2606.17165
Pesaran, Timmermann (1992). A Simple Nonparametric Test of Predictive Performance. Journal of Business & Economic Statistics. doi:10.1080/07350015.1992.10509922
Petrov, Serapio-García, Rentfrow (2024). Limited Ability of LLMs to Simulate Human Psychological Behaviours: A Psychometric Analysis. arXiv:2405.07248
Pfänder (2026). The Silicon Sample Benchmark: Predicting a Megastudy before the Results Are Revealed. link
Pielou (1966). The Measurement of Diversity in Different Types of Biological Collections. Journal of Theoretical Biology. doi:10.1016/0022-5193(66)90013-0
Pitman (1937). Significance Tests Which May Be Applied to Samples from Any Populations. Supplement to the Journal of the Royal Statistical Society. doi:10.2307/2984124
Poole-Dayan, Wu, Sorensen, Pei, Bakker (2026). Benchmarking Overton Pluralism in LLMs. International Conference on Learning Representations (ICLR). arXiv:2512.01351
Rao, Yerukola, Shah, Reinecke, Sap (2025). NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). doi:10.18653/v1/2025.naacl-long.120
Ren, Ye, Fang, Zhang, Song (2024). ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2024.acl-long.111
Rennard, Xypolopoulos (2026). Large Language Models Simulate Intersectional Synthetic Identities with a Budget of One to Two Dimensions. arXiv:2608.23005
Robert, Escoufier (1976). A Unifying Tool for Linear Multivariate Statistical Methods: The RV-Coefficient. Applied Statistics. doi:10.2307/2347233
Rubin (1987). Multiple Imputation for Nonresponse in Surveys. Wiley. doi:10.1002/9780470316696
Rubner, Tomasi, Guibas (2000). The Earth Mover's Distance as a Metric for Image Retrieval. International Journal of Computer Vision. doi:10.1023/A:1026543900054
Röttger, Hofmann, Pyatkin, Hinck, Kirk, Schütze, Hovy (2024). Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2024.acl-long.816
Samuel, Zou, Zhou, Chaudhari, Kalyan, Rajpurohit, Deshpande, Narasimhan, Murahari (2025). PersonaGym: Evaluating Persona Agents and LLMs. Findings of the Association for Computational Linguistics: EMNLP 2025. doi:10.18653/v1/2025.findings-emnlp.368
Santurkar, Durmus, Ladhak, Lee, Liang, Hashimoto (2023). Whose Opinions Do Language Models Reflect?. Proceedings of the 40th International Conference on Machine Learning. arXiv:2303.17548
Satterthwaite (1946). An Approximate Distribution of Estimates of Variance Components. Biometrics Bulletin. doi:10.2307/3002019
Scheffé (1947). A Useful Convergence Theorem for Probability Distributions. The Annals of Mathematical Statistics. doi:10.1214/aoms/1177730390
Scholz, Stephens (1987). K-Sample Anderson–Darling Tests. Journal of the American Statistical Association. doi:10.1080/01621459.1987.10478517
Schuirmann (1987). A Comparison of the Two One-Sided Tests Procedure and the Power Approach for Assessing the Equivalence of Average Bioavailability. Journal of Pharmacokinetics and Biopharmaceutics. doi:10.1007/BF01068419
Schwarz (1978). Estimating the Dimension of a Model. The Annals of Statistics. doi:10.1214/aos/1176344136
Sen, Lutz, Rogers, Garcia, Strohmaier (2025). Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs. Findings of the Association for Computational Linguistics: ACL 2025. arXiv:2511.01864
Sen, Ahnert, von der Heyde, Lasser, Wei, Strohmaier (2026). Total Simulated Survey Error: Designing and Diagnosing Survey Responses from Large Language Models. arXiv:2609.10280
Serapio-García, Safdari, Crepy, Sun, Fitz, Romero, Abdulhai, Faust, Matarić (2025). A Psychometric Framework for Evaluating and Shaping Personality Traits in Large Language Models. Nature Machine Intelligence. doi:10.1038/s42256-025-01115-6
Shankar, S P, Margapuri, Mazumder, Kumaraguru, Chakraborty (2026). Mind the Gap: Pitfalls of LLM Alignment with Asian Public Opinion. Proceedings of the International AAAI Conference on Web and Social Media. doi:10.1609/icwsm.v20i1.42740
Shannon (1948). A Mathematical Theory of Communication. Bell System Technical Journal. doi:10.1002/j.1538-7305.1948.tb01338.x
Shi, Haupt (2026). The Collapse of Heterogeneity in Silicon Philosophers. arXiv:2604.23575
Shrout, Fleiss (1979). Intraclass Correlations: Uses in Assessing Rater Reliability. Psychological Bulletin. doi:10.1037/0033-2909.86.2.420
Simpson (1949). Measurement of Diversity. Nature. doi:10.1038/163688a0
Skobelev, Fithian, Han (2026). Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs. arXiv:2609.16454
Smilde, Kiers, Bijlsma, Rubingh, van Erk (2009). Matrix Correlations for High-Dimensional Data: The Modified RV-Coefficient. Bioinformatics. doi:10.1093/bioinformatics/btn634
Smirnov (1948). Table for Estimating the Goodness of Fit of Empirical Distributions. The Annals of Mathematical Statistics. doi:10.1214/aoms/1177730256
Smouse, Long, Sokal (1986). Multiple Regression and Correlation Extensions of the Mantel Test of Matrix Correspondence. Systematic Zoology. doi:10.2307/2413122
Spearman (1904). The Proof and Measurement of Association between Two Things. The American Journal of Psychology. doi:10.2307/1412159
Spearman (1910). Correlation Calculated from Faulty Data. British Journal of Psychology. doi:10.1111/j.2044-8295.1910.tb00206.x
Suh, Jahanparast, Moon, Kang, Chang (2025). Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/2025.acl-long.1028
Sun, Lee, Nan, Zhao, Lee, Jansen, Kim (2024). Random Silicon Sampling: Simulating Human Sub-Population Opinion Using a Large Language Model Based on Group-Level Demographic Information. arXiv:2402.18144
Swain, Ballard (1991). Color Indexing. International Journal of Computer Vision. doi:10.1007/BF00130487
Székely, Rizzo, Bakirov (2007). Measuring and Testing Dependence by Correlation of Distances. The Annals of Statistics. doi:10.1214/009053607000000505
Székely, Rizzo (2013). Energy Statistics: A Class of Statistics Based on Distances. Journal of Statistical Planning and Inference. doi:10.1016/j.jspi.2013.03.018
Takemoto (2024). The Moral Machine Experiment on Large Language Models. Royal Society Open Science. doi:10.1098/rsos.231393
Tao, Viberg, Baker, Kizilcec (2024). Cultural Bias and Cultural Alignment of Large Language Models. PNAS Nexus. doi:10.1093/pnasnexus/pgae346
Tjuatja, Chen, Wu, Talwalkar, Neubig (2024). Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey Design. Transactions of the Association for Computational Linguistics. doi:10.1162/tacl_a_00685
Torgerson (1952). Multidimensional Scaling: I. Theory and Method. Psychometrika. doi:10.1007/BF02288916
Toubia, Gui, Peng, Merlau, Li, Chen (2025). Database Report: Twin-2K-500: A Data Set for Building Digital Twins of over 2,000 People Based on Their Answers to over 500 Questions. Marketing Science. doi:10.1287/mksc.2025.0262
Tucker (1951). A Method for Synthesis of Factor Analysis Studies. doi:10.21236/AD0047524
Vallender (1974). Calculation of the Wasserstein Distance Between Probability Distributions on the Line. Theory of Probability & Its Applications. doi:10.1137/1118101
Vaserstein (1969). Markov Processes over Denumerable Products of Spaces, Describing Large Systems of Automata. Problemy Peredachi Informatsii. No DOI; Math-Net.Ru ppi1811
Villani (2009). Optimal Transport: Old and New. Springer. doi:10.1007/978-3-540-71050-9
von der Heyde, Haensch, Wenz (2026). Vox Populi, Vox AI? Using Large Language Models to Estimate German Vote Choice. Social Science Computer Review. doi:10.1177/08944393251337014
von Mises (1928). Wahrscheinlichkeit, Statistik und Wahrheit. Springer. doi:10.1007/978-3-662-36230-3
Wang, Ma, Hu, Weber-Genzel, Röttger, Kreuter, Hovy, Plank (2024). ``My Answer is C'': First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models. Findings of the Association for Computational Linguistics: ACL 2024. doi:10.18653/v1/2024.findings-acl.441
Wang, Morgenstern, Dickerson (2025). Large Language Models That Replace Human Participants Can Harmfully Misportray and Flatten Identity Groups. Nature Machine Intelligence. doi:10.1038/s42256-025-00986-z
Wang, Hunt, Tang, Joseph (2026). When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability. arXiv:2609.07987
Welch (1947). The Generalization of `Student's' Problem When Several Different Population Variances Are Involved. Biometrika. doi:10.1093/biomet/34.1-2.28
Wherry (1931). A New Formula for Predicting the Shrinkage of the Coefficient of Multiple Correlation. The Annals of Mathematical Statistics. doi:10.1214/aoms/1177732951
Wilcoxon (1945). Individual Comparisons by Ranking Methods. Biometrics Bulletin. doi:10.2307/3001968
Williams, Weeber, Padó, Akbik (2026). Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs. Findings of the Association for Computational Linguistics: ACL 2026. doi:10.18653/v1/2026.findings-acl.236
Ye, Yoganarasimhan (2026). Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys. arXiv:2604.17267
Zhang, Lin, Sun, Qi, Yang, Chen, Lyu, Mou, Chen, Luo, Huang, Tang, Wei (2024). ElectionSim: Massive Population Election Simulation Powered by Large Language Model Driven Agents. arXiv:2410.20746
Zhao, Mondal, Tandon, Dillion, Gray, Gu (2024). WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). arXiv:2404.16308
Zhao, Dang, Grover (2024). Group Preference Optimization: Few-Shot Alignment of Large Language Models. International Conference on Learning Representations (ICLR). arXiv:2310.11523
Zheng, Zhou, Meng, Zhou, Huang (2024). Large Language Models Are Not Robust Multiple Choice Selectors. International Conference on Learning Representations (ICLR). arXiv:2309.03882