Measures¶
The registry lists 126 measures. Each names the work it comes from and the work that applied it to language-model survey responses; the status says whether a measure ships in 0.1, is planned, or is documented only. Python, R and the command line (popfidelity registry) read the same registry.
Population Fidelity Score¶
| Measure | Formula | Range | Origin | Applied to LLMs | Python / R | Status |
|---|---|---|---|---|---|---|
| Population Fidelity Score | (S_acc * S_adapt * S_struct)^(1/3); 0 if any component is 0 |
[0, 1], higher | da Silva et al. (2026); Klugman et al. (2011) | da Silva et al. (2026) | popfidelity.score / pfs_score |
0.1 |
| Accuracy component | S_acc = 1 - mean_i nEMD(s_i, m_i) |
[0, 1], higher | Boelaert et al. (2025); da Silva et al. (2026) | da Silva et al. (2026) | popfidelity.score_groups / pfs_score_groups |
0.1 |
| Adaptability component | A = median_{i<j} nEMD(m_i, m_j) / median_{i<j} nEMD(s_i, s_j); S_adapt = min(A, 1/A) |
[0, 1], higher | Boelaert et al. (2025); da Silva et al. (2026) | da Silva et al. (2026) | popfidelity.score_groups / pfs_score_groups |
0.1 |
| Structure component | S_struct = max(0, Spearman(d_s, d_m)) over the condensed pairwise nEMD vectors; 0 if the model is flat |
[0, 1], higher | Spearman (1904); Mantel (1967); da Silva et al. (2026) | da Silva et al. (2026) | popfidelity.score_groups / pfs_score_groups |
0.1 |
| Centre alignment | S_center = 1 - nEMD(mean_i s_i, mean_i m_i); reported beside PFS, not inside it |
[0, 1], higher | Boelaert et al. (2025) | Boelaert et al. (2025); da Silva et al. (2026) | popfidelity.score_groups / pfs_score_groups |
0.1 |
| Limiting component | argmin of S_acc, S_adapt, S_struct, skipping undefined components; the first wins ties |
accuracy | adaptability | structure, diagnostic | da Silva et al. (2026) | da Silva et al. (2026) | popfidelity.binding_component / pfs_binding_component |
0.1 |
| Signed PFS | sign(rho) * (S_acc * S_adapt * \|rho\|)^(1/3) |
[-1, 1], higher | da Silva et al. (2026) | da Silva et al. (2026) | popfidelity.variant('unclipped_structure') / pfs_variant('unclipped_structure') |
0.1 |
| PFS sensitivity variants | mean dispersion; one-sided min(A, 1); Pearson or Kendall structure; arithmetic mean; respondent-weighted E and C |
[0, 1], higher | da Silva et al. (2026) | da Silva et al. (2026) | popfidelity.VARIANTS / pfs_variant |
0.1 |
| Paired change in PFS | PFS(tuned) - PFS(base) with both scored on their common cells |
[-1, 1], higher | da Silva et al. (2026) | da Silva et al. (2026) | popfidelity.compare / pfs_compare |
0.1 |
| Cross-question PFS | geometric mean of per-question PFS; arithmetic means of components |
[0, 1], higher | da Silva et al. (2026); Klugman et al. (2011) | da Silva et al. (2026) | popfidelity.across_questions / pfs_across_questions |
0.1 |
Distances between answer distributions¶
| Measure | Formula | Range | Origin | Applied to LLMs | Python / R | Status |
|---|---|---|---|---|---|---|
| Normalised earth mover's distance | sum_k \|F_p(k) - F_q(k)\| / (K - 1) |
[0, 1], lower | Monge (1781); Kantorovitch (1958); Vaserstein (1969); Vallender (1974); Rubner et al. (2000) | Santurkar et al. (2023); Boelaert et al. (2025); da Silva et al. (2026) | popfidelity.metrics.distance(metric='nemd') / metric_distance(metric = 'nemd') |
0.1 |
| Earth mover's distance | sum_k \|F_p(k) - F_q(k)\| with unit spacing |
[0, K - 1], lower | Rubner et al. (2000); Villani (2009) | Suh et al. (2025); Cao et al. (2025) | popfidelity.metrics.distance(metric='emd') / metric_distance(metric = 'emd') |
0.1 |
| 1-Wasserstein distance with option positions | sum_{k<K} \|F_p(k) - F_q(k)\| (x_{k+1} - x_k); normalised by x_K - x_1 |
[0, x_K - x_1], lower | Vaserstein (1969); Vallender (1974); Villani (2009) | Santurkar et al. (2023); Zhao et al. (2024) | popfidelity.metrics.distance(metric='wasserstein') / metric_distance(metric = 'wasserstein') |
0.1 |
| Total variation distance | 1/2 sum_k \|p_k - q_k\|; overlap and histogram intersection are 1 - TV, Bray-Curtis equals TV |
[0, 1], lower | Scheffé (1947); Levin et al. (2009); Gibbs and Su (2002) | Meister et al. (2025); Hu et al. (2026) | popfidelity.metrics.distance(metric='total_variation') / metric_distance(metric = 'total_variation') |
0.1 |
| Overlap coefficient | sum_k min(p_k, q_k) = 1 - TV(p, q) |
[0, 1], higher | Inman and Bradley (1989); Swain and Ballard (1991); Bray and Curtis (1957) | Pfänder (2026) | 1 - popfidelity.metrics.distance(metric='total_variation') / 1 - metric_distance(metric = 'total_variation') |
0.1 |
| Kullback-Leibler divergence | sum_k p_k log(p_k / q_k) after adding epsilon to every share; direction and log base explicit |
[0, inf), lower | Kullback and Leibler (1951); Ali and Silvey (1966) | Dominguez-Olmedo et al. (2024); Sun et al. (2024); Libovick\'y (2026) | popfidelity.metrics.distance(metric='kl') / metric_distance(metric = 'kl') |
0.1 |
| Jensen-Shannon divergence | 1/2 KL(p \|\| m) + 1/2 KL(q \|\| m), m = (p + q) / 2 |
[0, 1] in bits, lower | Lin (1991) | Cao et al. (2025); DataViking-Tech (2026) | popfidelity.metrics.distance(metric='js_divergence') / metric_distance(metric = 'js_divergence') |
0.1 |
| Jensen-Shannon distance | sqrt(JSD(p, q)), a metric |
[0, 1] in bits, lower | Lin (1991); Endres and Schindelin (2003) | Durmus et al. (2024); Zhao et al. (2024) | popfidelity.metrics.distance(metric='js_distance') / metric_distance(metric = 'js_distance') |
0.1 |
| Hellinger distance | sqrt(1 - sum_k sqrt(p_k q_k)) |
[0, 1], lower | Hellinger (1909) | Shankar et al. (2026) | popfidelity.metrics.distance(metric='hellinger') / metric_distance(metric = 'hellinger') |
0.1 |
| Bhattacharyya coefficient and distance | BC = sum_k sqrt(p_k q_k); D_B = -ln BC |
[0, 1] and [0, inf), coefficient higher, distance lower | Bhattacharyya (1946); Kailath (1967) | – | popfidelity.metrics.distance(metric='bhattacharyya_distance') / metric_distance(metric = 'bhattacharyya_distance') |
0.1 |
| Kolmogorov-Smirnov distance | max_k \|F_p(k) - F_q(k)\|; similarity 1 - D |
[0, 1], lower | Kolmogorov (1933); Smirnov (1948); Conover (1972) | Maier et al. (2025); Hong et al. (2025) | popfidelity.metrics.distance(metric='kolmogorov_smirnov') / metric_distance(metric = 'kolmogorov_smirnov') |
0.1 |
| Discrete Cramer-von Mises distance | sum_k (F_p(k) - F_q(k))^2 (p_k + q_k) / 2 |
[0, 1], lower | Cramér (1928); von Mises (1928); Anderson (1962); Choulakian et al. (1994) | – | popfidelity.metrics.distance(metric='cramer_von_mises') / metric_distance(metric = 'cramer_von_mises') |
0.1 |
| Anderson-Darling statistic for discrete scales | CDF gaps weighted by 1 / (H (1 - H)) of the pooled CDF, tie-aware |
[0, inf), lower | Anderson and Darling (1952); Scholz and Stephens (1987) | – | – | planned 0.2 |
| Chi-square divergence | sum_k (p_k - q_k)^2 / q_k |
[0, inf), lower | Pearson (1900); Ali and Silvey (1966) | Huang et al. (2025) | popfidelity.metrics.distance(metric='chi_square_divergence') / metric_distance(metric = 'chi_square_divergence') |
0.1 |
| Pearson chi-square test and Cramer's V | X^2 = sum (O - E)^2 / E over cells by options; V = sqrt(X^2 / (N min(r - 1, c - 1))) |
V in [0, 1], diagnostic | Pearson (1900); Cramér (1946) | Sun et al. (2024); Argyle et al. (2023) | popfidelity.dispersion.association / dispersion_association |
0.1 |
| Maximum mean discrepancy | sqrt((p - q)' K (p - q)) with a Gaussian kernel over option indices |
[0, inf), lower | Gretton et al. (2012) | Hu et al. (2025); Jia et al. (2026) | popfidelity.metrics.distance(metric='mmd') / metric_distance(metric = 'mmd') |
0.1 |
| Energy distance | 2 sum_{k<K} (F_p(k) - F_q(k))^2 (x_{k+1} - x_k) |
[0, inf), lower | Székely and Rizzo (2013) | – | popfidelity.metrics.distance(metric='energy_distance') / metric_distance(metric = 'energy_distance') |
0.1 |
| Sliced Wasserstein distance of joint answer profiles | mean over random projections of the 1-D Wasserstein distance |
[0, inf), lower | Bonneel et al. (2015) | Hu et al. (2025) | – | planned 0.2 |
| Absolute difference of mean answers | \|sum_k x_k q_k - sum_k x_k p_k\| |
[0, x_K - x_1], lower | – | Bisbee et al. (2024) | popfidelity.metrics.distance(metric='mean_difference') / metric_distance(metric = 'mean_difference') |
0.1 |
| Squared difference of rescaled mean answers | ((mu_q - mu_p) / (x_K - x_1))^2 |
[0, 1], lower | – | Libovick\'y (2026) | popfidelity.metrics.distance(metric='rescaled_squared_mean_difference') / metric_distance(metric = 'rescaled_squared_mean_difference') |
0.1 |
| Shannon entropy | -sum_k p_k log p_k |
[0, log K], diagnostic | Shannon (1948) | Dominguez-Olmedo et al. (2024) | popfidelity.metrics.entropy / metric_entropy |
0.1 |
| Normalised entropy | H(p) / ln K |
[0, 1], diagnostic | Shannon (1948); Pielou (1966) | Dominguez-Olmedo et al. (2024) | popfidelity.metrics.normalised_entropy / metric_normalised_entropy |
0.1 |
| Effective number of options | exp(H(p)) with natural logarithms |
[1, K], diagnostic | Hill (1973) | Ahn et al. (2026) | popfidelity.metrics.effective_categories / metric_effective_categories |
0.1 |
| Collision probability | sum_k p_k^2 |
[1/K, 1], diagnostic | Simpson (1949) | Skobelev et al. (2026) | popfidelity.metrics.collision / metric_collision |
0.1 |
| Modal answer agreement | 1[argmax p = argmax q], the first maximum on ties |
{0, 1}, higher | – | Feng et al. (2024); AlKhamissi et al. (2024) | popfidelity.metrics.distance(metric='top_agreement') / metric_distance(metric = 'top_agreement') |
0.1 |
Alignment scores from the literature¶
| Measure | Formula | Range | Origin | Applied to LLMs | Python / R | Status |
|---|---|---|---|---|---|---|
| OpinionQA representativeness | mean_q [1 - W1(D_m(q), D_g(q)) / (N - 1)] |
[0, 1], higher | Santurkar et al. (2023) | Santurkar et al. (2023); Zhao et al. (2024) | popfidelity.alignment.opinionqa_representativeness / align_opinionqa_representativeness |
0.1 |
| OpinionQA steerability | mean_q max_c [1 - W1(D_m(q; c), D_g(q)) / (N - 1)] over steering prompts c |
[0, 1], higher | Santurkar et al. (2023) | Santurkar et al. (2023) | popfidelity.alignment.opinionqa_steerability / align_opinionqa_steerability |
0.1 |
| OpinionQA consistency | share of topics whose best-aligned group is the overall best group |
[0, 1], higher | Santurkar et al. (2023) | Santurkar et al. (2023) | popfidelity.alignment.opinionqa_consistency / align_opinionqa_consistency |
0.1 |
| GlobalOpinionQA similarity | mean_q [1 - JS distance(P_m(q), P_c(q))] |
[0, 1], higher | Durmus et al. (2024); Lin (1991) | Durmus et al. (2024); Zhao et al. (2024) | popfidelity.alignment.globalopinionqa_similarity / align_globalopinionqa_similarity |
0.1 |
| Distributional alignment | mean over groups and questions of TV(y, y_hat), with uniform and majority baselines |
[0, 1], lower | Meister et al. (2025) | Meister et al. (2025) | popfidelity.alignment.meister_alignment / align_meister |
0.1 |
| SimBench score | 100 (1 - TVD(P, Q) / mean TVD(P, uniform)) |
(-inf, 100], higher | Hu et al. (2026) | Hu et al. (2026) | popfidelity.alignment.simbench_score / align_simbench |
0.1 |
| SubPOP Wasserstein distance with bounds | sum_i \|F_H(i) - F_theta(i)\|; uniform upper bound; human-bootstrap noise floor |
[0, K - 1], lower | Suh et al. (2025) | Suh et al. (2025) | popfidelity.alignment.subpop_distance / align_subpop_distance |
0.1 |
| Human resampling noise floor | mean EMD between a distribution and its multinomial resamples |
[0, K - 1], reference | Suh et al. (2025); Efron (1979) | Suh et al. (2025); Williams et al. (2026) | popfidelity.alignment.subpop_noise_floor / align_subpop_noise_floor |
0.1 |
| Wasserstein threshold curve | share of questions with W1 on [0, 1]-rescaled answers at most t, t = 0, 0.05, ..., 1 |
[0, 1], higher | Zhao et al. (2024) | Zhao et al. (2024) | popfidelity.alignment.threshold_curve / align_threshold_curve |
0.1 |
| Machine Bias quality bands | nEMD right-closed at 0.05, 0.10, 0.15 and 0.30 |
very good to very bad, lower band | Boelaert et al. (2025) | Boelaert et al. (2025) | popfidelity.structure.quality_table / structure_quality_table |
0.1 |
| Machine Bias centre and social models | log1p(nEMD_i) ~ log1p(nEMD(s_i, model centre)) against log1p(nEMD_i) ~ demographics, by adjusted R^2, BIC and F tests |
adjusted R^2, diagnostic | Boelaert et al. (2025); Wherry (1931); Schwarz (1978) | Boelaert et al. (2025) | – | planned 0.2 |
| A-bias | \|P(answer is label A) - 1/k\| |
[0, 1 - 1/k], lower | Dominguez-Olmedo et al. (2024) | Dominguez-Olmedo et al. (2024) | popfidelity.alignment.a_bias / align_a_bias |
0.1 |
| Order-randomised answer distribution | answer distribution averaged over permutations of the option-to-label assignment |
distribution, diagnostic | Dominguez-Olmedo et al. (2024) | Dominguez-Olmedo et al. (2024) | – | planned 0.2 |
| SynthBench parity score | mean of 1 - JSD, (1 + Kendall tau_b) / 2, conditioning fidelity, 1 - CV and refusal parity |
[0, 1], higher | DataViking-Tech (2026) | DataViking-Tech (2026) | – | planned 0.2 |
Dispersion and homogenisation¶
| Measure | Formula | Range | Origin | Applied to LLMs | Python / R | Status |
|---|---|---|---|---|---|---|
| Standard-deviation ratio | SD of model answers / SD of survey answers, pooled or within cells |
[0, inf), closer to 1 | – | Bisbee et al. (2024); Peng et al. (2025); Lukauskas and Sarkauskait.e (2026) | popfidelity.dispersion.sd_ratio / dispersion_sd_ratio |
0.1 |
| Normalised response variance | Var(answer) / (x_K - x_1)^2 |
[0, 1/4], diagnostic | Williams et al. (2026) | Williams et al. (2026) | popfidelity.dispersion.normalised_variance / dispersion_normalised_variance |
0.1 |
| Gap-inflation factor | (max_g - min_g share at or above an option) for the model over the survey |
[0, inf), closer to 1 | Chen et al. (2026) | Chen et al. (2026) | popfidelity.dispersion.gap_inflation / dispersion_gap_inflation |
0.1 |
| Stereotyping index | eta^2_model - eta^2_survey and V_model - V_survey |
[-1, 1], closer to 0 | Pearson (1900); Cramér (1946) | Chen et al. (2026) | popfidelity.dispersion.stereotyping / dispersion_stereotyping |
0.1 |
| Entropy, effective-category and collision ratios | H(q) / H(p); exp H(q) / exp H(p); sum q^2 / sum p^2 |
[0, inf), closer to 1 | Shannon (1948); Hill (1973); Simpson (1949) | Ahn et al. (2026); Skobelev et al. (2026) | popfidelity.dispersion.diversity_ratios / dispersion_diversity_ratios |
0.1 |
| Support collapse | options with zero model share and the survey share on them |
options and [0, 1], lower | Doudkin (2026) | Doudkin (2026) | popfidelity.dispersion.support_collapse / dispersion_support_collapse |
0.1 |
| Modal concentration and collapse rule | max_k q_k; collapsed when it exceeds the survey's by 0.15 |
[0, 1], closer to the survey | Ozkan (2026) | Ozkan (2026) | popfidelity.dispersion.modal_collapse / dispersion_modal_collapse |
0.1 |
| Variation ratio | 1 - max_k p_k |
[0, 1 - 1/K], diagnostic | Freeman (1965) | Li et al. (2025) | 1 - popfidelity.metrics.modal_concentration / 1 - metric_modal_concentration |
0.1 |
| Aggregation consistency | distance between a directly simulated group and the weighted aggregate of its simulated subgroups |
[0, 1], lower | Li et al. (2025) | Li et al. (2025) | – | planned 0.2 |
| Describe-sample gap | TV(sampled, human) - TV(verbalised, human) |
[-1, 1], closer to 0 | Jang et al. (2026) | Jang et al. (2026) | – | planned 0.2 |
| Kernel-of-truth exaggeration | E_B(a \| X+) = (1 + gamma) E(a \| X+) - gamma E(a \| X-) |
(-inf, inf), closer to 0 | Bordalo et al. (2016) | Jeoung et al. (2025) | – | planned 0.2 |
Structure across subgroups and questions¶
| Measure | Formula | Range | Origin | Applied to LLMs | Python / R | Status |
|---|---|---|---|---|---|---|
| Mantel test | correlation of the two condensed vectors with a permutation null over cells |
[-1, 1] and a p-value, higher | Mantel (1967); Smouse et al. (1986); Legendre and Fortin (2010) | da Silva et al. (2026); Shi and Haupt (2026) | popfidelity.structure.mantel / structure_mantel |
0.1 |
| PFS structure with a permutation p-value | Mantel test of the survey and model pairwise nEMD matrices |
[-1, 1] and a p-value, higher | Mantel (1967); da Silva et al. (2026) | da Silva et al. (2026) | popfidelity.structure.structure_test / structure_test |
0.1 |
| Spearman, Pearson and Kendall tau-b | Pearson on values, Pearson on average ranks, Knight's tau-b with ties |
[-1, 1], higher | Pearson (1895); Spearman (1904); Kendall (1938); Kendall (1945); Knight (1966) | Masoud et al. (2025); da Silva et al. (2026) | popfidelity.structure.correlation / structure_correlation |
0.1 |
| Classical multidimensional scaling | B = -1/2 J D^2 J; coordinates from the leading eigenvectors |
coordinates, diagnostic | Torgerson (1952); Gower (1966) | Boelaert et al. (2025) | – | planned 0.2 |
| Procrustes analysis and PROTEST | Gower's m^2 after the optimal rotation, scaling and translation, with a permutation test |
[0, 1], lower | Gower (1975); Jackson (1995); Peres-Neto and Jackson (2001) | Lam et al. (2026) | – | planned 0.2 |
| RV coefficient | tr(XX'YY') / sqrt(tr((XX')^2) tr((YY')^2)); modified RV for many variables |
[0, 1], higher | Robert and Escoufier (1976); Smilde et al. (2009) | Shi and Haupt (2026); Lukauskas and Sarkauskait.e (2026) | – | planned 0.2 |
| Tucker's congruence coefficient | sum x y / sqrt(sum x^2 sum y^2) |
[-1, 1], higher | Tucker (1951); Lorenzo-Seva and ten Berge (2006) | Lukauskas and Sarkauskait.e (2026); Lam et al. (2026) | – | planned 0.2 |
| Question-correlation structure | Pearson and RMSE over upper triangles of question-question correlations; permutation floor, split-half ceiling |
[-1, 1] and RMSE, higher | Williams et al. (2026) | Williams et al. (2026); Moon et al. (2024) | – | planned 0.2 |
| Self-correlation distance | \|\|C_model - C_human\|\|_F with C the question-question correlation matrix |
[0, inf), lower | Libovick\'y (2026) | Libovick\'y (2026) | – | planned 0.2 |
| Lin's concordance correlation | 2 s_xy / (s_x^2 + s_y^2 + (mean x - mean y)^2) |
[-1, 1], higher | Lin (1989) | Choi et al. (2026) | – | planned 0.2 |
| Regression coefficient recovery | share of coefficients equal to the survey's, sign flips, average marginal effects |
shares, higher | Clogg et al. (1995); Rubin (1987) | Bisbee et al. (2024); von der Heyde et al. (2026) | – | planned 0.2 |
| Tetrachoric correlation and item-pair Cramer's V | latent-normal correlation of 2x2 tables; V for every item pair, human against model |
[-1, 1], higher | Pearson (1900); Cramér (1946) | Argyle et al. (2023) | – | planned 0.2 |
| Independence-assumption footprint | V_synthetic / V_reference with bias-corrected Cramer's V |
[0, inf), closer to 1 | Cramér (1946) | Bae (2026) | – | planned 0.2 |
| Persona-effect ceiling | marginal R^2 of annotation ~ persona + (1 \| item) |
[0, 1], reference | Nakagawa and Schielzeth (2013) | Hu and Collier (2024) | – | planned 0.2 |
| Intersectional collapse index | how much a pair's tilt follows its dominant single feature rather than both features |
[0, 1], lower | Rennard and Xypolopoulos (2026) | Rennard and Xypolopoulos (2026) | – | planned 0.2 |
| Convex-hull consistency | share of claims whose unconditioned opinion lies in the hull of the group opinions |
[0, 1], higher | Neumann et al. (2026) | Neumann et al. (2026) | – | planned 0.2 |
| Persona-conditioned informativeness | mean over items of max(local Moran's I, 0) |
[0, inf), higher | An et al. (2026) | An et al. (2026) | – | planned 0.2 |
| Cultural-map distance and trajectory | Euclidean distance on the Inglehart-Welzel map; lag and magnitude ratio of change |
[0, inf), lower | Tao et al. (2024) | Tao et al. (2024); Daryani et al. (2026) | – | planned 0.2 |
| Hofstede cultural alignment test | Kendall tau between VSM13 dimension rankings of countries, model against survey |
[-1, 1], higher | Kendall (1938) | Masoud et al. (2025) | – | planned 0.2 |
| WEIRDness gradient | correlation over countries of model-country similarity with cultural distance from the United States |
[-1, 1], closer to 0 | Atari et al. (2023) | Atari et al. (2023) | – | planned 0.2 |
| Heterogeneity-collapse battery | variance ratio, entropy, Mantel and RV between correlation matrices, PCA variance explained |
varies, closer to humans | Mantel (1967); Robert and Escoufier (1976) | Shi and Haupt (2026); Lee and Wang (2026) | – | planned 0.2 |
| Distance correlation | dCov(X, Y) / sqrt(dVar(X) dVar(Y)) from double-centred distance matrices |
[0, 1], closer to humans | Székely et al. (2007) | Ahnert et al. (2026) | – | planned 0.2 |
| Construct coherence | construct JSD over joint answers; mean absolute gaps in Cronbach's alpha and AVE |
varies, closer to humans | Cronbach (1951); Fornell and Larcker (1981); Lin (1991) | Lukauskas and Sarkauskait.e (2026) | – | planned 0.3 |
| Item-parameter recovery | correlation of 2PL discrimination and difficulty fitted on model and human answers |
[-1, 1], higher | – | Liu et al. (2025) | – | planned 0.3 |
Reliability, steerability and elicitation artefacts¶
| Measure | Formula | Range | Origin | Applied to LLMs | Python / R | Status |
|---|---|---|---|---|---|---|
| Intraclass correlation | McGraw and Wong one-way, consistency and agreement forms, single or average |
(-inf, 1], higher | Shrout and Fleiss (1979); McGraw and Wong (1996) | da Silva et al. (2026); Lukauskas and Sarkauskait.e (2026) | popfidelity.reliability.icc / reliability_icc |
0.1 |
| Krippendorff's alpha | 1 - (n - 1) sum o_ck delta_ck / sum n_c n_k delta_ck |
(-inf, 1], higher | Krippendorff (1970); Hayes and Krippendorff (2007) | – | popfidelity.reliability.krippendorff_alpha / reliability_krippendorff_alpha |
0.1 |
| Fleiss' kappa | (P_bar - P_e) / (1 - P_e) |
(-inf, 1], higher | Fleiss (1971) | Ahnert et al. (2026) | popfidelity.reliability.fleiss_kappa / reliability_fleiss_kappa |
0.1 |
| Cronbach's alpha | k / (k - 1) (1 - sum Var(item) / Var(total)) |
(-inf, 1], higher | Cronbach (1951) | Serapio-García et al. (2025); Lukauskas and Sarkauskait.e (2026) | popfidelity.reliability.cronbach_alpha / reliability_cronbach_alpha |
0.1 |
| Spearman-Brown prophecy | k r / (1 + (k - 1) r) |
(-inf, 1], reference | Spearman (1910); Brown (1910) | Williams et al. (2026) | popfidelity.reliability.spearman_brown / reliability_spearman_brown |
0.1 |
| Pooled t interval over replicate runs | t(1 - a/2, sum df) * sqrt(sum df_i s_i^2 / sum df_i) / sqrt(n_i) |
half-widths, diagnostic | Satterthwaite (1946); Welch (1947) | da Silva et al. (2026) | popfidelity.across_replicates / pfs_across_replicates |
0.1 |
| Noise-to-signal ratio | (sd_seed + sd_prompt) / sd_signal; above 1 the position is noise |
[0, inf), lower | Nguyen and Ahmad (2026) | Nguyen and Ahmad (2026) | popfidelity.reliability.noise_to_signal / reliability_noise_to_signal |
0.1 |
| Psychometric validity battery | alpha, omega, convergent and discriminant correlations, CFA fit, psychometric similarity score |
varies, closer to humans | Cronbach (1951) | Serapio-García et al. (2025); Petrov et al. (2024); Lukauskas and Sarkauskait.e (2026) | – | planned 0.3 |
Elicitation artefacts¶
| Measure | Formula | Range | Origin | Applied to LLMs | Python / R | Status |
|---|---|---|---|---|---|---|
| First-token versus text mismatch rate | share of items whose first-token argmax differs from the parsed text answer |
[0, 1], lower | Wang et al. (2024) | Wang et al. (2024) | popfidelity.reliability.mismatch_rate / reliability_mismatch_rate |
0.1 |
| Option-ID selection bias | standard deviation over option IDs of recall; PriDe debiasing prior |
[0, 1], lower | Zheng et al. (2024) | Zheng et al. (2024) | – | planned 0.2 |
| Response-bias shift | mean change of the bias-relevant answers between question versions; perturbations should not move it |
[-1, 1], matches humans | Tjuatja et al. (2024) | Tjuatja et al. (2024) | – | planned 0.2 |
| Prompt steerability indices | [W(p, p_hat) - W(p_k, p_hat)] / W(p_hat+, p_hat-) |
[-1, 1], higher | Miehling et al. (2025) | Miehling et al. (2025) | – | planned 0.2 |
| Sociodemographic prompting sensitivity and placebo test | share of labels changed by a profile; variance under cultural cues against placebo cues |
[0, 1], diagnostic | Beck et al. (2024) | Beck et al. (2024); Mukherjee et al. (2024) | – | planned 0.2 |
| Format instability and steerability ratio | weighted mean difference across formats; LLM-human over human-human distance |
[0, inf), lower | Khan et al. (2025) | Khan et al. (2025) | – | planned 0.2 |
| Value consistency and stance reliability | multi-distribution JSD to a centroid; bootstrap reliability of stances; rank-order stability |
[0, 1], lower | Lin (1991) | Moore et al. (2024); Ceron et al. (2024); Kovač et al. (2024); Röttger et al. (2024) | – | planned 0.2 |
Individual-level prediction¶
| Measure | Formula | Range | Origin | Applied to LLMs | Python / R | Status |
|---|---|---|---|---|---|---|
| Hard and soft accuracy | mean 1(y_hat = y); mean (1 - \|y_hat - y\| / (K - 1)) |
[0, 1], higher | – | AlKhamissi et al. (2024); Peng et al. (2025) | – | planned 0.3 |
| Accuracy against a test-retest ceiling | agent accuracy / the person's own retest accuracy |
[0, 1+], higher | Spearman (1904) | Park et al. (2024); Toubia et al. (2025) | – | planned 0.3 |
| Performance parity gap | best-group minus worst-group performance |
[0, 1], lower | Dwork et al. (2012); Hardt et al. (2016) | Park et al. (2024) | – | planned 0.3 |
| Log loss, Brier and ranked probability scores | -log q_y; sum (q_k - 1[y = k])^2; sum_k (Q_k - 1[y <= k])^2 / (K - 1) |
[0, inf), lower | Good (1952); Brier (1950); Epstein (1969); Murphy (1971); Gneiting and Raftery (2007) | Chen et al. (2026); Binz et al. (2025) | – | planned 0.3 |
| Expected calibration error | sum_b \|B_b\| / N \|acc(B_b) - conf(B_b)\| |
[0, 1], lower | Naeini et al. (2015); Guo et al. (2017) | – | – | planned 0.3 |
| ROC-AUC and pseudo-R^2 | area under the ROC curve; 1 - NLL_model / NLL_random |
[0, 1], higher | – | Kim and Lee (2023); Binz et al. (2025) | – | planned 0.3 |
| Match rate, F1 and election agreement | share of matching choices; per-party F1; share of states called correctly |
[0, 1], higher | – | von der Heyde et al. (2026); Zhang et al. (2024) | – | planned 0.3 |
Effect-level and causal validity¶
| Measure | Formula | Range | Origin | Applied to LLMs | Python / R | Status |
|---|---|---|---|---|---|---|
| Effect-size correlation with disattenuation | r / sqrt(rel_model rel_human); sign agreement |
[-1, 1], higher | Spearman (1904) | Ashokkumar et al. (2026); Pfänder (2026) | – | planned 0.3 |
| Type S and type M errors | P(wrong sign \| significant); E\|estimate\| / \|true effect\| |
[0, 1] and [0, inf), lower | Gelman and Carlin (2014); Pesaran and Timmermann (1992) | Cui et al. (2025) | – | planned 0.3 |
| Replication rate, coverage, inflation and false positives | same sign and significant; interval coverage; Fisher-z ratio; significance on human nulls |
[0, 1], higher | Schuirmann (1987) | Cui et al. (2025); Aher et al. (2023); Horton et al. (2023) | – | planned 0.3 |
| AMCE distance and conjoint battery | \|\|AMCE_model - AMCE_human\|\|_2 with sign agreement and coverage |
[0, inf), lower | – | Takemoto (2024); Ahmad and Takemoto (2025); Hung et al. (2026) | – | planned 0.3 |
| Treatment-effect recovery and negative controls | tau_model / tau_human; TV of a negative-control outcome across arms |
[0, inf), closer to 1 | – | Persson et al. (2026); Lin et al. (2026); Gui and Toubia (2023); Li and Ji (2026) | – | planned 0.3 |
| Willingness-to-pay recovery | model willingness to pay over human willingness to pay |
ratio, closer to 1 | – | Brand et al. (2023) | – | reference only |
Inference and uncertainty¶
| Measure | Formula | Range | Origin | Applied to LLMs | Python / R | Status |
|---|---|---|---|---|---|---|
| Bootstrap percentile intervals | percentile interval of a score over resamples of cells with replacement |
intervals, narrower | Efron (1979); Efron and Tibshirani (1993) | Pfänder (2026); Williams et al. (2026) | popfidelity.uncertainty.bootstrap / uncertainty_bootstrap |
0.1 |
| Permutation null of cell correspondence | scores after shuffling which cell each model distribution belongs to |
p-value, lower p | Pitman (1937) | Williams et al. (2026) | popfidelity.uncertainty.permutation_null / uncertainty_permutation_null |
0.1 |
| Null predictions | survey centre, leave-one-out, uniform, majority and permuted cells; above_null = (score - null) / (1 - null) |
scores, reference | da Silva et al. (2026); Meister et al. (2025) | da Silva et al. (2026); Boelaert et al. (2025); Dominguez-Olmedo et al. (2024) | popfidelity.baselines.null_scores / baseline_null_scores |
0.1 |
| Holm, TOST, Wilcoxon and cluster-robust tests | Holm step-down; two one-sided tests; signed-rank test; sandwich standard errors |
p-values, diagnostic | Holm (1979); Schuirmann (1987); Wilcoxon (1945); Liang and Zeger (1986) | da Silva et al. (2026) | – | planned 0.2 |
| Prediction-powered inference | theta = mean f(X_unlabelled) - mean(f(X) - Y); PPI++ tunes a power parameter; DSL |
estimates, narrower | Angelopoulos et al. (2023); Angelopoulos et al. (2023); Egami et al. (2023) | Krsteski et al. (2026); Broska et al. (2025) | – | planned 0.3 |
| Effective human sample size | ESS gain (Var_human / Var_method - 1); largest k with miscoverage at most gamma alpha |
[0, inf), higher | Kish (1965) | Krsteski et al. (2026); Huang et al. (2025); Gao et al. (2026) | – | planned 0.3 |
| Rectification difficulty and residualised correlation | Var(Y)(1 - rho^2); rho_T = Corr(Y - E[Y\|T], Y_hat - E[Y_hat\|T]) |
[0, inf), lower | – | Ye and Yoganarasimhan (2026); Wang et al. (2026) | – | planned 0.3 |
| Sim-to-real quantile curve | quantile function of a discrepancy proxy, summarised by CVaR |
curve, lower | – | Iyengar et al. (2025) | – | planned 0.3 |
Free-text measures (documented only)¶
| Measure | Formula | Range | Origin | Applied to LLMs | Python / R | Status |
|---|---|---|---|---|---|---|
| Flattening and coverage of free text | n-gram rarity, embedding dispersion, distinct options and Vendi score per identity |
varies, closer to humans | Friedman and Dieng (2023) | Wang et al. (2025) | – | reference only |
| Semantic diversity and hivemind homogeneity | mean pairwise cosine distance of response embeddings |
[0, 2], closer to humans | – | Liu et al. (2024) | – | reference only |
| Persona consistency and PersonaScore | prompt-to-line, line-to-line and Q&A consistency; rubric scores of evaluator models |
[0, 1] or 1 to 5, higher | – | Abdulhai et al. (2025); Samuel et al. (2025) | – | reference only |
| Computational Turing test | held-out accuracy of a classifier telling human from model replies |
[0.5, 1], closer to 0.5 | – | Pagan et al. (2025); Mei et al. (2024) | – | reference only |
| Caricature: individuation and exaggeration | classifier separability of persona outputs and normalised similarity to persona-topic axes |
[0, 1], lower | – | Cheng et al. (2023); Cheng et al. (2023) | – | reference only |
| Overton pluralism score | share of human viewpoint clusters a response represents |
[0, 1], higher | – | Poole-Dayan et al. (2026) | – | reference only |