Skip to content

Measures

The registry lists 126 measures. Each names the work it comes from and the work that applied it to language-model survey responses; the status says whether a measure ships in 0.1, is planned, or is documented only. Python, R and the command line (popfidelity registry) read the same registry.

Population Fidelity Score

Measure Formula Range Origin Applied to LLMs Python / R Status
Population Fidelity Score (S_acc * S_adapt * S_struct)^(1/3); 0 if any component is 0 [0, 1], higher da Silva et al. (2026); Klugman et al. (2011) da Silva et al. (2026) popfidelity.score / pfs_score 0.1
Accuracy component S_acc = 1 - mean_i nEMD(s_i, m_i) [0, 1], higher Boelaert et al. (2025); da Silva et al. (2026) da Silva et al. (2026) popfidelity.score_groups / pfs_score_groups 0.1
Adaptability component A = median_{i<j} nEMD(m_i, m_j) / median_{i<j} nEMD(s_i, s_j); S_adapt = min(A, 1/A) [0, 1], higher Boelaert et al. (2025); da Silva et al. (2026) da Silva et al. (2026) popfidelity.score_groups / pfs_score_groups 0.1
Structure component S_struct = max(0, Spearman(d_s, d_m)) over the condensed pairwise nEMD vectors; 0 if the model is flat [0, 1], higher Spearman (1904); Mantel (1967); da Silva et al. (2026) da Silva et al. (2026) popfidelity.score_groups / pfs_score_groups 0.1
Centre alignment S_center = 1 - nEMD(mean_i s_i, mean_i m_i); reported beside PFS, not inside it [0, 1], higher Boelaert et al. (2025) Boelaert et al. (2025); da Silva et al. (2026) popfidelity.score_groups / pfs_score_groups 0.1
Limiting component argmin of S_acc, S_adapt, S_struct, skipping undefined components; the first wins ties accuracy | adaptability | structure, diagnostic da Silva et al. (2026) da Silva et al. (2026) popfidelity.binding_component / pfs_binding_component 0.1
Signed PFS sign(rho) * (S_acc * S_adapt * \|rho\|)^(1/3) [-1, 1], higher da Silva et al. (2026) da Silva et al. (2026) popfidelity.variant('unclipped_structure') / pfs_variant('unclipped_structure') 0.1
PFS sensitivity variants mean dispersion; one-sided min(A, 1); Pearson or Kendall structure; arithmetic mean; respondent-weighted E and C [0, 1], higher da Silva et al. (2026) da Silva et al. (2026) popfidelity.VARIANTS / pfs_variant 0.1
Paired change in PFS PFS(tuned) - PFS(base) with both scored on their common cells [-1, 1], higher da Silva et al. (2026) da Silva et al. (2026) popfidelity.compare / pfs_compare 0.1
Cross-question PFS geometric mean of per-question PFS; arithmetic means of components [0, 1], higher da Silva et al. (2026); Klugman et al. (2011) da Silva et al. (2026) popfidelity.across_questions / pfs_across_questions 0.1

Distances between answer distributions

Measure Formula Range Origin Applied to LLMs Python / R Status
Normalised earth mover's distance sum_k \|F_p(k) - F_q(k)\| / (K - 1) [0, 1], lower Monge (1781); Kantorovitch (1958); Vaserstein (1969); Vallender (1974); Rubner et al. (2000) Santurkar et al. (2023); Boelaert et al. (2025); da Silva et al. (2026) popfidelity.metrics.distance(metric='nemd') / metric_distance(metric = 'nemd') 0.1
Earth mover's distance sum_k \|F_p(k) - F_q(k)\| with unit spacing [0, K - 1], lower Rubner et al. (2000); Villani (2009) Suh et al. (2025); Cao et al. (2025) popfidelity.metrics.distance(metric='emd') / metric_distance(metric = 'emd') 0.1
1-Wasserstein distance with option positions sum_{k<K} \|F_p(k) - F_q(k)\| (x_{k+1} - x_k); normalised by x_K - x_1 [0, x_K - x_1], lower Vaserstein (1969); Vallender (1974); Villani (2009) Santurkar et al. (2023); Zhao et al. (2024) popfidelity.metrics.distance(metric='wasserstein') / metric_distance(metric = 'wasserstein') 0.1
Total variation distance 1/2 sum_k \|p_k - q_k\|; overlap and histogram intersection are 1 - TV, Bray-Curtis equals TV [0, 1], lower Scheffé (1947); Levin et al. (2009); Gibbs and Su (2002) Meister et al. (2025); Hu et al. (2026) popfidelity.metrics.distance(metric='total_variation') / metric_distance(metric = 'total_variation') 0.1
Overlap coefficient sum_k min(p_k, q_k) = 1 - TV(p, q) [0, 1], higher Inman and Bradley (1989); Swain and Ballard (1991); Bray and Curtis (1957) Pfänder (2026) 1 - popfidelity.metrics.distance(metric='total_variation') / 1 - metric_distance(metric = 'total_variation') 0.1
Kullback-Leibler divergence sum_k p_k log(p_k / q_k) after adding epsilon to every share; direction and log base explicit [0, inf), lower Kullback and Leibler (1951); Ali and Silvey (1966) Dominguez-Olmedo et al. (2024); Sun et al. (2024); Libovick\'y (2026) popfidelity.metrics.distance(metric='kl') / metric_distance(metric = 'kl') 0.1
Jensen-Shannon divergence 1/2 KL(p \|\| m) + 1/2 KL(q \|\| m), m = (p + q) / 2 [0, 1] in bits, lower Lin (1991) Cao et al. (2025); DataViking-Tech (2026) popfidelity.metrics.distance(metric='js_divergence') / metric_distance(metric = 'js_divergence') 0.1
Jensen-Shannon distance sqrt(JSD(p, q)), a metric [0, 1] in bits, lower Lin (1991); Endres and Schindelin (2003) Durmus et al. (2024); Zhao et al. (2024) popfidelity.metrics.distance(metric='js_distance') / metric_distance(metric = 'js_distance') 0.1
Hellinger distance sqrt(1 - sum_k sqrt(p_k q_k)) [0, 1], lower Hellinger (1909) Shankar et al. (2026) popfidelity.metrics.distance(metric='hellinger') / metric_distance(metric = 'hellinger') 0.1
Bhattacharyya coefficient and distance BC = sum_k sqrt(p_k q_k); D_B = -ln BC [0, 1] and [0, inf), coefficient higher, distance lower Bhattacharyya (1946); Kailath (1967) – popfidelity.metrics.distance(metric='bhattacharyya_distance') / metric_distance(metric = 'bhattacharyya_distance') 0.1
Kolmogorov-Smirnov distance max_k \|F_p(k) - F_q(k)\|; similarity 1 - D [0, 1], lower Kolmogorov (1933); Smirnov (1948); Conover (1972) Maier et al. (2025); Hong et al. (2025) popfidelity.metrics.distance(metric='kolmogorov_smirnov') / metric_distance(metric = 'kolmogorov_smirnov') 0.1
Discrete Cramer-von Mises distance sum_k (F_p(k) - F_q(k))^2 (p_k + q_k) / 2 [0, 1], lower Cramér (1928); von Mises (1928); Anderson (1962); Choulakian et al. (1994) – popfidelity.metrics.distance(metric='cramer_von_mises') / metric_distance(metric = 'cramer_von_mises') 0.1
Anderson-Darling statistic for discrete scales CDF gaps weighted by 1 / (H (1 - H)) of the pooled CDF, tie-aware [0, inf), lower Anderson and Darling (1952); Scholz and Stephens (1987) – – planned 0.2
Chi-square divergence sum_k (p_k - q_k)^2 / q_k [0, inf), lower Pearson (1900); Ali and Silvey (1966) Huang et al. (2025) popfidelity.metrics.distance(metric='chi_square_divergence') / metric_distance(metric = 'chi_square_divergence') 0.1
Pearson chi-square test and Cramer's V X^2 = sum (O - E)^2 / E over cells by options; V = sqrt(X^2 / (N min(r - 1, c - 1))) V in [0, 1], diagnostic Pearson (1900); Cramér (1946) Sun et al. (2024); Argyle et al. (2023) popfidelity.dispersion.association / dispersion_association 0.1
Maximum mean discrepancy sqrt((p - q)' K (p - q)) with a Gaussian kernel over option indices [0, inf), lower Gretton et al. (2012) Hu et al. (2025); Jia et al. (2026) popfidelity.metrics.distance(metric='mmd') / metric_distance(metric = 'mmd') 0.1
Energy distance 2 sum_{k<K} (F_p(k) - F_q(k))^2 (x_{k+1} - x_k) [0, inf), lower Székely and Rizzo (2013) – popfidelity.metrics.distance(metric='energy_distance') / metric_distance(metric = 'energy_distance') 0.1
Sliced Wasserstein distance of joint answer profiles mean over random projections of the 1-D Wasserstein distance [0, inf), lower Bonneel et al. (2015) Hu et al. (2025) – planned 0.2
Absolute difference of mean answers \|sum_k x_k q_k - sum_k x_k p_k\| [0, x_K - x_1], lower – Bisbee et al. (2024) popfidelity.metrics.distance(metric='mean_difference') / metric_distance(metric = 'mean_difference') 0.1
Squared difference of rescaled mean answers ((mu_q - mu_p) / (x_K - x_1))^2 [0, 1], lower – Libovick\'y (2026) popfidelity.metrics.distance(metric='rescaled_squared_mean_difference') / metric_distance(metric = 'rescaled_squared_mean_difference') 0.1
Shannon entropy -sum_k p_k log p_k [0, log K], diagnostic Shannon (1948) Dominguez-Olmedo et al. (2024) popfidelity.metrics.entropy / metric_entropy 0.1
Normalised entropy H(p) / ln K [0, 1], diagnostic Shannon (1948); Pielou (1966) Dominguez-Olmedo et al. (2024) popfidelity.metrics.normalised_entropy / metric_normalised_entropy 0.1
Effective number of options exp(H(p)) with natural logarithms [1, K], diagnostic Hill (1973) Ahn et al. (2026) popfidelity.metrics.effective_categories / metric_effective_categories 0.1
Collision probability sum_k p_k^2 [1/K, 1], diagnostic Simpson (1949) Skobelev et al. (2026) popfidelity.metrics.collision / metric_collision 0.1
Modal answer agreement 1[argmax p = argmax q], the first maximum on ties {0, 1}, higher – Feng et al. (2024); AlKhamissi et al. (2024) popfidelity.metrics.distance(metric='top_agreement') / metric_distance(metric = 'top_agreement') 0.1

Alignment scores from the literature

Measure Formula Range Origin Applied to LLMs Python / R Status
OpinionQA representativeness mean_q [1 - W1(D_m(q), D_g(q)) / (N - 1)] [0, 1], higher Santurkar et al. (2023) Santurkar et al. (2023); Zhao et al. (2024) popfidelity.alignment.opinionqa_representativeness / align_opinionqa_representativeness 0.1
OpinionQA steerability mean_q max_c [1 - W1(D_m(q; c), D_g(q)) / (N - 1)] over steering prompts c [0, 1], higher Santurkar et al. (2023) Santurkar et al. (2023) popfidelity.alignment.opinionqa_steerability / align_opinionqa_steerability 0.1
OpinionQA consistency share of topics whose best-aligned group is the overall best group [0, 1], higher Santurkar et al. (2023) Santurkar et al. (2023) popfidelity.alignment.opinionqa_consistency / align_opinionqa_consistency 0.1
GlobalOpinionQA similarity mean_q [1 - JS distance(P_m(q), P_c(q))] [0, 1], higher Durmus et al. (2024); Lin (1991) Durmus et al. (2024); Zhao et al. (2024) popfidelity.alignment.globalopinionqa_similarity / align_globalopinionqa_similarity 0.1
Distributional alignment mean over groups and questions of TV(y, y_hat), with uniform and majority baselines [0, 1], lower Meister et al. (2025) Meister et al. (2025) popfidelity.alignment.meister_alignment / align_meister 0.1
SimBench score 100 (1 - TVD(P, Q) / mean TVD(P, uniform)) (-inf, 100], higher Hu et al. (2026) Hu et al. (2026) popfidelity.alignment.simbench_score / align_simbench 0.1
SubPOP Wasserstein distance with bounds sum_i \|F_H(i) - F_theta(i)\|; uniform upper bound; human-bootstrap noise floor [0, K - 1], lower Suh et al. (2025) Suh et al. (2025) popfidelity.alignment.subpop_distance / align_subpop_distance 0.1
Human resampling noise floor mean EMD between a distribution and its multinomial resamples [0, K - 1], reference Suh et al. (2025); Efron (1979) Suh et al. (2025); Williams et al. (2026) popfidelity.alignment.subpop_noise_floor / align_subpop_noise_floor 0.1
Wasserstein threshold curve share of questions with W1 on [0, 1]-rescaled answers at most t, t = 0, 0.05, ..., 1 [0, 1], higher Zhao et al. (2024) Zhao et al. (2024) popfidelity.alignment.threshold_curve / align_threshold_curve 0.1
Machine Bias quality bands nEMD right-closed at 0.05, 0.10, 0.15 and 0.30 very good to very bad, lower band Boelaert et al. (2025) Boelaert et al. (2025) popfidelity.structure.quality_table / structure_quality_table 0.1
Machine Bias centre and social models log1p(nEMD_i) ~ log1p(nEMD(s_i, model centre)) against log1p(nEMD_i) ~ demographics, by adjusted R^2, BIC and F tests adjusted R^2, diagnostic Boelaert et al. (2025); Wherry (1931); Schwarz (1978) Boelaert et al. (2025) – planned 0.2
A-bias \|P(answer is label A) - 1/k\| [0, 1 - 1/k], lower Dominguez-Olmedo et al. (2024) Dominguez-Olmedo et al. (2024) popfidelity.alignment.a_bias / align_a_bias 0.1
Order-randomised answer distribution answer distribution averaged over permutations of the option-to-label assignment distribution, diagnostic Dominguez-Olmedo et al. (2024) Dominguez-Olmedo et al. (2024) – planned 0.2
SynthBench parity score mean of 1 - JSD, (1 + Kendall tau_b) / 2, conditioning fidelity, 1 - CV and refusal parity [0, 1], higher DataViking-Tech (2026) DataViking-Tech (2026) – planned 0.2

Dispersion and homogenisation

Measure Formula Range Origin Applied to LLMs Python / R Status
Standard-deviation ratio SD of model answers / SD of survey answers, pooled or within cells [0, inf), closer to 1 – Bisbee et al. (2024); Peng et al. (2025); Lukauskas and Sarkauskait.e (2026) popfidelity.dispersion.sd_ratio / dispersion_sd_ratio 0.1
Normalised response variance Var(answer) / (x_K - x_1)^2 [0, 1/4], diagnostic Williams et al. (2026) Williams et al. (2026) popfidelity.dispersion.normalised_variance / dispersion_normalised_variance 0.1
Gap-inflation factor (max_g - min_g share at or above an option) for the model over the survey [0, inf), closer to 1 Chen et al. (2026) Chen et al. (2026) popfidelity.dispersion.gap_inflation / dispersion_gap_inflation 0.1
Stereotyping index eta^2_model - eta^2_survey and V_model - V_survey [-1, 1], closer to 0 Pearson (1900); Cramér (1946) Chen et al. (2026) popfidelity.dispersion.stereotyping / dispersion_stereotyping 0.1
Entropy, effective-category and collision ratios H(q) / H(p); exp H(q) / exp H(p); sum q^2 / sum p^2 [0, inf), closer to 1 Shannon (1948); Hill (1973); Simpson (1949) Ahn et al. (2026); Skobelev et al. (2026) popfidelity.dispersion.diversity_ratios / dispersion_diversity_ratios 0.1
Support collapse options with zero model share and the survey share on them options and [0, 1], lower Doudkin (2026) Doudkin (2026) popfidelity.dispersion.support_collapse / dispersion_support_collapse 0.1
Modal concentration and collapse rule max_k q_k; collapsed when it exceeds the survey's by 0.15 [0, 1], closer to the survey Ozkan (2026) Ozkan (2026) popfidelity.dispersion.modal_collapse / dispersion_modal_collapse 0.1
Variation ratio 1 - max_k p_k [0, 1 - 1/K], diagnostic Freeman (1965) Li et al. (2025) 1 - popfidelity.metrics.modal_concentration / 1 - metric_modal_concentration 0.1
Aggregation consistency distance between a directly simulated group and the weighted aggregate of its simulated subgroups [0, 1], lower Li et al. (2025) Li et al. (2025) – planned 0.2
Describe-sample gap TV(sampled, human) - TV(verbalised, human) [-1, 1], closer to 0 Jang et al. (2026) Jang et al. (2026) – planned 0.2
Kernel-of-truth exaggeration E_B(a \| X+) = (1 + gamma) E(a \| X+) - gamma E(a \| X-) (-inf, inf), closer to 0 Bordalo et al. (2016) Jeoung et al. (2025) – planned 0.2

Structure across subgroups and questions

Measure Formula Range Origin Applied to LLMs Python / R Status
Mantel test correlation of the two condensed vectors with a permutation null over cells [-1, 1] and a p-value, higher Mantel (1967); Smouse et al. (1986); Legendre and Fortin (2010) da Silva et al. (2026); Shi and Haupt (2026) popfidelity.structure.mantel / structure_mantel 0.1
PFS structure with a permutation p-value Mantel test of the survey and model pairwise nEMD matrices [-1, 1] and a p-value, higher Mantel (1967); da Silva et al. (2026) da Silva et al. (2026) popfidelity.structure.structure_test / structure_test 0.1
Spearman, Pearson and Kendall tau-b Pearson on values, Pearson on average ranks, Knight's tau-b with ties [-1, 1], higher Pearson (1895); Spearman (1904); Kendall (1938); Kendall (1945); Knight (1966) Masoud et al. (2025); da Silva et al. (2026) popfidelity.structure.correlation / structure_correlation 0.1
Classical multidimensional scaling B = -1/2 J D^2 J; coordinates from the leading eigenvectors coordinates, diagnostic Torgerson (1952); Gower (1966) Boelaert et al. (2025) – planned 0.2
Procrustes analysis and PROTEST Gower's m^2 after the optimal rotation, scaling and translation, with a permutation test [0, 1], lower Gower (1975); Jackson (1995); Peres-Neto and Jackson (2001) Lam et al. (2026) – planned 0.2
RV coefficient tr(XX'YY') / sqrt(tr((XX')^2) tr((YY')^2)); modified RV for many variables [0, 1], higher Robert and Escoufier (1976); Smilde et al. (2009) Shi and Haupt (2026); Lukauskas and Sarkauskait.e (2026) – planned 0.2
Tucker's congruence coefficient sum x y / sqrt(sum x^2 sum y^2) [-1, 1], higher Tucker (1951); Lorenzo-Seva and ten Berge (2006) Lukauskas and Sarkauskait.e (2026); Lam et al. (2026) – planned 0.2
Question-correlation structure Pearson and RMSE over upper triangles of question-question correlations; permutation floor, split-half ceiling [-1, 1] and RMSE, higher Williams et al. (2026) Williams et al. (2026); Moon et al. (2024) – planned 0.2
Self-correlation distance \|\|C_model - C_human\|\|_F with C the question-question correlation matrix [0, inf), lower Libovick\'y (2026) Libovick\'y (2026) – planned 0.2
Lin's concordance correlation 2 s_xy / (s_x^2 + s_y^2 + (mean x - mean y)^2) [-1, 1], higher Lin (1989) Choi et al. (2026) – planned 0.2
Regression coefficient recovery share of coefficients equal to the survey's, sign flips, average marginal effects shares, higher Clogg et al. (1995); Rubin (1987) Bisbee et al. (2024); von der Heyde et al. (2026) – planned 0.2
Tetrachoric correlation and item-pair Cramer's V latent-normal correlation of 2x2 tables; V for every item pair, human against model [-1, 1], higher Pearson (1900); Cramér (1946) Argyle et al. (2023) – planned 0.2
Independence-assumption footprint V_synthetic / V_reference with bias-corrected Cramer's V [0, inf), closer to 1 Cramér (1946) Bae (2026) – planned 0.2
Persona-effect ceiling marginal R^2 of annotation ~ persona + (1 \| item) [0, 1], reference Nakagawa and Schielzeth (2013) Hu and Collier (2024) – planned 0.2
Intersectional collapse index how much a pair's tilt follows its dominant single feature rather than both features [0, 1], lower Rennard and Xypolopoulos (2026) Rennard and Xypolopoulos (2026) – planned 0.2
Convex-hull consistency share of claims whose unconditioned opinion lies in the hull of the group opinions [0, 1], higher Neumann et al. (2026) Neumann et al. (2026) – planned 0.2
Persona-conditioned informativeness mean over items of max(local Moran's I, 0) [0, inf), higher An et al. (2026) An et al. (2026) – planned 0.2
Cultural-map distance and trajectory Euclidean distance on the Inglehart-Welzel map; lag and magnitude ratio of change [0, inf), lower Tao et al. (2024) Tao et al. (2024); Daryani et al. (2026) – planned 0.2
Hofstede cultural alignment test Kendall tau between VSM13 dimension rankings of countries, model against survey [-1, 1], higher Kendall (1938) Masoud et al. (2025) – planned 0.2
WEIRDness gradient correlation over countries of model-country similarity with cultural distance from the United States [-1, 1], closer to 0 Atari et al. (2023) Atari et al. (2023) – planned 0.2
Heterogeneity-collapse battery variance ratio, entropy, Mantel and RV between correlation matrices, PCA variance explained varies, closer to humans Mantel (1967); Robert and Escoufier (1976) Shi and Haupt (2026); Lee and Wang (2026) – planned 0.2
Distance correlation dCov(X, Y) / sqrt(dVar(X) dVar(Y)) from double-centred distance matrices [0, 1], closer to humans Székely et al. (2007) Ahnert et al. (2026) – planned 0.2
Construct coherence construct JSD over joint answers; mean absolute gaps in Cronbach's alpha and AVE varies, closer to humans Cronbach (1951); Fornell and Larcker (1981); Lin (1991) Lukauskas and Sarkauskait.e (2026) – planned 0.3
Item-parameter recovery correlation of 2PL discrimination and difficulty fitted on model and human answers [-1, 1], higher – Liu et al. (2025) – planned 0.3

Reliability, steerability and elicitation artefacts

Measure Formula Range Origin Applied to LLMs Python / R Status
Intraclass correlation McGraw and Wong one-way, consistency and agreement forms, single or average (-inf, 1], higher Shrout and Fleiss (1979); McGraw and Wong (1996) da Silva et al. (2026); Lukauskas and Sarkauskait.e (2026) popfidelity.reliability.icc / reliability_icc 0.1
Krippendorff's alpha 1 - (n - 1) sum o_ck delta_ck / sum n_c n_k delta_ck (-inf, 1], higher Krippendorff (1970); Hayes and Krippendorff (2007) – popfidelity.reliability.krippendorff_alpha / reliability_krippendorff_alpha 0.1
Fleiss' kappa (P_bar - P_e) / (1 - P_e) (-inf, 1], higher Fleiss (1971) Ahnert et al. (2026) popfidelity.reliability.fleiss_kappa / reliability_fleiss_kappa 0.1
Cronbach's alpha k / (k - 1) (1 - sum Var(item) / Var(total)) (-inf, 1], higher Cronbach (1951) Serapio-García et al. (2025); Lukauskas and Sarkauskait.e (2026) popfidelity.reliability.cronbach_alpha / reliability_cronbach_alpha 0.1
Spearman-Brown prophecy k r / (1 + (k - 1) r) (-inf, 1], reference Spearman (1910); Brown (1910) Williams et al. (2026) popfidelity.reliability.spearman_brown / reliability_spearman_brown 0.1
Pooled t interval over replicate runs t(1 - a/2, sum df) * sqrt(sum df_i s_i^2 / sum df_i) / sqrt(n_i) half-widths, diagnostic Satterthwaite (1946); Welch (1947) da Silva et al. (2026) popfidelity.across_replicates / pfs_across_replicates 0.1
Noise-to-signal ratio (sd_seed + sd_prompt) / sd_signal; above 1 the position is noise [0, inf), lower Nguyen and Ahmad (2026) Nguyen and Ahmad (2026) popfidelity.reliability.noise_to_signal / reliability_noise_to_signal 0.1
Psychometric validity battery alpha, omega, convergent and discriminant correlations, CFA fit, psychometric similarity score varies, closer to humans Cronbach (1951) Serapio-García et al. (2025); Petrov et al. (2024); Lukauskas and Sarkauskait.e (2026) – planned 0.3

Elicitation artefacts

Measure Formula Range Origin Applied to LLMs Python / R Status
First-token versus text mismatch rate share of items whose first-token argmax differs from the parsed text answer [0, 1], lower Wang et al. (2024) Wang et al. (2024) popfidelity.reliability.mismatch_rate / reliability_mismatch_rate 0.1
Option-ID selection bias standard deviation over option IDs of recall; PriDe debiasing prior [0, 1], lower Zheng et al. (2024) Zheng et al. (2024) – planned 0.2
Response-bias shift mean change of the bias-relevant answers between question versions; perturbations should not move it [-1, 1], matches humans Tjuatja et al. (2024) Tjuatja et al. (2024) – planned 0.2
Prompt steerability indices [W(p, p_hat) - W(p_k, p_hat)] / W(p_hat+, p_hat-) [-1, 1], higher Miehling et al. (2025) Miehling et al. (2025) – planned 0.2
Sociodemographic prompting sensitivity and placebo test share of labels changed by a profile; variance under cultural cues against placebo cues [0, 1], diagnostic Beck et al. (2024) Beck et al. (2024); Mukherjee et al. (2024) – planned 0.2
Format instability and steerability ratio weighted mean difference across formats; LLM-human over human-human distance [0, inf), lower Khan et al. (2025) Khan et al. (2025) – planned 0.2
Value consistency and stance reliability multi-distribution JSD to a centroid; bootstrap reliability of stances; rank-order stability [0, 1], lower Lin (1991) Moore et al. (2024); Ceron et al. (2024); Kovač et al. (2024); Röttger et al. (2024) – planned 0.2

Individual-level prediction

Measure Formula Range Origin Applied to LLMs Python / R Status
Hard and soft accuracy mean 1(y_hat = y); mean (1 - \|y_hat - y\| / (K - 1)) [0, 1], higher – AlKhamissi et al. (2024); Peng et al. (2025) – planned 0.3
Accuracy against a test-retest ceiling agent accuracy / the person's own retest accuracy [0, 1+], higher Spearman (1904) Park et al. (2024); Toubia et al. (2025) – planned 0.3
Performance parity gap best-group minus worst-group performance [0, 1], lower Dwork et al. (2012); Hardt et al. (2016) Park et al. (2024) – planned 0.3
Log loss, Brier and ranked probability scores -log q_y; sum (q_k - 1[y = k])^2; sum_k (Q_k - 1[y <= k])^2 / (K - 1) [0, inf), lower Good (1952); Brier (1950); Epstein (1969); Murphy (1971); Gneiting and Raftery (2007) Chen et al. (2026); Binz et al. (2025) – planned 0.3
Expected calibration error sum_b \|B_b\| / N \|acc(B_b) - conf(B_b)\| [0, 1], lower Naeini et al. (2015); Guo et al. (2017) – – planned 0.3
ROC-AUC and pseudo-R^2 area under the ROC curve; 1 - NLL_model / NLL_random [0, 1], higher – Kim and Lee (2023); Binz et al. (2025) – planned 0.3
Match rate, F1 and election agreement share of matching choices; per-party F1; share of states called correctly [0, 1], higher – von der Heyde et al. (2026); Zhang et al. (2024) – planned 0.3

Effect-level and causal validity

Measure Formula Range Origin Applied to LLMs Python / R Status
Effect-size correlation with disattenuation r / sqrt(rel_model rel_human); sign agreement [-1, 1], higher Spearman (1904) Ashokkumar et al. (2026); Pfänder (2026) – planned 0.3
Type S and type M errors P(wrong sign \| significant); E\|estimate\| / \|true effect\| [0, 1] and [0, inf), lower Gelman and Carlin (2014); Pesaran and Timmermann (1992) Cui et al. (2025) – planned 0.3
Replication rate, coverage, inflation and false positives same sign and significant; interval coverage; Fisher-z ratio; significance on human nulls [0, 1], higher Schuirmann (1987) Cui et al. (2025); Aher et al. (2023); Horton et al. (2023) – planned 0.3
AMCE distance and conjoint battery \|\|AMCE_model - AMCE_human\|\|_2 with sign agreement and coverage [0, inf), lower – Takemoto (2024); Ahmad and Takemoto (2025); Hung et al. (2026) – planned 0.3
Treatment-effect recovery and negative controls tau_model / tau_human; TV of a negative-control outcome across arms [0, inf), closer to 1 – Persson et al. (2026); Lin et al. (2026); Gui and Toubia (2023); Li and Ji (2026) – planned 0.3
Willingness-to-pay recovery model willingness to pay over human willingness to pay ratio, closer to 1 – Brand et al. (2023) – reference only

Inference and uncertainty

Measure Formula Range Origin Applied to LLMs Python / R Status
Bootstrap percentile intervals percentile interval of a score over resamples of cells with replacement intervals, narrower Efron (1979); Efron and Tibshirani (1993) Pfänder (2026); Williams et al. (2026) popfidelity.uncertainty.bootstrap / uncertainty_bootstrap 0.1
Permutation null of cell correspondence scores after shuffling which cell each model distribution belongs to p-value, lower p Pitman (1937) Williams et al. (2026) popfidelity.uncertainty.permutation_null / uncertainty_permutation_null 0.1
Null predictions survey centre, leave-one-out, uniform, majority and permuted cells; above_null = (score - null) / (1 - null) scores, reference da Silva et al. (2026); Meister et al. (2025) da Silva et al. (2026); Boelaert et al. (2025); Dominguez-Olmedo et al. (2024) popfidelity.baselines.null_scores / baseline_null_scores 0.1
Holm, TOST, Wilcoxon and cluster-robust tests Holm step-down; two one-sided tests; signed-rank test; sandwich standard errors p-values, diagnostic Holm (1979); Schuirmann (1987); Wilcoxon (1945); Liang and Zeger (1986) da Silva et al. (2026) – planned 0.2
Prediction-powered inference theta = mean f(X_unlabelled) - mean(f(X) - Y); PPI++ tunes a power parameter; DSL estimates, narrower Angelopoulos et al. (2023); Angelopoulos et al. (2023); Egami et al. (2023) Krsteski et al. (2026); Broska et al. (2025) – planned 0.3
Effective human sample size ESS gain (Var_human / Var_method - 1); largest k with miscoverage at most gamma alpha [0, inf), higher Kish (1965) Krsteski et al. (2026); Huang et al. (2025); Gao et al. (2026) – planned 0.3
Rectification difficulty and residualised correlation Var(Y)(1 - rho^2); rho_T = Corr(Y - E[Y\|T], Y_hat - E[Y_hat\|T]) [0, inf), lower – Ye and Yoganarasimhan (2026); Wang et al. (2026) – planned 0.3
Sim-to-real quantile curve quantile function of a discrepancy proxy, summarised by CVaR curve, lower – Iyengar et al. (2025) – planned 0.3

Free-text measures (documented only)

Measure Formula Range Origin Applied to LLMs Python / R Status
Flattening and coverage of free text n-gram rarity, embedding dispersion, distinct options and Vendi score per identity varies, closer to humans Friedman and Dieng (2023) Wang et al. (2025) – reference only
Semantic diversity and hivemind homogeneity mean pairwise cosine distance of response embeddings [0, 2], closer to humans – Liu et al. (2024) – reference only
Persona consistency and PersonaScore prompt-to-line, line-to-line and Q&A consistency; rubric scores of evaluator models [0, 1] or 1 to 5, higher – Abdulhai et al. (2025); Samuel et al. (2025) – reference only
Computational Turing test held-out accuracy of a classifier telling human from model replies [0.5, 1], closer to 0.5 – Pagan et al. (2025); Mei et al. (2024) – reference only
Caricature: individuation and exaggeration classifier separability of persona outputs and normalised similarity to persona-topic axes [0, 1], lower – Cheng et al. (2023); Cheng et al. (2023) – reference only
Overton pluralism score share of human viewpoint clusters a response represents [0, 1], higher – Poole-Dayan et al. (2026) – reference only