Skip to contents

Population Fidelity

pfs_across_questions()
Scores averaged over questions: geometric means for PFS, arithmetic means for everything else
pfs_across_replicates()
Replicate runs averaged per condition, with run-to-run SDs and pooled t-interval half-widths
pfs_binding_component()
The smallest of accuracy, adaptability and structure in each row, skipping undefined ones
pfs_compare()
A tuned and a base condition scored on the cells all three share, with base values and changes
pfs_delta_column()
Name of the tuned-minus-base change of a metric
pfs_score()
Population Fidelity Score, its components and diagnostics on the cells both sides share
pfs_score_cells()
Each cell's normalised EMD, accuracy and Boelaert et al. quality band
pfs_score_groups()
Population Fidelity of the pooled population and of every facet level, one row per subgroup
pfs_variant()
One specification of the score; the default is the Population Fidelity Score of da Silva et al. (2026)
pfs_variants()
Every named specification of the score from the paper's sensitivity analysis

Cells, questions and facets

cells_align()
Groups of cells restricted to their common cells, in sorted order; options must agree
cells_by_question()
Every question of a long table as cells, keyed by question id
cells_common_labels()
Labels present in every group of cells, sorted
cells_from_long()
Cells of one question from a long table with one row per cell and option
cells_from_matrix()
Cells of one question from a matrix of answer shares
cells_from_wide()
Cells of one question from a table with one column per option
cells_subset()
These cells restricted to the given labels, in that order
cells_to_long()
The canonical long table of cells, with constant identity columns
facets_for_cells()
The facets of the given cells, in their order; unknown cells raise an error
facets_from_frame()
Facets from a table with one row per cell and one column per demographic family
facets_from_labels()
Facets parsed from cell labels with a spec's pattern and derived families
facets_from_respondents()
Facets read off respondent rows; every respondent of a cell must share each level
facets_from_spec_frame()
Facets from a per-cell table with a spec's families and level orders
facets_to_frame()
One row per cell and one column per demographic family
read_questions()
Read a questions YAML spec into questions keyed by id
read_facets()
Read a facets YAML spec: family order, level orders and how to read families off cell labels

Reading, writing and aggregating

read_table()
A CSV, TSV, Parquet or JSON-lines table, chosen by the file suffix
write_table()
Write a table as CSV, or Parquet by suffix, creating the folder; numbers keep 17 digits
read_cells()
Every question of a long cells table, keyed by question, after keeping rows that match the filters
write_cells()
Write cells as one long table with constant identity columns such as source or mode
read_records()
Every record under a run directory or in one JSON-lines file, the latest per prompt, flattened
aggregate_fa()
Each cell's full-answer shares: parsed answers joined to respondents, averaged over valid answers
aggregate_ntp()
Each cell's next-token shares: every respondent takes its profile's distribution, averaged per cell
aggregate_records()
Cells from elicited records: full answers per respondent, or next-token shares per profile
aggregate_survey()
Each cell's survey answer shares: one-hot answers in option order, averaged over valid answers
retain_cells()
The cells kept for scoring: present in the survey and with min_valid valid answers in every model
report_markdown()
A short Markdown report: the population rows, then the subgroups with the lowest PFS
plot_components()
Dot plot of accuracy, adaptability, structure and PFS for every subgroup of a scored table

Distances and the measure catalogue

metric_benchmarks()
Benchmarks whose official scores the catalogue documents
metric_collision()
Simpson's collision probability, the sum of squared shares, of each row (Simpson, 1949)
metric_distance()
A distance or divergence between matching rows; one-row inputs broadcast
metric_effective_categories()
Effective number of options exp(H) of each row (Hill, 1973)
metric_entropy()
Shannon entropy of each row in the given log base
metric_modal_concentration()
Share of the most common answer in each row (Ozkan, 2026)
metric_names()
Names of every distance and divergence the core computes
metric_normalised_entropy()
Entropy of each row over its maximum log K (Dominguez-Olmedo et al., 2024)
metric_pairwise()
The metric between every pair of rows: the condensed vector of dist() or the square matrix
metric_registry()
Every catalogued measure with its family, tier, formula, origin and LLM-application references
popfidelity_core_version()
Version of the compiled core

Alignment scores from the literature

align_a_bias()
A-bias |P(label A) - 1/k| of answers under randomised option order (Dominguez-Olmedo et al., 2024)
align_globalopinionqa_similarity()
GlobalOpinionQA similarity per row: 1 - JS distance (Durmus et al., 2024); the log base is a choice
align_meister()
Mean total variation over groups and questions with uniform and majority baselines (Meister et al., 2025)
align_opinionqa_alignment()
OpinionQA alignment per row: 1 - W1 / (N - 1) on the ordinal scale (Santurkar et al., 2023)
align_opinionqa_consistency()
OpinionQA consistency from a topics-by-groups alignment table: the best group and how often it wins
align_opinionqa_representativeness()
OpinionQA representativeness: alignment averaged over questions
align_opinionqa_steerability()
OpinionQA steerability: per question the best alignment over steering prompts, then the mean
align_relative_improvement()
Share of the gap from a zero-shot distance to the human noise floor that a model closes
align_simbench()
SimBench score 100 (1 - TVD / TVD_uniform) against the dataset-mean uniform distance (Hu et al., 2026)
align_subpop_distance()
SubPOP's unnormalised Wasserstein distance per row (Suh et al., 2025)
align_subpop_noise_floor()
Mean distance between a human distribution and its multinomial resamples, SubPOP's lower bound
align_threshold_curve()
Share of questions whose distance is within each threshold (Zhao et al., 2024)

Dispersion and structure

dispersion_association()
Pearson's chi-square test of independence between cells and answers, with Cramer's V
dispersion_diversity_ratios()
Entropy, effective-category and collision ratios of the pooled model over the pooled survey answers
dispersion_gap_inflation()
Model between-level gap in the share answering at or above an option over the survey gap (Chen et al., 2026)
dispersion_modal_collapse()
Per cell, modal concentration of each side and whether the model exceeds the survey by margin (Ozkan, 2026)
dispersion_normalised_variance()
Answer variance over the squared scale range, per cell (Williams et al., 2026)
dispersion_sd_ratio()
Model answer SD over survey answer SD, pooled or averaged within cells (Bisbee et al., 2024)
dispersion_stereotyping()
Between-cell variance shares and Cramer's V, model minus survey (Chen et al., 2026); needs counts
dispersion_support_collapse()
Per cell, options the model never gives and the survey share on them (Doudkin, 2026)
structure_correlation()
A correlation and its two-sided t-test p-value; Kendall's tau-b has no p-value here
structure_mantel()
Mantel test of two condensed distance matrices on the same cells, permuting cells (Mantel, 1967)
structure_quality_bands()
Boelaert et al.'s labels for cell errors, right-closed at 0.05, 0.10, 0.15 and 0.30
structure_quality_table()
Percentage of cells in each quality band, the Machine Bias summary
structure_test()
The PFS structure correlation with a Mantel permutation p-value instead of the naive t test

Baselines, uncertainty and reliability

baseline_above_null()
(score - null) / (1 - null): zero at the null, one at perfect agreement
baseline_leave_one_out()
Every cell predicted by the mean of the other cells
baseline_majority()
Every cell predicted by all mass on its own modal answer (Meister et al., 2025)
baseline_null_scores()
The population-level scores of every null prediction of the survey
baseline_permute_cells()
The survey distributions shuffled across cells, breaking which cell holds which answers
baseline_survey_center()
Every cell predicted by the equal-weight survey centre: perfect centre, no variation, PFS 0
baseline_uniform()
Every cell predicted by the uniform distribution
uncertainty_bootstrap()
Percentile intervals from resampling cells with replacement; each draw is seeded on its own
uncertainty_mean_interval()
Mean of a few run values and the half-width t(1 - a/2, n - 1) s / sqrt(n)
uncertainty_permutation_null()
Scores after shuffling model cells: the null of no cell correspondence, with one-sided p-values in which near-ties count as extreme
uncertainty_pooled_t_interval()
Half-widths of t intervals from one deviation pooled over conditions by degrees of freedom
uncertainty_t_quantile()
Quantile of Student's t distribution, computed in the core so Python and R agree
reliability_cronbach_alpha()
Cronbach's alpha of a respondents-by-items table (Cronbach, 1951)
reliability_fleiss_kappa()
Fleiss' kappa from a subjects-by-categories table of rater counts (Fleiss, 1971)
reliability_icc()
Intraclass correlation of a subjects-by-runs table (Shrout and Fleiss, 1979; McGraw and Wong, 1996)
reliability_krippendorff_alpha()
Krippendorff's alpha of a units-by-coders table with NA for missing codes (Krippendorff, 1970)
reliability_mismatch_rate()
Share of items whose two answer indices differ, such as first-token versus text (Wang et al., 2024)
reliability_noise_to_signal()
(sd_seed + sd_prompt) / sd_signal; above one the signal is lost in noise (Nguyen and Ahmad, 2026)
reliability_run_spread()
Mean and sample SD of repeated run values
reliability_spearman_brown()
Reliability of a measure factor times as long (Spearman, 1910; Brown, 1910)

Examples

example_toy()
The four-cell binary example of the Population Fidelity paper: the survey and models M1 to M3
example_path()
Path to a file of the bundled examples