Skip to content

popfidelity

A language model that answers survey questions can match the average answer of a population and still erase the differences between the people in it. popfidelity measures how well a model's answers represent a human population, group by group, and gives Python, R and the command line the same numbers from one Rust core.

Its flagship measure is the Population Fidelity Score (PFS) of da Silva et al. (2026). PFS compares the survey and the model on cells, the subpopulations a survey reports, and checks three conditions:

Component Question it answers Score
Accuracy Is each cell's model distribution close to the survey's? one minus the mean normalised earth mover's distance
Adaptability Do the model's cells differ from one another as much as the survey's? min(A, 1/A) of the ratio of median pairwise distances
Structure Do the cells that differ in the survey also differ in the model? Spearman correlation of the two pairwise-distance vectors, clipped at zero

PFS is their geometric mean, so a model fails as soon as one condition fails. The package reports it for the pooled population, every demographic subgroup, paired tuned-against-base comparisons, replicate runs and sets of questions.

Beside PFS sit the measures the literature uses to compare synthetic and human samples (distances, the OpinionQA, GlobalOpinionQA, Meister, SimBench and SubPOP scores, dispersion and homogenisation, the Mantel test, null baselines, bootstrap intervals and reliability), each with the work it comes from and the work that applied it to language models. A runner elicits next-token and full-answer distributions from hosted APIs, OpenAI-compatible servers, Ollama and local transformers.

Get started See the examples