A language model that answers survey questions can match the average answer of a population and still erase the differences between the people in it. The Population Fidelity Score (PFS) of da Silva et al. (2026) checks three conditions on cells, the subpopulations a survey reports, each with a distribution over the answer options:
- accuracy: each cell’s model distribution is close to the survey’s, one minus the mean normalised earth mover’s distance (nEMD);
-
adaptability: the model’s cells differ from one
another as much as the survey’s do, the ratio of the median pairwise
nEMD of the two sides, scored as
min(A, 1/A); - structure: the cells that differ in the survey are the ones that differ in the model, the Spearman correlation of the two pairwise-distance vectors, clipped at zero.
PFS is their geometric mean, so a model fails as soon as one condition fails.
The paper’s toy example
Four cells answer a binary question. Three models have the same accuracy and the same centre, and differ in how they spread the cells.
toy <- example_toy()
toy$survey$shares
#> yes no
#> A 0.2 0.8
#> B 0.4 0.6
#> C 0.6 0.4
#> D 0.8 0.2
tied <- pfs_variant(round_pairs = 10)
scores <- do.call(rbind, lapply(c("M1", "M2", "M3"), function(name) {
pfs_score(toy$survey, toy[[name]], tied)
}))
cbind(model = c("M1", "M2", "M3"), scores[, c("score_accuracy", "score_adaptability", "score_structure", "pfs", "binding_component")])
#> model score_accuracy score_adaptability score_structure pfs
#> 1 M1 0.8 0.5 1 0.7368063
#> 2 M2 0.8 1.0 1 0.9283178
#> 3 M3 0.8 1.0 0 0.0000000
#> binding_component
#> 1 adaptability
#> 2 accuracy
#> 3 structureM1 compresses the differences between cells,
M2 keeps them and M3 scrambles which cell is
which. Rounding the pairwise distances to ten decimals restores the
exact ties the example is built on.
Subgroups and paired comparisons
Scores are reported for the pooled population and for every level of every demographic family. Facets come from a table with one row per cell, or from a YAML spec that reads families off the cell labels.
set.seed(1)
labels <- sprintf("cell%02d", 1:24)
draw <- function() {
values <- matrix(stats::rgamma(24 * 4, shape = 2), nrow = 24)
shares <- values / rowSums(values)
dimnames(shares) <- list(labels, c("Never", "Sometimes", "Often", "Always"))
shares
}
survey <- cells_from_matrix(draw(), question = "habit")
centre <- colMeans(survey$shares)
shrunk <- cells_from_matrix(
sweep(survey$shares, 2, centre) * 0.4 + matrix(centre, 24, 4, byrow = TRUE, dimnames = dimnames(survey$shares)),
question = "habit"
)
noisy <- cells_from_matrix(draw(), question = "habit")
facets <- facets_from_frame(data.frame(
cell_id = labels,
age = rep(c("18-34", "35-54", "55+"), 8),
region = rep(c("North", "South"), each = 12)
))
groups <- pfs_score_groups(survey, shrunk, facets)
groups[, c("group", "level", "n_cells", "score_accuracy", "score_adaptability", "score_structure", "pfs")]
#> # A tibble: 6 × 7
#> group level n_cells score_accuracy score_adaptability score_structure pfs
#> <chr> <chr> <int> <dbl> <dbl> <dbl> <dbl>
#> 1 populat… all 24 0.936 0.400 1 0.721
#> 2 age 18-34 8 0.936 0.4 1 0.721
#> 3 age 35-54 8 0.921 0.400 1 0.717
#> 4 age 55+ 8 0.950 0.400 1 0.724
#> 5 region North 12 0.930 0.400 1 0.719
#> 6 region South 12 0.942 0.400 1 0.722A model shrunk toward the survey centre keeps the arrangement of the
cells, so its structure stays perfect while its adaptability falls to
the shrinkage factor. pfs_compare() scores a tuned and a
base condition on the cells they share and adds the changes:
paired <- pfs_compare(survey, shrunk, noisy)
paired[, c("group", "level", "pfs", "base_pfs", "delta_pfs_vs_base")]
#> # A tibble: 1 × 5
#> group level pfs base_pfs delta_pfs_vs_base
#> <chr> <chr> <dbl> <dbl> <dbl>
#> 1 population all 0.721 0 0.721Other measures
The package also computes the distances and scores of the literature on synthetic survey samples. Each measure is catalogued with the paper it comes from and the paper that applied it to language models:
registry <- metric_registry()
head(registry[registry$tier == "1", c("id", "name", "origin", "applied", "r")])
#> # A tibble: 6 × 5
#> id name origin applied r
#> <chr> <chr> <chr> <chr> <chr>
#> 1 pfs Population Fidelity Score dasilva2026populatio… dasilv… pfs_…
#> 2 pfs_accuracy Accuracy component boelaert2025machine,… dasilv… pfs_…
#> 3 pfs_adaptability Adaptability component boelaert2025machine,… dasilv… pfs_…
#> 4 pfs_structure Structure component spearman1904proof, m… dasilv… pfs_…
#> 5 pfs_center Centre alignment boelaert2025machine boelae… pfs_…
#> 6 pfs_binding Limiting component dasilva2026population dasilv… pfs_…
metric_distance(survey, shrunk, "js_distance")[1:3]
#> [1] 0.1397188 0.2077028 0.2116162
align_meister(survey, shrunk)
#> $total_variation
#> [1] 0.117823
#>
#> $uniform_baseline
#> [1] 0.1987829
#>
#> $majority_baseline
#> [1] 0.5848184
structure_test(survey, shrunk, permutations = 199)
#> $statistic
#> [1] 1
#>
#> $p_value
#> [1] 0.005
#>
#> $permutations
#> [1] 199
#>
#> $null_mean
#> [1] -0.003257428
#>
#> $null_sd
#> [1] 0.08044845Data from files
read_cells() reads the long table that the Python
package and the popfidelity command line write, one row per
cell and option, and read_records() reads the raw records
of an elicitation run, which aggregate_records() turns into
cells.
questions <- read_questions(example_path("toy", "questions.yaml"))
survey_cells <- read_cells(example_path("toy", "survey_cells.csv"), questions)
survey_cells$toy
#> <popfidelity_cells> toy: 4 cells, 2 options (yes, no)