Skip to contents

A language model that answers survey questions can match the average answer of a population and still erase the differences between the people in it. The Population Fidelity Score (PFS) of da Silva et al. (2026) checks three conditions on cells, the subpopulations a survey reports, each with a distribution over the answer options:

  • accuracy: each cell’s model distribution is close to the survey’s, one minus the mean normalised earth mover’s distance (nEMD);
  • adaptability: the model’s cells differ from one another as much as the survey’s do, the ratio of the median pairwise nEMD of the two sides, scored as min(A, 1/A);
  • structure: the cells that differ in the survey are the ones that differ in the model, the Spearman correlation of the two pairwise-distance vectors, clipped at zero.

PFS is their geometric mean, so a model fails as soon as one condition fails.

The paper’s toy example

Four cells answer a binary question. Three models have the same accuracy and the same centre, and differ in how they spread the cells.

toy <- example_toy()
toy$survey$shares
#>   yes  no
#> A 0.2 0.8
#> B 0.4 0.6
#> C 0.6 0.4
#> D 0.8 0.2
tied <- pfs_variant(round_pairs = 10)
scores <- do.call(rbind, lapply(c("M1", "M2", "M3"), function(name) {
  pfs_score(toy$survey, toy[[name]], tied)
}))
cbind(model = c("M1", "M2", "M3"), scores[, c("score_accuracy", "score_adaptability", "score_structure", "pfs", "binding_component")])
#>   model score_accuracy score_adaptability score_structure       pfs
#> 1    M1            0.8                0.5               1 0.7368063
#> 2    M2            0.8                1.0               1 0.9283178
#> 3    M3            0.8                1.0               0 0.0000000
#>   binding_component
#> 1      adaptability
#> 2          accuracy
#> 3         structure

M1 compresses the differences between cells, M2 keeps them and M3 scrambles which cell is which. Rounding the pairwise distances to ten decimals restores the exact ties the example is built on.

Subgroups and paired comparisons

Scores are reported for the pooled population and for every level of every demographic family. Facets come from a table with one row per cell, or from a YAML spec that reads families off the cell labels.

set.seed(1)
labels <- sprintf("cell%02d", 1:24)
draw <- function() {
  values <- matrix(stats::rgamma(24 * 4, shape = 2), nrow = 24)
  shares <- values / rowSums(values)
  dimnames(shares) <- list(labels, c("Never", "Sometimes", "Often", "Always"))
  shares
}
survey <- cells_from_matrix(draw(), question = "habit")
centre <- colMeans(survey$shares)
shrunk <- cells_from_matrix(
  sweep(survey$shares, 2, centre) * 0.4 + matrix(centre, 24, 4, byrow = TRUE, dimnames = dimnames(survey$shares)),
  question = "habit"
)
noisy <- cells_from_matrix(draw(), question = "habit")
facets <- facets_from_frame(data.frame(
  cell_id = labels,
  age = rep(c("18-34", "35-54", "55+"), 8),
  region = rep(c("North", "South"), each = 12)
))
groups <- pfs_score_groups(survey, shrunk, facets)
groups[, c("group", "level", "n_cells", "score_accuracy", "score_adaptability", "score_structure", "pfs")]
#> # A tibble: 6 × 7
#>   group    level n_cells score_accuracy score_adaptability score_structure   pfs
#>   <chr>    <chr>   <int>          <dbl>              <dbl>           <dbl> <dbl>
#> 1 populat… all        24          0.936              0.400               1 0.721
#> 2 age      18-34       8          0.936              0.4                 1 0.721
#> 3 age      35-54       8          0.921              0.400               1 0.717
#> 4 age      55+         8          0.950              0.400               1 0.724
#> 5 region   North      12          0.930              0.400               1 0.719
#> 6 region   South      12          0.942              0.400               1 0.722

A model shrunk toward the survey centre keeps the arrangement of the cells, so its structure stays perfect while its adaptability falls to the shrinkage factor. pfs_compare() scores a tuned and a base condition on the cells they share and adds the changes:

paired <- pfs_compare(survey, shrunk, noisy)
paired[, c("group", "level", "pfs", "base_pfs", "delta_pfs_vs_base")]
#> # A tibble: 1 × 5
#>   group      level   pfs base_pfs delta_pfs_vs_base
#>   <chr>      <chr> <dbl>    <dbl>             <dbl>
#> 1 population all   0.721        0             0.721

Other measures

The package also computes the distances and scores of the literature on synthetic survey samples. Each measure is catalogued with the paper it comes from and the paper that applied it to language models:

registry <- metric_registry()
head(registry[registry$tier == "1", c("id", "name", "origin", "applied", "r")])
#> # A tibble: 6 × 5
#>   id               name                      origin                applied r    
#>   <chr>            <chr>                     <chr>                 <chr>   <chr>
#> 1 pfs              Population Fidelity Score dasilva2026populatio… dasilv… pfs_…
#> 2 pfs_accuracy     Accuracy component        boelaert2025machine,… dasilv… pfs_…
#> 3 pfs_adaptability Adaptability component    boelaert2025machine,… dasilv… pfs_…
#> 4 pfs_structure    Structure component       spearman1904proof, m… dasilv… pfs_…
#> 5 pfs_center       Centre alignment          boelaert2025machine   boelae… pfs_…
#> 6 pfs_binding      Limiting component        dasilva2026population dasilv… pfs_…
metric_distance(survey, shrunk, "js_distance")[1:3]
#> [1] 0.1397188 0.2077028 0.2116162
align_meister(survey, shrunk)
#> $total_variation
#> [1] 0.117823
#> 
#> $uniform_baseline
#> [1] 0.1987829
#> 
#> $majority_baseline
#> [1] 0.5848184
structure_test(survey, shrunk, permutations = 199)
#> $statistic
#> [1] 1
#> 
#> $p_value
#> [1] 0.005
#> 
#> $permutations
#> [1] 199
#> 
#> $null_mean
#> [1] -0.003257428
#> 
#> $null_sd
#> [1] 0.08044845

Data from files

read_cells() reads the long table that the Python package and the popfidelity command line write, one row per cell and option, and read_records() reads the raw records of an elicitation run, which aggregate_records() turns into cells.

questions <- read_questions(example_path("toy", "questions.yaml"))
survey_cells <- read_cells(example_path("toy", "survey_cells.csv"), questions)
survey_cells$toy
#> <popfidelity_cells> toy: 4 cells, 2 options (yes, no)