Skip to content

Concepts

Cells, options and facets

A question has ordered answer options. A cell is a subpopulation, such as women aged 18 to 29 with a college degree, and its distribution is the share of the cell choosing each option. The survey and the model each give one distribution per cell; scores compare them on the cells both share.

Facets say which level of each demographic family a cell belongs to. A facets YAML spec lists the families with their level order and, when cell labels encode the levels, a regular expression with named groups that reads them:

schema_version: 1
families:
  - name: gender
    levels: [Female, Male]
  - name: age
    levels: [18-29, 30-44, 45-64, 65+]
labels:
  pattern: "^(?P<gender>[^|]+)\\|(?P<age>[^|]+)$"

The Population Fidelity Score

With survey distributions s_i and model distributions m_i over K options:

  • nEMD is the earth mover's distance on the ordinal scale divided by K - 1, sum_k |F_s(k) - F_m(k)| / (K - 1).
  • Accuracy is 1 - mean_i nEMD(s_i, m_i).
  • Adaptability compares the median pairwise nEMD between cells: A = D_model / D_survey, scored min(A, 1/A). Below one, the model compresses the differences between cells; above one, it exaggerates them.
  • Structure is the Spearman correlation between the survey's and the model's pairwise distances, clipped at zero. A model that gives every cell the same answer scores zero.
  • PFS is the geometric mean of the three. The smallest component is the binding one.
  • Centre alignment, 1 - nEMD(mean s, mean m), is reported beside PFS, not inside it: a model can match the population average and still fail.

Adaptability needs two cells and structure three. Scores of a subgroup use only the cells in it; paired comparisons use only the cells the tuned, base and survey tables share.

Variants

The paper's sensitivity analysis varies every choice: mean instead of median dispersion, a one-sided adaptability score, Pearson or Kendall structure, signed structure, an arithmetic mean, and respondent-weighted errors. pf.VARIANTS in Python and pfs_variants() in R list them; pass several to score_groups to get one row per variant.

Elicitation modes

Next-token probabilities (ntp) read the model's probability of each option's first token after the interview prompt and renormalise them over the options. Full answers (fa) sample a short answer per respondent and parse it. A cell's model distribution averages its respondents. Endpoints without log-probabilities fall back to full answers after a 32-prompt capability probe, as in the paper.