Concepts¶
Cells, options and facets¶
A question has ordered answer options. A cell is a subpopulation, such as women aged 18 to 29 with a college degree, and its distribution is the share of the cell choosing each option. The survey and the model each give one distribution per cell; scores compare them on the cells both share.
Facets say which level of each demographic family a cell belongs to. A facets YAML spec lists the families with their level order and, when cell labels encode the levels, a regular expression with named groups that reads them:
schema_version: 1
families:
- name: gender
levels: [Female, Male]
- name: age
levels: [18-29, 30-44, 45-64, 65+]
labels:
pattern: "^(?P<gender>[^|]+)\\|(?P<age>[^|]+)$"
The Population Fidelity Score¶
With survey distributions s_i and model distributions m_i over K options:
- nEMD is the earth mover's distance on the ordinal scale divided by
K - 1,sum_k |F_s(k) - F_m(k)| / (K - 1). - Accuracy is
1 - mean_i nEMD(s_i, m_i). - Adaptability compares the median pairwise nEMD between cells:
A = D_model / D_survey, scoredmin(A, 1/A). Below one, the model compresses the differences between cells; above one, it exaggerates them. - Structure is the Spearman correlation between the survey's and the model's pairwise distances, clipped at zero. A model that gives every cell the same answer scores zero.
- PFS is the geometric mean of the three. The smallest component is the binding one.
- Centre alignment,
1 - nEMD(mean s, mean m), is reported beside PFS, not inside it: a model can match the population average and still fail.
Adaptability needs two cells and structure three. Scores of a subgroup use only the cells in it; paired comparisons use only the cells the tuned, base and survey tables share.
Variants¶
The paper's sensitivity analysis varies every choice: mean instead of median
dispersion, a one-sided adaptability score, Pearson or Kendall structure, signed
structure, an arithmetic mean, and respondent-weighted errors. pf.VARIANTS in
Python and pfs_variants() in R list them; pass several to score_groups to get
one row per variant.
Elicitation modes¶
Next-token probabilities (ntp) read the model's probability of each option's
first token after the interview prompt and renormalise them over the options.
Full answers (fa) sample a short answer per respondent and parse it. A cell's
model distribution averages its respondents. Endpoints without log-probabilities
fall back to full answers after a 32-prompt capability probe, as in the paper.