Python API¶
The top-level package re-exports the scoring entry points; the submodules hold the other measures.
cells ¶
Cells
dataclass
¶
Response distributions of one question: one row per cell (subpopulation), one column per option.
subset ¶
subset(labels: Sequence[str]) -> Cells
These cells restricted to the given labels, in that order.
to_long ¶
to_long(**identity: Any) -> pd.DataFrame
The canonical long table: cell_id, question, option, share, n plus constant identity
columns.
from_long
classmethod
¶
from_long(
frame: DataFrame,
*,
options: Sequence[str] | None = None,
question: str | None = None,
cell: str = "cell_id",
option: str = "option",
share: str = "share",
count: str | None = "n",
) -> Cells
Cells from a long table, one row per cell and option, keeping the given option order.
from_wide
classmethod
¶
from_wide(
frame: DataFrame,
options: Sequence[str],
*,
cell: str = "cell_id",
count: str | None = None,
question: str = "",
normalise: bool = False,
) -> Cells
Cells from a table with one column per option.
common_labels ¶
common_labels(*groups: Cells) -> list[str]
Labels present in every group of cells, sorted.
align ¶
align(*groups: Cells) -> tuple[Cells, ...]
The groups restricted to their common cells, in sorted order; options must agree.
cells_by_question ¶
cells_by_question(
frame: DataFrame,
*,
questions: Iterable[str] | None = None,
options: dict[str, Sequence[str]] | None = None,
) -> dict[str, Cells]
Every question of a long table as cells, keyed by question id.
facets ¶
FacetFamily
dataclass
¶
One demographic family: level names in reporting order and the level of each cell (None when missing).
Facets
dataclass
¶
FacetSpec
dataclass
¶
A facets YAML spec: family order, level orders and how to read families off cell labels.
facets_from_frame ¶
facets_from_frame(
frame: DataFrame,
*,
cell: str = "cell_id",
families: Sequence[str] | None = None,
orders: Mapping[str, Sequence[str]] | None = None,
missing: Iterable[str] = MISSING,
) -> Facets
Facets from a table with one row per cell and one column per family.
facets_from_respondents ¶
facets_from_respondents(
respondents: DataFrame,
*,
cell: str = "cell_id",
families: Sequence[str],
orders: Mapping[str, Sequence[str]] | None = None,
missing: Iterable[str] = MISSING,
) -> Facets
Facets read off respondent rows; every respondent of a cell must share each family's level.
facet_spec_from_mapping ¶
facet_spec_from_mapping(
raw: Mapping[str, Any],
) -> FacetSpec
Builds a facets spec from its parsed YAML.
questions ¶
AnswerOption
dataclass
¶
One answer option: its survey label, the full-answer string the parser accepts, its token and position.
Question
dataclass
¶
A closed survey question with ordered answer options and the text each elicitation mode shows.
positions
property
¶
positions: list[float]
Option positions on the answer scale, unit spacing unless the spec says otherwise.
rendered_options ¶
rendered_options() -> str
The options as the prompt lists them, such as 'A. Very happy' or 'B. Quite happy'.
prompt_text ¶
prompt_text(mode: Mode) -> str
The question as asked under next-token (ntp) or full-answer (fa) elicitation.
normalise ¶
normalise(value: Any) -> str | None
A raw survey answer as one of the option labels, or None when missing.
label_index ¶
label_index(value: Any) -> int | None
Index of the option whose label matches a survey answer, None when missing.
canonical_index ¶
canonical_index(answer: str | None) -> int | None
Index of the option whose full-answer string matches, ignoring case, or None.
question_from_mapping ¶
question_from_mapping(raw: Mapping[str, Any]) -> Question
Builds a question from one entry of a questions spec.
questions_from_mapping ¶
questions_from_mapping(
raw: Mapping[str, Any],
) -> dict[str, Question]
Builds every question of a parsed questions spec, keyed by id.
read_questions ¶
read_questions(path: str | Path) -> dict[str, Question]
Reads a questions YAML spec into questions keyed by id.
variants ¶
pfs ¶
score_groups ¶
score_groups(
survey: Cells,
model: Cells,
facets: Facets | None = None,
*,
variant: Variant = DEFAULT,
variants: Sequence[Variant] | None = None,
identity: Mapping[str, Any] | None = None,
threads: int | None = None,
) -> pd.DataFrame
Population Fidelity of the pooled population and of every facet level, one row per subgroup.
score ¶
score(
survey: Cells,
model: Cells,
variant: Variant = DEFAULT,
*,
threads: int | None = None,
) -> dict[str, Any]
Population Fidelity Score, its components and diagnostics on the cells both sides share.
score_cells ¶
score_cells(
survey: Cells, model: Cells, **identity: Any
) -> pd.DataFrame
Each cell's normalised EMD, accuracy and Boelaert et al. quality band.
compare ¶
compare(
survey: Cells,
tuned: Cells,
base: Cells,
facets: Facets | None = None,
*,
variant: Variant = DEFAULT,
variants: Sequence[Variant] | None = None,
identity: Mapping[str, Any] | None = None,
threads: int | None = None,
) -> pd.DataFrame
A tuned and a base condition scored on the cells all three share, with base values and changes.
binding_component ¶
binding_component(
frame: DataFrame, prefix: str = ""
) -> pd.Series
The smallest of accuracy, adaptability and structure in each row, skipping undefined ones.
across_questions ¶
across_questions(
frame: DataFrame, keys: Sequence[str]
) -> pd.DataFrame
Scores averaged over questions: geometric means for PFS, arithmetic means for everything else.
across_replicates ¶
across_replicates(
frame: DataFrame,
keys: Sequence[str],
*,
spread: Sequence[str] = (),
pool_by: Sequence[str] | None = None,
replicate: str = "replicate",
confidence: float = 0.95,
) -> pd.DataFrame
Replicate runs averaged per condition, with run-to-run SDs and pooled t-interval half- widths.
metrics ¶
distance ¶
distance(
reference: Any,
compared: Any,
metric: str = "nemd",
*,
epsilon: float | None = None,
base: float | None = None,
bandwidth: float | None = None,
positions: Any = None,
) -> FloatArray
A distance or divergence between matching rows; one-row inputs broadcast. See NAMES and the
registry.
pairwise ¶
pairwise(
cells: Any,
metric: str = "nemd",
*,
form: Literal["condensed", "square"] = "condensed",
epsilon: float | None = None,
base: float | None = None,
bandwidth: float | None = None,
positions: Any = None,
) -> FloatArray
The metric between every pair of rows: scipy's condensed vector or the square matrix.
entropy ¶
entropy(cells: Any, base: float = 2.0) -> FloatArray
Shannon entropy of each row in the given log base.
normalised_entropy ¶
normalised_entropy(cells: Any) -> FloatArray
Entropy of each row over its maximum ln K (Dominguez-Olmedo et al., 2024).
effective_categories ¶
effective_categories(cells: Any) -> FloatArray
Effective number of options exp(H) of each row (Hill, 1973).
collision ¶
collision(cells: Any) -> FloatArray
Simpson's collision probability sum p_k^2 of each row (Simpson, 1949).
modal_concentration ¶
modal_concentration(cells: Any) -> FloatArray
Share of the most common answer in each row (Ozkan, 2026).
registry ¶
registry() -> pd.DataFrame
Every catalogued measure with its family, tier, formula, origin and LLM-application references.
alignment ¶
opinionqa_alignment ¶
opinionqa_alignment(
human: Any, model: Any, positions: Any = None
) -> FloatArray
OpinionQA alignment per row: 1 - W1 / (N - 1) on the ordinal scale (Santurkar et al.,
2023).
opinionqa_representativeness ¶
opinionqa_representativeness(
pairs: Sequence[tuple[Any, Any] | tuple[Any, Any, Any]],
) -> float
OpinionQA representativeness: alignment averaged over questions, each (human, model[,
positions]).
opinionqa_steerability ¶
opinionqa_steerability(
human: Sequence[Any],
steered: Sequence[Any],
positions: Sequence[Any] | None = None,
) -> float
OpinionQA steerability: per question the best alignment over steering prompts, then the mean.
opinionqa_consistency ¶
opinionqa_consistency(
alignment: DataFrame,
) -> tuple[str, float]
OpinionQA consistency from a topics-by-groups alignment table: the best group and how often it wins.
globalopinionqa_similarity ¶
globalopinionqa_similarity(
human: Any, model: Any, base: float = 2.0
) -> FloatArray
GlobalOpinionQA similarity per row: 1 - JS distance (Durmus et al., 2024); the log base is
a choice.
meister_alignment ¶
meister_alignment(
human: Any, model: Any
) -> dict[str, float]
Mean total variation over groups and questions with its uniform and majority baselines (Meister et al., 2025).
simbench_score ¶
simbench_score(
human: Any, model: Any, *, per_item: bool = False
) -> float
SimBench score 100 (1 - TVD / TVD_uniform) against the dataset-mean uniform distance (Hu et
al., 2026).
subpop_distance ¶
subpop_distance(human: Any, model: Any) -> FloatArray
SubPOP's unnormalised Wasserstein distance per row (Suh et al., 2025).
subpop_noise_floor ¶
subpop_noise_floor(
human: Any,
respondents: int,
*,
draws: int = 1000,
seed: int = 0,
) -> float
Mean distance between a human distribution and its multinomial resamples: SubPOP's lower bound.
relative_improvement ¶
relative_improvement(
model: float, zero_shot: float, floor: float
) -> float
Share of the gap from a zero-shot distance to the human noise floor that a model closes.
threshold_curve ¶
threshold_curve(
distances: Any, thresholds: Any = None
) -> pd.DataFrame
Share of questions whose distance is within each threshold (Zhao et al., 2024).
a_bias ¶
a_bias(label_shares: Any) -> float
A-bias |P(label A) - 1/k| of answers under randomised option order (Dominguez-Olmedo et
al., 2024).
dispersion ¶
sd_ratio ¶
sd_ratio(
survey: Cells,
model: Cells,
positions: Any = None,
*,
within: bool = False,
) -> float
Model answer SD over survey answer SD, pooled or averaged within cells (Bisbee et al., 2024).
normalised_variance ¶
normalised_variance(
cells: Cells, positions: Any = None
) -> FloatArray
Answer variance over the squared scale range, per cell (Williams et al., 2026).
gap_inflation ¶
gap_inflation(
survey: Cells,
model: Cells,
facets: Facets,
family: str,
*,
from_option: int,
) -> dict[str, float]
Model between-level gap in the share answering at or above an option over the survey gap (Chen et al., 2026).
diversity_ratios ¶
diversity_ratios(
survey: Cells, model: Cells
) -> dict[str, float]
Entropy, effective-category and collision ratios of the pooled model over the pooled survey answers.
support_collapse ¶
support_collapse(
survey: Cells, model: Cells, *, threshold: float = 0.0
) -> pd.DataFrame
Per cell, options the model never gives and the survey share on them (Doudkin, 2026).
modal_collapse ¶
modal_collapse(
survey: Cells, model: Cells, *, margin: float = 0.15
) -> pd.DataFrame
Per cell, modal concentration of each side and whether the model exceeds the survey by
margin (Ozkan, 2026).
stereotyping ¶
stereotyping(
survey: Cells, model: Cells, positions: Any = None
) -> dict[str, float]
Between-cell variance shares and Cramer's V, model minus survey (Chen et al., 2026); needs counts.
association ¶
association(cells: Cells) -> dict[str, float]
Pearson's chi-square test of independence between cells and answers, with Cramer's V.
collapsed_rows ¶
collapsed_rows(cells: Any) -> FloatArray
Each row's modal share, a quick view of mode collapse.
structure ¶
mantel ¶
mantel(
first: Any,
second: Any,
*,
method: Method = "spearman",
permutations: int = 999,
alternative: Alternative = "greater",
seed: int = 0,
threads: int | None = None,
) -> dict[str, float]
Mantel test of two condensed distance matrices on the same cells, permuting cells (Mantel, 1967).
structure_test ¶
structure_test(
survey: Cells,
model: Cells,
*,
method: Method = "spearman",
permutations: int = 999,
alternative: Alternative = "greater",
seed: int = 0,
threads: int | None = None,
) -> dict[str, float]
The PFS structure correlation with a Mantel permutation p-value instead of the naive t test.
correlation ¶
correlation(
first: Any, second: Any, method: Method = "spearman"
) -> tuple[float, float]
A correlation and its two-sided t-test p-value; Kendall's tau-b has no p-value here.
quality_bands ¶
quality_bands(errors: Any) -> list[str]
Boelaert et al.'s labels for cell errors, right-closed at 0.05, 0.10, 0.15 and 0.30.
quality_table ¶
quality_table(survey: Cells, model: Cells) -> pd.DataFrame
Percentage of cells in each quality band, the Machine Bias summary.
rank_average ¶
rank_average(values: Any) -> np.ndarray
Average ranks starting at one, as scipy.stats.rankdata gives them.
baselines ¶
survey_center ¶
survey_center(survey: Cells) -> Cells
Every cell predicted by the equal-weight survey centre: perfect centre, no variation, PFS 0.
leave_one_out ¶
leave_one_out(survey: Cells) -> Cells
Every cell predicted by the mean of the other cells.
majority ¶
majority(survey: Cells) -> Cells
Every cell predicted by all mass on its own modal answer (Meister et al., 2025).
permute_cells ¶
permute_cells(survey: Cells, seed: int = 0) -> Cells
The survey distributions shuffled across cells, breaking which cell holds which answers.
above_null ¶
above_null(score: float, null: float) -> float
(score - null) / (1 - null): zero at the null, one at perfect agreement.
null_scores ¶
null_scores(
survey: Cells, variant: Variant = DEFAULT, seed: int = 0
) -> pd.DataFrame
The population-level scores of every null prediction of the survey.
uncertainty ¶
bootstrap ¶
bootstrap(
survey: Cells,
model: Cells,
variant: Variant = DEFAULT,
*,
draws: int = 1000,
seed: int = 0,
confidence: float = 0.95,
threads: int | None = None,
) -> pd.DataFrame
Percentile intervals from resampling cells with replacement; each draw is seeded on its own.
permutation_null ¶
permutation_null(
survey: Cells,
model: Cells,
variant: Variant = DEFAULT,
*,
draws: int = 999,
seed: int = 0,
threads: int | None = None,
) -> pd.DataFrame
Scores after shuffling model cells: the null of no cell correspondence, with one-sided p-values in which near-ties count as extreme.
pooled_t_interval ¶
pooled_t_interval(
deviations: Any, runs: Any, confidence: float = 0.95
) -> np.ndarray
Half-widths of t intervals from one deviation pooled over conditions by degrees of freedom.
mean_interval ¶
mean_interval(
values: Any, confidence: float = 0.95
) -> tuple[float, float]
Mean of a few run values and the half-width t(1 - a/2, n - 1) s / sqrt(n).
student_t_quantile ¶
student_t_quantile(
probability: float, freedom: float
) -> float
Quantile of Student's t distribution, computed in the core so Python and R agree.
reliability ¶
icc ¶
icc(
ratings: Any,
model: IccModel = "agreement",
*,
average: bool = False,
) -> float
Intraclass correlation of a subjects-by-runs table (Shrout and Fleiss, 1979; McGraw and Wong, 1996).
cronbach_alpha ¶
cronbach_alpha(scores: Any) -> float
Cronbach's alpha of a respondents-by-items table (Cronbach, 1951).
spearman_brown ¶
spearman_brown(
reliability: float, factor: float = 2.0
) -> float
Reliability of a measure factor times as long (Spearman, 1910; Brown, 1910).
krippendorff_alpha ¶
krippendorff_alpha(
data: Any, level: Level = "ordinal"
) -> float
Krippendorff's alpha of a units-by-coders table with NaN for missing codes (Krippendorff, 1970).
fleiss_kappa ¶
fleiss_kappa(counts: Any) -> float
Fleiss' kappa from a subjects-by-categories table of rater counts (Fleiss, 1971).
noise_to_signal ¶
noise_to_signal(
seed_sd: float, prompt_sd: float, signal_sd: float
) -> float
(sd_seed + sd_prompt) / sd_signal; above one the signal is lost in noise (Nguyen and Ahmad,
2026).
mismatch_rate ¶
mismatch_rate(first: Any, second: Any) -> float
Share of items whose two answer indices differ, such as first-token versus text (Wang et al., 2024).
run_spread ¶
run_spread(values: Any) -> tuple[float, float]
Mean and sample SD of repeated run values.
aggregate ¶
survey_cells ¶
survey_cells(
respondents: DataFrame,
question: Question,
*,
cell: str = "cell_id",
answer: str | None = None,
weight: str | None = None,
) -> Cells
Each cell's survey answer shares: one-hot answers in option order, averaged over valid answers.
fa_cells ¶
fa_cells(
answers: DataFrame,
respondents: DataFrame,
question: Question,
*,
cell: str = "cell_id",
respondent: str = "respondent_id",
answer: str = "answer",
) -> Cells
Each cell's full-answer shares: parsed answers joined to respondents, averaged over valid answers.
ntp_cells ¶
ntp_cells(
profiles: DataFrame,
respondents: DataFrame,
question: Question,
*,
cell: str = "cell_id",
profile: str = "profile_id",
columns: Sequence[str] | None = None,
) -> Cells
Each cell's next-token shares: every respondent takes its profile's distribution, averaged per cell.
retain ¶
retain(
survey: Cells,
models: Mapping[str, Cells],
*,
min_valid: int = MIN_VALID_ANSWERS,
) -> tuple[Cells, dict[str, Cells]]
The cells kept for scoring: present in the survey and with min_valid valid answers in every
model.
record_cells ¶
record_cells(
records: DataFrame,
respondents: DataFrame,
question: Question,
mode: str,
*,
cell: str = "cell_id",
respondent: str = "respondent_id",
profile: str = "profile_id",
) -> Cells
Cells from elicited records: full answers per respondent, or next-token shares per profile.
io ¶
read_records ¶
read_records(path: str | Path) -> pd.DataFrame
Every record under a run directory or in one JSON-lines file, the latest per prompt, flattened.
read_table ¶
read_table(path: str | Path) -> pd.DataFrame
A CSV, TSV, Parquet or JSON-lines table, chosen by the file suffix; label columns stay text and numbers read back exactly.
write_table ¶
write_table(frame: DataFrame, path: str | Path) -> None
Writes a table as CSV or Parquet by suffix, creating the folder.
read_cells ¶
read_cells(
path: str | Path,
questions: Mapping[str, Question] | None = None,
**filters: Any,
) -> dict[str, Cells]
Every question of a long cells table, keyed by question, after keeping rows that match
filters.
write_cells ¶
write_cells(
cells: Iterable[Cells],
path: str | Path,
**identity: Any,
) -> None
Writes cells as one long table with constant identity columns such as source or mode.