Skip to content

Python API

The top-level package re-exports the scoring entry points; the submodules hold the other measures.

cells

Cells dataclass

Response distributions of one question: one row per cell (subpopulation), one column per option.

subset

subset(labels: Sequence[str]) -> Cells

These cells restricted to the given labels, in that order.

to_long

to_long(**identity: Any) -> pd.DataFrame

The canonical long table: cell_id, question, option, share, n plus constant identity columns.

from_long classmethod

from_long(
    frame: DataFrame,
    *,
    options: Sequence[str] | None = None,
    question: str | None = None,
    cell: str = "cell_id",
    option: str = "option",
    share: str = "share",
    count: str | None = "n",
) -> Cells

Cells from a long table, one row per cell and option, keeping the given option order.

from_wide classmethod

from_wide(
    frame: DataFrame,
    options: Sequence[str],
    *,
    cell: str = "cell_id",
    count: str | None = None,
    question: str = "",
    normalise: bool = False,
) -> Cells

Cells from a table with one column per option.

common_labels

common_labels(*groups: Cells) -> list[str]

Labels present in every group of cells, sorted.

align

align(*groups: Cells) -> tuple[Cells, ...]

The groups restricted to their common cells, in sorted order; options must agree.

cells_by_question

cells_by_question(
    frame: DataFrame,
    *,
    questions: Iterable[str] | None = None,
    options: dict[str, Sequence[str]] | None = None,
) -> dict[str, Cells]

Every question of a long table as cells, keyed by question id.

facets

FacetFamily dataclass

One demographic family: level names in reporting order and the level of each cell (None when missing).

codes

codes() -> IntArray

Each cell's level index, -1 when missing.

Facets dataclass

Which level of every demographic family each cell belongs to.

for_cells

for_cells(labels: Sequence[str]) -> Facets

The facets of the given cells, in their order; unknown cells raise.

specs

specs() -> list[tuple[str, list[str], IntArray]]

The families as the compiled core reads them.

to_frame

to_frame() -> pd.DataFrame

One row per cell, one column per family.

FacetSpec dataclass

A facets YAML spec: family order, level orders and how to read families off cell labels.

from_labels

from_labels(labels: Iterable[str]) -> Facets

Facets parsed from cell labels with the spec's pattern and derived families.

from_frame

from_frame(
    frame: DataFrame, *, cell: str = "cell_id"
) -> Facets

Facets from a per-cell table, with the spec's families and level orders.

facets_from_frame

facets_from_frame(
    frame: DataFrame,
    *,
    cell: str = "cell_id",
    families: Sequence[str] | None = None,
    orders: Mapping[str, Sequence[str]] | None = None,
    missing: Iterable[str] = MISSING,
) -> Facets

Facets from a table with one row per cell and one column per family.

facets_from_respondents

facets_from_respondents(
    respondents: DataFrame,
    *,
    cell: str = "cell_id",
    families: Sequence[str],
    orders: Mapping[str, Sequence[str]] | None = None,
    missing: Iterable[str] = MISSING,
) -> Facets

Facets read off respondent rows; every respondent of a cell must share each family's level.

facet_spec_from_mapping

facet_spec_from_mapping(
    raw: Mapping[str, Any],
) -> FacetSpec

Builds a facets spec from its parsed YAML.

read_facets

read_facets(path: str | Path) -> FacetSpec

Reads a facets YAML spec.

questions

AnswerOption dataclass

One answer option: its survey label, the full-answer string the parser accepts, its token and position.

Question dataclass

A closed survey question with ordered answer options and the text each elicitation mode shows.

labels property

labels: tuple[str, ...]

Option labels in response order.

positions property

positions: list[float]

Option positions on the answer scale, unit spacing unless the spec says otherwise.

canonical property

canonical: tuple[str, ...]

The full-answer strings in response order.

tokens property

tokens: tuple[str, ...]

The next-token strings in response order.

rendered_options

rendered_options() -> str

The options as the prompt lists them, such as 'A. Very happy' or 'B. Quite happy'.

prompt_text

prompt_text(mode: Mode) -> str

The question as asked under next-token (ntp) or full-answer (fa) elicitation.

normalise

normalise(value: Any) -> str | None

A raw survey answer as one of the option labels, or None when missing.

label_index

label_index(value: Any) -> int | None

Index of the option whose label matches a survey answer, None when missing.

canonical_index

canonical_index(answer: str | None) -> int | None

Index of the option whose full-answer string matches, ignoring case, or None.

question_from_mapping

question_from_mapping(raw: Mapping[str, Any]) -> Question

Builds a question from one entry of a questions spec.

questions_from_mapping

questions_from_mapping(
    raw: Mapping[str, Any],
) -> dict[str, Question]

Builds every question of a parsed questions spec, keyed by id.

read_questions

read_questions(path: str | Path) -> dict[str, Question]

Reads a questions YAML spec into questions keyed by id.

variants

Variant dataclass

One specification of the score; the default is the Population Fidelity Score of da Silva et al. (2026).

spec

spec() -> VariantSpec

The variant as the compiled core reads it.

with_name

with_name(name: str) -> Variant

The same specification under another name.

variant

variant(name: str) -> Variant

A named specification of the paper's sensitivity analysis, such as mean_dispersion.

pfs

delta_column

delta_column(metric: str) -> str

Name of the tuned-minus-base change of a metric.

score_groups

score_groups(
    survey: Cells,
    model: Cells,
    facets: Facets | None = None,
    *,
    variant: Variant = DEFAULT,
    variants: Sequence[Variant] | None = None,
    identity: Mapping[str, Any] | None = None,
    threads: int | None = None,
) -> pd.DataFrame

Population Fidelity of the pooled population and of every facet level, one row per subgroup.

score

score(
    survey: Cells,
    model: Cells,
    variant: Variant = DEFAULT,
    *,
    threads: int | None = None,
) -> dict[str, Any]

Population Fidelity Score, its components and diagnostics on the cells both sides share.

score_cells

score_cells(
    survey: Cells, model: Cells, **identity: Any
) -> pd.DataFrame

Each cell's normalised EMD, accuracy and Boelaert et al. quality band.

compare

compare(
    survey: Cells,
    tuned: Cells,
    base: Cells,
    facets: Facets | None = None,
    *,
    variant: Variant = DEFAULT,
    variants: Sequence[Variant] | None = None,
    identity: Mapping[str, Any] | None = None,
    threads: int | None = None,
) -> pd.DataFrame

A tuned and a base condition scored on the cells all three share, with base values and changes.

binding_component

binding_component(
    frame: DataFrame, prefix: str = ""
) -> pd.Series

The smallest of accuracy, adaptability and structure in each row, skipping undefined ones.

across_questions

across_questions(
    frame: DataFrame, keys: Sequence[str]
) -> pd.DataFrame

Scores averaged over questions: geometric means for PFS, arithmetic means for everything else.

across_replicates

across_replicates(
    frame: DataFrame,
    keys: Sequence[str],
    *,
    spread: Sequence[str] = (),
    pool_by: Sequence[str] | None = None,
    replicate: str = "replicate",
    confidence: float = 0.95,
) -> pd.DataFrame

Replicate runs averaged per condition, with run-to-run SDs and pooled t-interval half- widths.

metrics

distance

distance(
    reference: Any,
    compared: Any,
    metric: str = "nemd",
    *,
    epsilon: float | None = None,
    base: float | None = None,
    bandwidth: float | None = None,
    positions: Any = None,
) -> FloatArray

A distance or divergence between matching rows; one-row inputs broadcast. See NAMES and the registry.

pairwise

pairwise(
    cells: Any,
    metric: str = "nemd",
    *,
    form: Literal["condensed", "square"] = "condensed",
    epsilon: float | None = None,
    base: float | None = None,
    bandwidth: float | None = None,
    positions: Any = None,
) -> FloatArray

The metric between every pair of rows: scipy's condensed vector or the square matrix.

entropy

entropy(cells: Any, base: float = 2.0) -> FloatArray

Shannon entropy of each row in the given log base.

normalised_entropy

normalised_entropy(cells: Any) -> FloatArray

Entropy of each row over its maximum ln K (Dominguez-Olmedo et al., 2024).

effective_categories

effective_categories(cells: Any) -> FloatArray

Effective number of options exp(H) of each row (Hill, 1973).

collision

collision(cells: Any) -> FloatArray

Simpson's collision probability sum p_k^2 of each row (Simpson, 1949).

modal_concentration

modal_concentration(cells: Any) -> FloatArray

Share of the most common answer in each row (Ozkan, 2026).

registry

registry() -> pd.DataFrame

Every catalogued measure with its family, tier, formula, origin and LLM-application references.

benchmarks

benchmarks() -> pd.DataFrame

Benchmarks whose official scores the catalogue documents.

alignment

opinionqa_alignment

opinionqa_alignment(
    human: Any, model: Any, positions: Any = None
) -> FloatArray

OpinionQA alignment per row: 1 - W1 / (N - 1) on the ordinal scale (Santurkar et al., 2023).

opinionqa_representativeness

opinionqa_representativeness(
    pairs: Sequence[tuple[Any, Any] | tuple[Any, Any, Any]],
) -> float

OpinionQA representativeness: alignment averaged over questions, each (human, model[, positions]).

opinionqa_steerability

opinionqa_steerability(
    human: Sequence[Any],
    steered: Sequence[Any],
    positions: Sequence[Any] | None = None,
) -> float

OpinionQA steerability: per question the best alignment over steering prompts, then the mean.

opinionqa_consistency

opinionqa_consistency(
    alignment: DataFrame,
) -> tuple[str, float]

OpinionQA consistency from a topics-by-groups alignment table: the best group and how often it wins.

globalopinionqa_similarity

globalopinionqa_similarity(
    human: Any, model: Any, base: float = 2.0
) -> FloatArray

GlobalOpinionQA similarity per row: 1 - JS distance (Durmus et al., 2024); the log base is a choice.

meister_alignment

meister_alignment(
    human: Any, model: Any
) -> dict[str, float]

Mean total variation over groups and questions with its uniform and majority baselines (Meister et al., 2025).

simbench_score

simbench_score(
    human: Any, model: Any, *, per_item: bool = False
) -> float

SimBench score 100 (1 - TVD / TVD_uniform) against the dataset-mean uniform distance (Hu et al., 2026).

subpop_distance

subpop_distance(human: Any, model: Any) -> FloatArray

SubPOP's unnormalised Wasserstein distance per row (Suh et al., 2025).

subpop_noise_floor

subpop_noise_floor(
    human: Any,
    respondents: int,
    *,
    draws: int = 1000,
    seed: int = 0,
) -> float

Mean distance between a human distribution and its multinomial resamples: SubPOP's lower bound.

relative_improvement

relative_improvement(
    model: float, zero_shot: float, floor: float
) -> float

Share of the gap from a zero-shot distance to the human noise floor that a model closes.

threshold_curve

threshold_curve(
    distances: Any, thresholds: Any = None
) -> pd.DataFrame

Share of questions whose distance is within each threshold (Zhao et al., 2024).

a_bias

a_bias(label_shares: Any) -> float

A-bias |P(label A) - 1/k| of answers under randomised option order (Dominguez-Olmedo et al., 2024).

dispersion

sd_ratio

sd_ratio(
    survey: Cells,
    model: Cells,
    positions: Any = None,
    *,
    within: bool = False,
) -> float

Model answer SD over survey answer SD, pooled or averaged within cells (Bisbee et al., 2024).

normalised_variance

normalised_variance(
    cells: Cells, positions: Any = None
) -> FloatArray

Answer variance over the squared scale range, per cell (Williams et al., 2026).

gap_inflation

gap_inflation(
    survey: Cells,
    model: Cells,
    facets: Facets,
    family: str,
    *,
    from_option: int,
) -> dict[str, float]

Model between-level gap in the share answering at or above an option over the survey gap (Chen et al., 2026).

diversity_ratios

diversity_ratios(
    survey: Cells, model: Cells
) -> dict[str, float]

Entropy, effective-category and collision ratios of the pooled model over the pooled survey answers.

support_collapse

support_collapse(
    survey: Cells, model: Cells, *, threshold: float = 0.0
) -> pd.DataFrame

Per cell, options the model never gives and the survey share on them (Doudkin, 2026).

modal_collapse

modal_collapse(
    survey: Cells, model: Cells, *, margin: float = 0.15
) -> pd.DataFrame

Per cell, modal concentration of each side and whether the model exceeds the survey by margin (Ozkan, 2026).

stereotyping

stereotyping(
    survey: Cells, model: Cells, positions: Any = None
) -> dict[str, float]

Between-cell variance shares and Cramer's V, model minus survey (Chen et al., 2026); needs counts.

association

association(cells: Cells) -> dict[str, float]

Pearson's chi-square test of independence between cells and answers, with Cramer's V.

collapsed_rows

collapsed_rows(cells: Any) -> FloatArray

Each row's modal share, a quick view of mode collapse.

structure

mantel

mantel(
    first: Any,
    second: Any,
    *,
    method: Method = "spearman",
    permutations: int = 999,
    alternative: Alternative = "greater",
    seed: int = 0,
    threads: int | None = None,
) -> dict[str, float]

Mantel test of two condensed distance matrices on the same cells, permuting cells (Mantel, 1967).

structure_test

structure_test(
    survey: Cells,
    model: Cells,
    *,
    method: Method = "spearman",
    permutations: int = 999,
    alternative: Alternative = "greater",
    seed: int = 0,
    threads: int | None = None,
) -> dict[str, float]

The PFS structure correlation with a Mantel permutation p-value instead of the naive t test.

correlation

correlation(
    first: Any, second: Any, method: Method = "spearman"
) -> tuple[float, float]

A correlation and its two-sided t-test p-value; Kendall's tau-b has no p-value here.

quality_bands

quality_bands(errors: Any) -> list[str]

Boelaert et al.'s labels for cell errors, right-closed at 0.05, 0.10, 0.15 and 0.30.

quality_table

quality_table(survey: Cells, model: Cells) -> pd.DataFrame

Percentage of cells in each quality band, the Machine Bias summary.

rank_average

rank_average(values: Any) -> np.ndarray

Average ranks starting at one, as scipy.stats.rankdata gives them.

baselines

survey_center

survey_center(survey: Cells) -> Cells

Every cell predicted by the equal-weight survey centre: perfect centre, no variation, PFS 0.

leave_one_out

leave_one_out(survey: Cells) -> Cells

Every cell predicted by the mean of the other cells.

uniform

uniform(survey: Cells) -> Cells

Every cell predicted by the uniform distribution.

majority

majority(survey: Cells) -> Cells

Every cell predicted by all mass on its own modal answer (Meister et al., 2025).

permute_cells

permute_cells(survey: Cells, seed: int = 0) -> Cells

The survey distributions shuffled across cells, breaking which cell holds which answers.

above_null

above_null(score: float, null: float) -> float

(score - null) / (1 - null): zero at the null, one at perfect agreement.

null_scores

null_scores(
    survey: Cells, variant: Variant = DEFAULT, seed: int = 0
) -> pd.DataFrame

The population-level scores of every null prediction of the survey.

uncertainty

bootstrap

bootstrap(
    survey: Cells,
    model: Cells,
    variant: Variant = DEFAULT,
    *,
    draws: int = 1000,
    seed: int = 0,
    confidence: float = 0.95,
    threads: int | None = None,
) -> pd.DataFrame

Percentile intervals from resampling cells with replacement; each draw is seeded on its own.

permutation_null

permutation_null(
    survey: Cells,
    model: Cells,
    variant: Variant = DEFAULT,
    *,
    draws: int = 999,
    seed: int = 0,
    threads: int | None = None,
) -> pd.DataFrame

Scores after shuffling model cells: the null of no cell correspondence, with one-sided p-values in which near-ties count as extreme.

pooled_t_interval

pooled_t_interval(
    deviations: Any, runs: Any, confidence: float = 0.95
) -> np.ndarray

Half-widths of t intervals from one deviation pooled over conditions by degrees of freedom.

mean_interval

mean_interval(
    values: Any, confidence: float = 0.95
) -> tuple[float, float]

Mean of a few run values and the half-width t(1 - a/2, n - 1) s / sqrt(n).

student_t_quantile

student_t_quantile(
    probability: float, freedom: float
) -> float

Quantile of Student's t distribution, computed in the core so Python and R agree.

reliability

icc

icc(
    ratings: Any,
    model: IccModel = "agreement",
    *,
    average: bool = False,
) -> float

Intraclass correlation of a subjects-by-runs table (Shrout and Fleiss, 1979; McGraw and Wong, 1996).

cronbach_alpha

cronbach_alpha(scores: Any) -> float

Cronbach's alpha of a respondents-by-items table (Cronbach, 1951).

spearman_brown

spearman_brown(
    reliability: float, factor: float = 2.0
) -> float

Reliability of a measure factor times as long (Spearman, 1910; Brown, 1910).

krippendorff_alpha

krippendorff_alpha(
    data: Any, level: Level = "ordinal"
) -> float

Krippendorff's alpha of a units-by-coders table with NaN for missing codes (Krippendorff, 1970).

fleiss_kappa

fleiss_kappa(counts: Any) -> float

Fleiss' kappa from a subjects-by-categories table of rater counts (Fleiss, 1971).

noise_to_signal

noise_to_signal(
    seed_sd: float, prompt_sd: float, signal_sd: float
) -> float

(sd_seed + sd_prompt) / sd_signal; above one the signal is lost in noise (Nguyen and Ahmad, 2026).

mismatch_rate

mismatch_rate(first: Any, second: Any) -> float

Share of items whose two answer indices differ, such as first-token versus text (Wang et al., 2024).

run_spread

run_spread(values: Any) -> tuple[float, float]

Mean and sample SD of repeated run values.

aggregate

survey_cells

survey_cells(
    respondents: DataFrame,
    question: Question,
    *,
    cell: str = "cell_id",
    answer: str | None = None,
    weight: str | None = None,
) -> Cells

Each cell's survey answer shares: one-hot answers in option order, averaged over valid answers.

fa_cells

fa_cells(
    answers: DataFrame,
    respondents: DataFrame,
    question: Question,
    *,
    cell: str = "cell_id",
    respondent: str = "respondent_id",
    answer: str = "answer",
) -> Cells

Each cell's full-answer shares: parsed answers joined to respondents, averaged over valid answers.

ntp_cells

ntp_cells(
    profiles: DataFrame,
    respondents: DataFrame,
    question: Question,
    *,
    cell: str = "cell_id",
    profile: str = "profile_id",
    columns: Sequence[str] | None = None,
) -> Cells

Each cell's next-token shares: every respondent takes its profile's distribution, averaged per cell.

retain

retain(
    survey: Cells,
    models: Mapping[str, Cells],
    *,
    min_valid: int = MIN_VALID_ANSWERS,
) -> tuple[Cells, dict[str, Cells]]

The cells kept for scoring: present in the survey and with min_valid valid answers in every model.

record_cells

record_cells(
    records: DataFrame,
    respondents: DataFrame,
    question: Question,
    mode: str,
    *,
    cell: str = "cell_id",
    respondent: str = "respondent_id",
    profile: str = "profile_id",
) -> Cells

Cells from elicited records: full answers per respondent, or next-token shares per profile.

io

read_records

read_records(path: str | Path) -> pd.DataFrame

Every record under a run directory or in one JSON-lines file, the latest per prompt, flattened.

read_table

read_table(path: str | Path) -> pd.DataFrame

A CSV, TSV, Parquet or JSON-lines table, chosen by the file suffix; label columns stay text and numbers read back exactly.

write_table

write_table(frame: DataFrame, path: str | Path) -> None

Writes a table as CSV or Parquet by suffix, creating the folder.

read_cells

read_cells(
    path: str | Path,
    questions: Mapping[str, Question] | None = None,
    **filters: Any,
) -> dict[str, Cells]

Every question of a long cells table, keyed by question, after keeping rows that match filters.

write_cells

write_cells(
    cells: Iterable[Cells],
    path: str | Path,
    **identity: Any,
) -> None

Writes cells as one long table with constant identity columns such as source or mode.

plot

components

components(
    groups: DataFrame, title: str | None = None
) -> Any

Dot plot of accuracy, adaptability, structure and PFS for every subgroup of a scored table; needs the plot extra.