Getting started¶
Install¶
pip install popfidelity
pip install "popfidelity[api]"
The first line installs the scoring library and the command line; the api
extra adds the clients for hosted models, local adds transformers and
parquet adds Parquet tables. PyPI has wheels for Linux, macOS and Windows on
Python 3.11 and newer; anywhere else pip builds from source, which needs a Rust
toolchain.
remotes::install_github("neemiasbsilva/popfidelity", subdir = "r")
The package compiles its Rust core, so it needs Cargo and rustc 1.85 or newer,
as SystemRequirements states.
cargo add popfidelity
The core crate has no dependencies and needs Rust 1.85 or newer.
git clone https://github.com/neemiasbsilva/popfidelity
cd popfidelity
uv sync --locked --group dev --extra api --extra parquet
./run.sh test
Score a model in Python¶
The paper's toy example has four cells answering yes or no. Model M2 keeps the survey's differences between cells but shifts every cell by 0.2.
import popfidelity as pf
from popfidelity.examples import toy
cells = toy()
survey, model = cells["survey"], cells["M2"]
scores = pf.score(survey, model, pf.Variant(round_pairs=10))
print(scores["pfs"], scores["score_accuracy"], scores["binding_component"])
pf.Cells holds one row per cell and one column per option. Build cells from a
long table with pf.Cells.from_long, read a whole file with
popfidelity.io.read_cells, or aggregate respondents with
popfidelity.aggregate.survey_cells.
Subgroups come from facets, the demographic families each cell belongs to:
spec = pf.read_facets("examples/ces/facets.yaml")
facets = spec.from_labels(survey.labels)
groups = pf.score_groups(survey, model, facets)
groups has one row for the pooled population and one for every level of every
family, with the components, their diagnostics and the binding component.
Score a model in R¶
library(popfidelity)
toy <- example_toy()
pfs_score(toy$survey, toy$M2, pfs_variant("tied", round_pairs = 10, based_on = "default"))
The R functions mirror the Python ones with a family prefix: pfs_, metric_,
align_, dispersion_, structure_, baseline_, uncertainty_ and
reliability_. Both packages call the same Rust code, and the parity check compares
their tables at 1e-12.
Score from the command line¶
popfidelity score \
--survey examples/ces/data/survey_cells.csv \
--model examples/ces/results/ollama-qwen3-vl-2b/model_cells.csv \
--questions examples/ces/questions.yaml \
--facets examples/ces/facets.yaml \
--min-valid 20 --out scores
popfidelity report --groups scores/groups.csv --out scores/report.md
score writes groups.csv, cells.csv and overall.csv, plus paired.csv with
--base-source and bootstrap.csv with --bootstrap.