Skip to content

Benchmarks

Benchmarks whose official scores the registry documents.

Benchmark Score Licence Reference
OpinionQA representativeness, steerability, consistency Pew terms; fetch at runtime Santurkar et al. (2023)
GlobalOpinionQA 1 - JS distance per country CC BY-NC-SA 4.0 Durmus et al. (2024)
SubPOP Wasserstein distance with human noise floor CC BY-NC-SA 4.0, gated Suh et al. (2025)
SimBench 100 (1 - TVD / TVD_uniform) CC BY-NC-SA 4.0 Hu et al. (2026)
WorldValuesBench Wasserstein threshold curve WVS terms; fetch at runtime Zhao et al. (2024)
Machine Bias replication package nEMD, quality bands, adaptability, centre model MIT code; WVS data under WVS terms Boelaert et al. (2025)
Twin-2K-500 accuracy and correlation against a test-retest ceiling CC BY 4.0 Toubia et al. (2025); Peng et al. (2025)
SocSci210 Wasserstein distance on rescaled answers CC BY 4.0 Kolluri et al. (2025)
Psych-101 held-out negative log-likelihood and pseudo-R^2 CC BY 4.0 and CC BY-ND 4.0 Binz et al. (2025)
PersonaGym PersonaScore CC BY-NC-SA 4.0 Samuel et al. (2025)
OvertonBench OvertonScore see source Poole-Dayan et al. (2026)
CulturalBench multiple-choice and true/false accuracy CC BY 4.0 Chiu et al. (2025)
NormAd acceptability accuracy with country, value or rule context CC BY 4.0 Rao et al. (2025)
ValueBench value orientation agreement MIT Ren et al. (2024)
PRISM preference and rating dataset CC BY 4.0 and CC BY-NC 4.0 Kirk et al. (2024)
Silicon Sample Benchmark direction agreement, correlations, RMSE, calibration see source Pfänder (2026)
SynthBench SynthBench parity score MIT DataViking-Tech (2026)