Benchmarks¶
Benchmarks whose official scores the registry documents.
| Benchmark | Score | Licence | Reference |
|---|---|---|---|
| OpinionQA | representativeness, steerability, consistency | Pew terms; fetch at runtime | Santurkar et al. (2023) |
| GlobalOpinionQA | 1 - JS distance per country | CC BY-NC-SA 4.0 | Durmus et al. (2024) |
| SubPOP | Wasserstein distance with human noise floor | CC BY-NC-SA 4.0, gated | Suh et al. (2025) |
| SimBench | 100 (1 - TVD / TVD_uniform) | CC BY-NC-SA 4.0 | Hu et al. (2026) |
| WorldValuesBench | Wasserstein threshold curve | WVS terms; fetch at runtime | Zhao et al. (2024) |
| Machine Bias replication package | nEMD, quality bands, adaptability, centre model | MIT code; WVS data under WVS terms | Boelaert et al. (2025) |
| Twin-2K-500 | accuracy and correlation against a test-retest ceiling | CC BY 4.0 | Toubia et al. (2025); Peng et al. (2025) |
| SocSci210 | Wasserstein distance on rescaled answers | CC BY 4.0 | Kolluri et al. (2025) |
| Psych-101 | held-out negative log-likelihood and pseudo-R^2 | CC BY 4.0 and CC BY-ND 4.0 | Binz et al. (2025) |
| PersonaGym | PersonaScore | CC BY-NC-SA 4.0 | Samuel et al. (2025) |
| OvertonBench | OvertonScore | see source | Poole-Dayan et al. (2026) |
| CulturalBench | multiple-choice and true/false accuracy | CC BY 4.0 | Chiu et al. (2025) |
| NormAd | acceptability accuracy with country, value or rule context | CC BY 4.0 | Rao et al. (2025) |
| ValueBench | value orientation agreement | MIT | Ren et al. (2024) |
| PRISM | preference and rating dataset | CC BY 4.0 and CC BY-NC 4.0 | Kirk et al. (2024) |
| Silicon Sample Benchmark | direction agreement, correlations, RMSE, calibration | see source | Pfänder (2026) |
| SynthBench | SynthBench parity score | MIT | DataViking-Tech (2026) |