Skip to content

Examples

Example Survey Model answers Shows
Toy the paper's toy cells models M1 to M3 equal accuracy, different fidelity
CES Cooperative Election Study 2024, 96 cells qwen3-vl:2b on Ollama; configurations for OpenAI, Anthropic, Google, vLLM, Hugging Face and transformers the whole pipeline
Twin-2K-500 Twin-2K-500 wave 4, 24 cells GPT-4.1-mini digital twins and a human retest a proprietary model scored without an API key
GlobalOpinionQA four World Values Survey items in 90 countries, fetched at run time qwen3-vl:2b on Ollama, one country per interview similarity against fidelity across countries
Machine Bias World Values Survey cells archived and fine-tuned models the paper's tables recomputed

In the CES example, a 2-billion-parameter model is close to every cell on average (accuracy 0.71 to 0.90) but does not order the cells as the survey does, so structure binds and PFS stays at 0.30 or below. In Twin-2K-500, people answering the same items again reach PFS 0.95; digital twins reach 0.66 with a full persona and 0.58 with demographics only, held back by adaptability.