Examples¶
| Example | Survey | Model answers | Shows |
|---|---|---|---|
| Toy | the paper's toy cells | models M1 to M3 | equal accuracy, different fidelity |
| CES | Cooperative Election Study 2024, 96 cells | qwen3-vl:2b on Ollama; configurations for OpenAI, Anthropic, Google, vLLM, Hugging Face and transformers | the whole pipeline |
| Twin-2K-500 | Twin-2K-500 wave 4, 24 cells | GPT-4.1-mini digital twins and a human retest | a proprietary model scored without an API key |
| GlobalOpinionQA | four World Values Survey items in 90 countries, fetched at run time | qwen3-vl:2b on Ollama, one country per interview | similarity against fidelity across countries |
| Machine Bias | World Values Survey cells | archived and fine-tuned models | the paper's tables recomputed |
In the CES example, a 2-billion-parameter model is close to every cell on average (accuracy 0.71 to 0.90) but does not order the cells as the survey does, so structure binds and PFS stays at 0.30 or below. In Twin-2K-500, people answering the same items again reach PFS 0.95; digital twins reach 0.66 with a full persona and 0.58 with demographics only, held back by adaptability.