Eliciting answers¶
popfidelity elicit --config run.yaml interviews a model once per respondent and
question and stores every request as a JSON-lines record that a rerun resumes
from. Each respondent's demographics become prior question-and-answer turns, so the
prompt reads like a survey interview.
Backends¶
backend.kind |
Endpoints | Next-token | Full answer |
|---|---|---|---|
openai_compatible |
presets openai, vllm, llamacpp, hf-router, tgi, openrouter, together, fireworks, deepseek, gemini-openai, or any base_url |
where the endpoint returns log-probabilities | yes |
ollama |
Ollama's native generate API, raw prompts by default | yes | yes |
anthropic |
Claude models | no | yes |
google |
Gemini models | no | yes |
transformers |
a local Hugging Face model, exact softmax over the option tokens | yes | yes |
Keys come from environment variables or a .env file in the working directory
(OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY, HF_TOKEN and others;
see .env.example).
A configuration¶
schema_version: 1
run: ollama-qwen3-vl-2b
backend:
kind: ollama
model: qwen3-vl:2b
base_url: ${OLLAMA_BASE_URL:-http://localhost:11434}
questions: ../questions.yaml
respondents: ../data/respondents_sample.csv
context: ../context.yaml
records: ../records
modes: [ntp, fa]
seed: 20240110
ntp: {top_logprobs: 20}
fa: {temperature: 0.7, max_tokens: 12, attempts: 20}
probe: {prompts: 32, min_mass: 0.10, degrade_to_fa: true}
Run it with --dry-run to count the requests, then with --limit 8 before a full
run: hosted models are billed per request.
Records and manifests¶
Every record holds the prompt and its SHA-256, the seed, the parsed answer or the
option shares with their probability mass, the attempts, and the status. A
manifest.json per question records the backend, its capabilities, the sampling
settings, the probe result, the respondent file's SHA-256 and the library version.
popfidelity aggregate records --config run.yaml --out model_cells.csv turns the
records into cells.