Overview
Abstract
This study examines how persona prompting shapes language generated by two multimodal large language models in urban perception, a setting for examining subjective interpretations of shared visual evidence. We organize outputs into three functional levels: descriptive grounding (captions), intermediate semantic layer (perception tags), and interpretive framing (justifications). Using approximately 60,000 persona-conditioned annotations per model from Qwen3-VL-8B and Gemma-4-E4B-it, we find that captions converge strongly across persona profiles and show only small attribute-associated differences. Justifications vary substantially more: economic status produces the largest difference in both models, with political orientation and personality also prominent. Paired image-level comparisons confirm larger justification than caption differences for these three attributes. For perception tags, personas sharing the same attribute level produce more similar tag sets than personas with different attribute levels, with the largest separation observed for economic status. Exploratory topic analysis further reveals persona-specific evaluative emphasis. Across models, profile-pair similarity patterns are strongly correlated for all three output types, although agreement is lowest for justifications. Overall, persona prompting affects interpretive framing more strongly than descriptive grounding.
Key Results
Both models show the same qualitative ordering: persona-profile separation is weakest for captions and strongest for justifications. DSI measures relative diagonal-over-off-diagonal separation in each 24 × 24 profile-similarity matrix.
| Output level | Qwen3-VL-8B | Gemma-4-E4B-it |
|---|---|---|
| Descriptive Captions | 2.3% | 3.7% |
| Intermediate Perception tags | 11.7% | 25.1% |
| Interpretive Justifications | 14.9% | 30.0% |
- Descriptions converge. Mean caption similarities remain approximately 0.87–0.91 across persona groups.
- Economic status is the strongest dimension. Its mean within-minus-cross justification cosine-similarity difference is +0.0617 for Qwen3-VL and +0.1015 for Gemma4.
- Cross-model consistency is high. Profile-pair similarity matrices correlate at Pearson r = 0.80–0.89.
Study at a Glance
Each model used the same images, prompts, persona profiles, and generation protocol. The 24 profiles combine gender (2), economic status (2), political orientation (2), and personality (3). For each image and persona dimension, the analysis compares annotations from agents sharing an attribute level with annotations from agents at different levels. Captions and justifications are compared with sentence-embedding cosine similarity; perception tags use exact-set Jaccard similarity.
Responsible Interpretation
These results describe model behavior under persona prompts; they do not estimate human demographic differences. The profiles are controlled prompting conditions rather than complete identities, and synthetic personas may reproduce biased or stereotypical associations. Human correspondence would require demographically matched human annotations.
Citation
If this work informs your research, please cite our paper at the Workshop on Pluralistic AI & NLP: Diversity-aware, Sociotechnical, Responsible Alignment (PANDORA 2026).
@inproceedings{silva2026persona,
title={Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation},
author={Neemias Buceli da Silva and Matt Ratto and Myriam Delgado and Rodrigo Minetto and Daniel Silver and Thiago H. Silva},
booktitle={EMNLP26 Workshop on Pluralistic AI {\&} NLP: Diversity-aware, Sociotechnical, Responsible Alignment (PANDORA)},
year={2026},
}