EMNLP 2026 Workshop PANDORA

Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation

The same city scene. Similar descriptions. Different interpretations.

Neemias B da Silva1,2Matt Ratto1Myriam Delgado2Rodrigo Minetto2Daniel Silver1Thiago H Silva1,2
1 University of Toronto, Canada2 Federal University of Technology – Parana (UTFPR), Brazil

Overview

Panel A moves from an urban scene and persona profile through an MLLM to captions, perception tags, and justifications with increasing interpretive abstraction. Panel B applies three personas to the same brick campus scene: captions describe similar visible content while tags and justifications differ in emphasis.
Overview. Persona-conditioned outputs are separated into descriptive grounding, an intermediate semantic layer, and interpretive framing. Captions converge on visible content, while perception tags and justifications vary in emphasis and evaluation.

Abstract

This study examines how persona prompting shapes language generated by two multimodal large language models in urban perception, a setting for examining subjective interpretations of shared visual evidence. We organize outputs into three functional levels: descriptive grounding (captions), intermediate semantic layer (perception tags), and interpretive framing (justifications). Using approximately 60,000 persona-conditioned annotations per model from Qwen3-VL-8B and Gemma-4-E4B-it, we find that captions converge strongly across persona profiles and show only small attribute-associated differences. Justifications vary substantially more: economic status produces the largest difference in both models, with political orientation and personality also prominent. Paired image-level comparisons confirm larger justification than caption differences for these three attributes. For perception tags, personas sharing the same attribute level produce more similar tag sets than personas with different attribute levels, with the largest separation observed for economic status. Exploratory topic analysis further reveals persona-specific evaluative emphasis. Across models, profile-pair similarity patterns are strongly correlated for all three output types, although agreement is lowest for justifications. Overall, persona prompting affects interpretive framing more strongly than descriptive grounding.

Key Results

Both models show the same qualitative ordering: persona-profile separation is weakest for captions and strongest for justifications. DSI measures relative diagonal-over-off-diagonal separation in each 24 × 24 profile-similarity matrix.

Diagonal Strength Index (DSI) by model and output type
Output levelQwen3-VL-8BGemma-4-E4B-it
Descriptive Captions2.3%3.7%
Intermediate Perception tags11.7%25.1%
Interpretive Justifications14.9%30.0%
  • Descriptions converge. Mean caption similarities remain approximately 0.87–0.91 across persona groups.
  • Economic status is the strongest dimension. Its mean within-minus-cross justification cosine-similarity difference is +0.0617 for Qwen3-VL and +0.1015 for Gemma4.
  • Cross-model consistency is high. Profile-pair similarity matrices correlate at Pearson r = 0.80–0.89.
Two-panel grouped bar chart comparing within-minus-cross similarity differences by persona dimension and output type for Qwen3-VL and Gemma4
Persona effect by dimension and modality. Bar length is the within-minus-cross mean; whiskers are 95% BCa bootstrap confidence intervals over 50 images. Captions and justifications use cosine similarity, while perception tags use Jaccard similarity and are not directly comparable. The gray ±0.01 band is an exploratory caption-effect reference, not an equivalence threshold. In both models, justification differences exceed caption differences for economic status, political orientation, and personality.

Study at a Glance

119,707persona-conditioned annotations
24persona profiles
50urban-scene images
2 MLLMsmatched protocol

Each model used the same images, prompts, persona profiles, and generation protocol. The 24 profiles combine gender (2), economic status (2), political orientation (2), and personality (3). For each image and persona dimension, the analysis compares annotations from agents sharing an attribute level with annotations from agents at different levels. Captions and justifications are compared with sentence-embedding cosine similarity; perception tags use exact-set Jaccard similarity.

Responsible Interpretation

These results describe model behavior under persona prompts; they do not estimate human demographic differences. The profiles are controlled prompting conditions rather than complete identities, and synthetic personas may reproduce biased or stereotypical associations. Human correspondence would require demographically matched human annotations.

Citation

If this work informs your research, please cite our paper at the Workshop on Pluralistic AI & NLP: Diversity-aware, Sociotechnical, Responsible Alignment (PANDORA 2026).

BibTeX
@inproceedings{silva2026persona,
  title={Persona Prompting in Multimodal Urban Perception: Descriptive Convergence and Interpretive Variation},
  author={Neemias Buceli da Silva and Matt Ratto and Myriam Delgado and Rodrigo Minetto and Daniel Silver and Thiago H. Silva},
  booktitle={EMNLP26 Workshop on Pluralistic AI {\&} NLP: Diversity-aware, Sociotechnical, Responsible Alignment (PANDORA)},
  year={2026},
}