conflicting interpretations in dialogue
A perspectivist account of grounding and misunderstanding in asymmetric dialogue, and what vision-language models make of it. Part of the DeMeVa project.
Two people working on a shared task rarely have access to the same information. They still have to talk as if they do. This project asks what happens in the gap: how participants converge on a shared interpretation, when they only appear to, and whether current models can tell the difference.
The work is part of DeMeVa (“Dealing with Meaning Variation”), in the sub-project on conflicting interpretations in dialogue.
A perspectivist scheme for MapTask
Dialogue corpora normally assign one gold referent to each referring expression. That convention makes misunderstanding unrepresentable: if the annotation records what an expression “really” referred to, it cannot record that the speaker and the addressee took it differently. We instead annotate speaker intent and addressee interpretation as separate labels, on the HCRC MapTask corpus, where the two participants hold deliberately mismatched maps (Li et al., 2026). The scheme was applied with LLM assistance to about 13,000 referring expressions.
Two findings follow. First, once lexical variants of the same landmark are unified, full misunderstandings are rare. Second, what does reliably produce divergent interpretations is a multiplicity discrepancy: a landmark that occurs more than once on one participant’s map but only once on the other’s. In those cases the dialogue often proceeds smoothly to a shared phrase while the two participants are pointing at different objects. Apparent grounding can mask referential misalignment.
What vision-language models do with it
The annotated data supports a second question: can a model judge whether two interlocutors have actually converged? We formulate this as a prediction task over MapTask dialogues and vary what the model sees (Li et al., 2026). Some vision-language models over-predict alignment. Giving them the authentic maps improves accuracy, but it does so by pushing them toward assuming shared understanding. Replacing the images with textual descriptions of the same maps produces the same bias, so the effect comes from task-relevant content rather than from the visual modality. In the models that show this pattern, static referential cues on the map are treated as evidence of agreement, in place of tracking how grounding develops through the dialogue history.
Related work in the project
Perspectivist modeling runs through the wider DeMeVa effort as well: our submission to the LeWiDi-2025 shared task compares in-context learning and label distribution learning as ways of modeling annotator perspectives rather than aggregating them away (Ignatev et al., 2025).
Next steps
The immediate goal is to extend the annotation scheme beyond MapTask to other collaborative dialogue corpora, and to establish which of the current findings survive a change of task. A corpus suitable for replication should record both interlocutors’ interpretations separately rather than task success alone, which few existing resources do. Collecting new dialogue data is a possibility if no existing corpus meets that requirement.
References
2026
- Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric DialogueOralIn Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL), Aug 2026
2025
- DeMeVa at LeWiDi-2025: Modeling Perspectives with In-Context Learning and Label Distribution LearningIn Proceedings of the 4th Workshop on Perspectivist Approaches to NLP (NLPerspectives), Nov 2025