Both participants ground the expression to the same or equivalent landmark.
LREC 2026 + SIGDIAL 2026
Tracking Common Ground
in MapTask
Who means what—and do they mean the same thing?
A perspectivist view of reference and misunderstanding.
GMMT · participant interpretations / SINS · vision–language model evaluation
NLP Group, Department of Information and Computing Sciences, Utrecht University
The same expression, different interpretations.
What differs between their views?
The giver describes a route. The follower draws it. Neither can see the other's map.
Click either map to enlarge; click again to close. Authentic HCRC maps; coloured outlines are explanatory overlays, not annotation predictions. A map discrepancy creates a challenge for reference—it does not by itself establish a misunderstanding.
Abstract
Understanding in dialogue is shaped by what each participant knows, sees, and takes an expression to mean. When their information differs, the same words—and even an acknowledgement—can conceal different referents. GMMT and SINS bring together two complementary studies of this problem: a perspectivist annotation resource that preserves speaker intention and addressee interpretation, and a model evaluation built on those paired perspectives. Together, they distinguish what could be shared from what participants have actually established through interaction.
GMMT (LREC 2026) captures how reference is understood from each participant's perspective. Its annotation scheme records the speaker's intended landmark and the addressee's interpretation, alongside five attributes describing reference and grounding. A scheme-constrained GPT-5 pipeline produces 13,077 annotated expressions across 128 MapTask dialogues, with a human-verified subset for reliability assessment. These annotations distinguish alignment, pending understanding, and silent misunderstanding; reference chains trace how understanding emerges, diverges, and is repaired. After lexical variants are unified, misunderstandings are uncommon overall, but repeated, identically named landmarks remain a recurring source of misalignment.
SINS (SIGDIAL 2026) builds on GMMT to ask whether vision–language models can judge when the two participants' interpretations match. We systematically vary dialogue context and access to map information. The clearest failure mode, observed in Qwen3-VL-8B-Instruct, concerns how evidence is used: analyses suggest that the model over-weights static map information about what could be shared and under-weights dialogue evidence about what participants have actually established together. Authentic maps can improve aggregate performance while encouraging premature judgments of alignment; textual map descriptions reproduce this tendency. Cross-model differences show that this bias is model-dependent, not universal.
Three understanding states
Each reference expression receives an understanding state, combining the participants' inferred referents with evidence of grounding in the dialogue. States describe individual references—not whole maps or dialogues. Following successive mentions reveals how understanding is established, remains unresolved, or diverges.
Grounding is not established in the annotation cascade. This includes existence checks, missing uptake, clarification, and unresolved references—not necessarily a misunderstanding.
Both participants ground the expression, but to different landmarks. Apparent agreement can conceal divergent interpretations.
Counts from the 13,077 released GPT-5 annotations, after unifying cross-map lexical variants. Percentages are rounded; these are not full-corpus human-verified counts.
One expression, two situated interpretations
Click an underlined phrase in the dialogue. The coloured map outlines and the two interpretation cards update for that exact reference. The opening parked-van case is a real exchange (q8nc6), not the papers' simplified dialogue.
Real participant maps
Click a landmark's icon or name to find its mentions. Click the map background to enlarge.
Find mentions of a landmark
—
Participant interpretations
Inspect annotation fields and source rationale
Rationale text is retained from the selected source, not independently verified by this website. Null means the field was not evaluated in the annotation cascade.
Separate task: does a SINS model judge these interpretations to match?
Interpretation matching · Yes / No
Qwen3-VL-8B-Instruct · dialogue start through current transactionThese are matching judgments, not landmark predictions. The SINS target comes from the released GPT-5 annotation layer and does not change when you inspect a human-verified annotation above. A dash means no saved prediction is available.
Follow the whole conversation
All 128 dialogues · 16 map pairs · 13,077 reference annotations
GPT-5 throughout; human-verified annotations for 3 dialogues (504 references).
Source and attribution. Maps and dialogue text are from the HCRC MapTask Corpus, displayed under CC BY 4.0 as stated on the official download page. Maps were converted to PNG and resized; transcripts were reconstructed from timed units and reformatted. Interactive overlays are project additions. Participant interpretations come from GMMT and matching judgments from SINS.
Potential common ground is not established common ground
The interesting failure is how the model weighs evidence—not simply whether adding maps improves its score.
What the maps make available
Landmark names, locations, and co-presence offer static clues to what the participants could be referring to.
Potential referential overlapWhat the dialogue establishes
Uptake, confirmation, clarification, and repair reveal whether the participants have grounded a reference—and in which landmark.
Interactional, incremental groundingThis is a behavioural interpretation of the controlled experiments, not a direct measurement of attention weights. The clearest evidence is from Qwen3-VL-8B-Instruct.
The hidden trade-off
Better at saying “same”.
Worse at noticing “not yet”.
With both maps, macro-F1 improves from .591 to .671. But this gain conceals a loss in detecting non-alignment, especially pending cases. The smaller drop on misunderstood cases alone is not statistically significant for both-maps; single-map conditions show significant drops.
Recall within each target class · same model and dialogue window (startT) · n = 13,077
Describing a map has a similar effect.
Textual landmark descriptions reproduce the alignment shift. Blank and shuffled maps instead push the model toward “No”. Task-relevant content, not the mere presence of an image, drives the pattern.
More certain—even when agreement is absent.
With both maps, calibration error on non-aligned cases rises from .235 to .403, while it improves on aligned cases. The model makes confident errors precisely where grounding is missing or divergent.
Mentioning it again is not confirming it.
Alignment predictions rise with repeated mentions, though longer chains also contain more aligned cases. Broader dialogue context helps less when maps are present: +.078 macro-F1, versus +.152 without maps (curL → startT). Together, these patterns are consistent with an over-reliance on static cues.
Scope matters: this is an overhearer task, not an interactive agent. The bias is not universal across VLMs; Qwen3-VL-4B shifts in the opposite direction. A higher “Yes” rate alone is not proof of over-alignment—the recall trade-off and class-specific calibration are crucial. Read the evidence and limitations in SINS ↗
Two connected datasets
GMMT supplies participant-level interpretations; SINS converts them into controlled common-ground judgments for multimodal models.
Grounded Misunderstandings in MapTask
Separate giver and follower landmark resolutions for 13,077 referring expressions, including aligned, pending, and misunderstood states.
- Unit
- Reference expression
- Views
- Giver + follower
- License
- CC BY 4.0
Seeing Is Not Sharing
A binary interpretation-matching task with controlled dialogue windows and map-access conditions for vision–language models.
- Gold
- Yes / No
- Factors
- Text × map access
- Data license
- CC BY 4.0
Read and cite the work
2026
Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
Nan Li, Albert Gatt, Massimo Poesio
2026
Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
Nan Li, Albert Gatt, Massimo Poesio
Cite GMMT
@inproceedings{li2026grounded,
title = {Grounded Misunderstandings in Asymmetric Dialogue:
A Perspectivist Annotation Scheme for {MapTask}},
author = {Li, Nan and Gatt, Albert and Poesio, Massimo},
booktitle = {Proceedings of the Fifteenth Language Resources
and Evaluation Conference (LREC 2026)},
month = {May},
year = {2026},
pages = {4988--5001},
publisher = {European Language Resources Association (ELRA)},
doi = {10.63317/59anbt78wyj7},
url = {https://lrec.elra.info/lrec2026-main-392}
}
Cite SINS
@inproceedings{li2026seeing,
title = {Seeing Is Not Sharing: Some Vision-Language Models
Overestimate Common Ground in Asymmetric Dialogue},
author = {Li, Nan and Gatt, Albert and Poesio, Massimo},
editor = {Choi, Jinho D. and Chen, Yun-Nung and Funakoshi,
Kotaro and Emami, Ali},
booktitle = {Proceedings of the 27th Annual Meeting of the
Special Interest Group on Discourse and Dialogue},
month = aug,
year = {2026},
pages = {694--710},
publisher = {Association for Computational Linguistics},
url = {https://aclanthology.org/2026.sigdial-1.49/}
}