LREC 2026 SIGDIAL 2026

Tracking Common Ground
in MapTask

Who means what—and do they mean the same thing?
A perspectivist view of reference and misunderstanding.

GMMT · participant interpretations / SINS · vision–language model evaluation

Nan Li · Albert Gatt · Massimo Poesio

NLP Group, Department of Information and Computing Sciences, Utrecht University

Abstract

Understanding in dialogue is shaped by what each participant knows, sees, and takes an expression to mean. When their information differs, the same words—and even an acknowledgement—can conceal different referents. GMMT and SINS bring together two complementary studies of this problem: a perspectivist annotation resource that preserves speaker intention and addressee interpretation, and a model evaluation built on those paired perspectives. Together, they distinguish what could be shared from what participants have actually established through interaction.

GMMT (LREC 2026) captures how reference is understood from each participant's perspective. Its annotation scheme records the speaker's intended landmark and the addressee's interpretation, alongside five attributes describing reference and grounding. A scheme-constrained GPT-5 pipeline produces 13,077 annotated expressions across 128 MapTask dialogues, with a human-verified subset for reliability assessment. These annotations distinguish alignment, pending understanding, and silent misunderstanding; reference chains trace how understanding emerges, diverges, and is repaired. After lexical variants are unified, misunderstandings are uncommon overall, but repeated, identically named landmarks remain a recurring source of misalignment.

SINS (SIGDIAL 2026) builds on GMMT to ask whether vision–language models can judge when the two participants' interpretations match. We systematically vary dialogue context and access to map information. The clearest failure mode, observed in Qwen3-VL-8B-Instruct, concerns how evidence is used: analyses suggest that the model over-weights static map information about what could be shared and under-weights dialogue evidence about what participants have actually established together. Authentic maps can improve aggregate performance while encouraging premature judgments of alignment; textual map descriptions reproduce this tendency. Cross-model differences show that this bias is model-dependent, not universal.

13,077reference expressions
128dialogues
16map pairs
2participant interpretations

Three understanding states

Each reference expression receives an understanding state, combining the participants' inferred referents with evidence of grounding in the dialogue. States describe individual references—not whole maps or dialogues. Following successive mentions reveals how understanding is established, remains unresolved, or diverges.

Aligned9,435 72.1%

Both participants ground the expression to the same or equivalent landmark.

Pending3,403 26.0%

Grounding is not established in the annotation cascade. This includes existence checks, missing uptake, clarification, and unresolved references—not necessarily a misunderstanding.

Misunderstood239 1.8%

Both participants ground the expression, but to different landmarks. Apparent agreement can conceal divergent interpretations.

Counts from the 13,077 released GPT-5 annotations, after unifying cross-map lexical variants. Percentages are rounded; these are not full-corpus human-verified counts.

Interactive demo

One expression, two situated interpretations

Click an underlined phrase in the dialogue. The coloured map outlines and the two interpretation cards update for that exact reference. The opening parked-van case is a real exchange (q8nc6), not the papers' simplified dialogue.

1 Choose a dialogue2 Click a phrase in context3 Compare the two interpretations
Interpretation annotation
Map pair m9

Real participant maps

HCRC MapTask · actual maps

Click a landmark's icon or name to find its mentions. Click the map background to enlarge.

Giver interpretation Follower interpretation Selected landmark

Find mentions of a landmark
Selected reference
—

Participant interpretations

Giver interpretation —
≠
Follower interpretation —
Inspect annotation fields and source rationale

Rationale text is retained from the selected source, not independently verified by this website. Null means the field was not evaluated in the annotation cascade.

Separate task: does a SINS model judge these interpretations to match?

Interpretation matching · Yes / No

Qwen3-VL-8B-Instruct · dialogue start through current transaction
SINS dataset target—
Dialogue only—
Dialogue + both maps—

These are matching judgments, not landmark predictions. The SINS target comes from the released GPT-5 annotation layer and does not change when you inspect a human-verified annotation above. A dash means no saved prediction is available.

Follow the whole conversation

All 128 dialogues · 16 map pairs · 13,077 reference annotations
GPT-5 throughout; human-verified annotations for 3 dialogues (504 references).

Browse all dialogues & annotations

Source and attribution. Maps and dialogue text are from the HCRC MapTask Corpus, displayed under CC BY 4.0 as stated on the official download page. Maps were converted to PNG and resized; transcripts were reconstructed from timed units and reformatted. Interactive overlays are project additions. Participant interpretations come from GMMT and matching judgments from SINS.

Enlarged map

Click the map again, click outside, or press Esc to close.

SINS · SIGDIAL 2026

Potential common ground is not established common ground

The interesting failure is how the model weighs evidence—not simply whether adding maps improves its score.

Over-weighted

What the maps make available

Landmark names, locations, and co-presence offer static clues to what the participants could be referring to.

Potential referential overlap
Under-weighted

What the dialogue establishes

Uptake, confirmation, clarification, and repair reveal whether the participants have grounded a reference—and in which landmark.

Interactional, incremental grounding

This is a behavioural interpretation of the controlled experiments, not a direct measurement of attention weights. The clearest evidence is from Qwen3-VL-8B-Instruct.

The hidden trade-off

Better at saying “same”.
Worse at noticing “not yet”.

With both maps, macro-F1 improves from .591 to .671. But this gain conceals a loss in detecting non-alignment, especially pending cases. The smaller drop on misunderstood cases alone is not statistically significant for both-maps; single-map conditions show significant drops.

01 / Content, not just vision

Describing a map has a similar effect.

Textual landmark descriptions reproduce the alignment shift. Blank and shuffled maps instead push the model toward “No”. Task-relevant content, not the mere presence of an image, drives the pattern.

02 / Confidence, not only answer bias

More certain—even when agreement is absent.

With both maps, calibration error on non-aligned cases rises from .235 to .403, while it improves on aligned cases. The model makes confident errors precisely where grounding is missing or divergent.

03 / History, not just repeated names

Mentioning it again is not confirming it.

Alignment predictions rise with repeated mentions, though longer chains also contain more aligned cases. Broader dialogue context helps less when maps are present: +.078 macro-F1, versus +.152 without maps (curL → startT). Together, these patterns are consistent with an over-reliance on static cues.

Scope matters: this is an overhearer task, not an interactive agent. The bias is not universal across VLMs; Qwen3-VL-4B shifts in the opposite direction. A higher “Yes” rate alone is not proof of over-alignment—the recall trade-off and class-specific calibration are crucial. Read the evidence and limitations in SINS ↗

Open resources

Two connected datasets

GMMT supplies participant-level interpretations; SINS converts them into controlled common-ground judgments for multimodal models.

GMMT
Annotation resource · LREC 2026

Grounded Misunderstandings in MapTask

Separate giver and follower landmark resolutions for 13,077 referring expressions, including aligned, pending, and misunderstood states.

Unit
Reference expression
Views
Giver + follower
License
CC BY 4.0
SINS
Model challenge · SIGDIAL 2026

Seeing Is Not Sharing

A binary interpretation-matching task with controlled dialogue windows and map-access conditions for vision–language models.

Gold
Yes / No
Factors
Text × map access
Data license
CC BY 4.0
Publications

Read and cite the work

LREC
2026

Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask

Nan Li, Albert Gatt, Massimo Poesio

SIGDIAL
2026

Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

Nan Li, Albert Gatt, Massimo Poesio

LREC 2026

Cite GMMT

@inproceedings{li2026grounded,
  title = {Grounded Misunderstandings in Asymmetric Dialogue:
    A Perspectivist Annotation Scheme for {MapTask}},
  author = {Li, Nan and Gatt, Albert and Poesio, Massimo},
  booktitle = {Proceedings of the Fifteenth Language Resources
    and Evaluation Conference (LREC 2026)},
  month = {May},
  year = {2026},
  pages = {4988--5001},
  publisher = {European Language Resources Association (ELRA)},
  doi = {10.63317/59anbt78wyj7},
  url = {https://lrec.elra.info/lrec2026-main-392}
}
SIGDIAL 2026

Cite SINS

@inproceedings{li2026seeing,
  title = {Seeing Is Not Sharing: Some Vision-Language Models
    Overestimate Common Ground in Asymmetric Dialogue},
  author = {Li, Nan and Gatt, Albert and Poesio, Massimo},
  editor = {Choi, Jinho D. and Chen, Yun-Nung and Funakoshi,
    Kotaro and Emami, Ali},
  booktitle = {Proceedings of the 27th Annual Meeting of the
    Special Interest Group on Discourse and Dialogue},
  month = aug,
  year = {2026},
  pages = {694--710},
  publisher = {Association for Computational Linguistics},
  url = {https://aclanthology.org/2026.sigdial-1.49/}
}