Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication
arXiv:2606.17372v3 Announce Type: replace-cross Abstract: Two recent studies \citep{jones2026llms, zeng2026lvlms} reach apparently contradictory conclusions about whether large vision-language models (LVLMs) can coordinate similarly to humans on efficient referring expressions.
We control for task differences between the studies while directly comparing their prompting styles.
We replicate the finding that models can coordinate efficient referring expressions when \textit{explicitly} prompted to do so, suggesting that other task differences are not responsible for divergent results.
However, we also find that the same models fail to infer the need for communicative efficiency from a more \textit{implicit} prompt, highlighting critical differences between how humans and AI systems communicate.