Unveiling the Visual Counting Bottleneck in Vision-Language Models

Proceedings of the 43rd International Conference on Machine Learning (ICML 2026), 2026

We investigate why vision-language models struggle to count beyond training distributions, tracing failures to the mapping from visual quantities to symbolic numbers.

Recommended citation: Xingzhou Pang, Yifan Hou, Junling Wang, Mrinmaya Sachan. (2026). "Unveiling the Visual Counting Bottleneck in Vision-Language Models." Proceedings of the 43rd International Conference on Machine Learning (ICML 2026). https://arxiv.org/abs/2605.30170