Unveiling the Visual Counting Bottleneck in Vision-Language Models
Proceedings of the 43rd International Conference on Machine Learning (ICML 2026), 2026
We investigate why vision-language models struggle to count beyond training distributions, tracing failures to the mapping from visual quantities to symbolic numbers.
Recommended citation: Xingzhou Pang, Yifan Hou, Junling Wang, Mrinmaya Sachan. (2026). "Unveiling the Visual Counting Bottleneck in Vision-Language Models." Proceedings of the 43rd International Conference on Machine Learning (ICML 2026). https://arxiv.org/abs/2605.30170
