Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction
Published in NeurIPS 2025 Multimodal Algorithmic Reasoning Workshop (ICML Under Review), 2025
| arXiv | GitHub |
Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex reasoning tasks requiring deep visual understanding. Inspired by human visual cognition, we propose Visual Abstract Thinking (VAT), a novel reasoning paradigm. VAT enables MLLMs to reason more effectively by leveraging visual abstractions—focusing on semantic and geometric concepts—rather than relying solely on explicit verbal Chain-of-Thought. Results show that VAT achieves a 17% improvement over the GPT-4o baseline and consistently outperforms Chain-of-Thought prompting on visual perception and reasoning tasks, while requiring fewer tokens and delivering higher performance.
Recommended citation: Dairu Liu, Ziyue Wang, Minyuan Ruan, Fuwen Luo, Chi Chen, Peng Li, Yang Liu. “Visual Abstract Thinking Empowers Multimodal Reasoning.” NeurIPS 2025 Multimodal Algorithmic Reasoning Workshop.
