TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding
Abstract
Lay Summary
Today’s AI systems that understand both images and text can solve problems step-by-step, but they have a big flaw: once they “look” at an image at the very start, they can’t look back during reasoning—they just think blindly using text. This makes them bad at complex visual tasks like geometry or science problems that need checking details multiple times. This paper introduces TVI-CoT, a new way to make AI reason more like humans: it uses special simple commands, i.e.,Think, Look, Answer, to switch back and forth between thinking in text and checking the image as needed. When stuck, the AI pauses to “Look” at the right part of the image, gathers new visual clues, and then keeps “Thinking” until it’s ready to “Answer.” Tests show this method makes AI much better at tricky image-based questions, especially math and science problems with diagrams, while staying efficient and easy to understand.