Multimodal Evidence Systems: VLMs, Figure Grounding, and Cross-Modal Retrieval
Text-only systems break as soon as the evidence stops being mostly text, which is exactly what happens in accessible travel planning when photos, floor plans, route maps, captions, and measurements all shape the answer.
Huang Tzu Lin