Skip to content
CategoryEngineering Journal

Document & Multimodal Intelligence

Beyond OCR: layout, tables, vision-language models, figure grounding, and cross-modal evidence.

2 posts