← all papers · overview

Flipvqa: Scaling Multi-modal Instruction Tuning Via Textbook-to-knowledge Synthesis

Abstract

Textbooks are among the richest repositories of human-verified reasoning knowledge, yet their complex layouts contain multi-column typesetting, cross-page question answer separation, and interleaved figures, make automated extraction of structured QA and VQA pairs extremely challenging. Existing alternatives either synthesize data from scratch, which lacks authentic problem contexts, or rely on costly expert annotation that cannot scale. We propose , an automated pipeline that resolves long-range logical dependencies and cross-page discontinuities in OCR-parsed documents, recovering coherent question--answer--figure associations even when answers reside in separate companion volumes. A subsequent multi-stage curation pipeline transforms these raw extractions into AI-ready supervision signals. Using FlipVQA-Miner, we construct , comprising 83K QA and VQA pairs spanning 11 academic disciplines, at a \times cost sa

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).