Reading Order Inference for Complex Document Layouts
复杂版面阅读顺序推断新突破,聚焦中世纪手稿多栏交错文本,提升OCR数字化准确率。
arXiv:2607.01018v1 Announce Type: cross Abstract: Reading order inference remains a critical bottleneck in the digitization of complex historical manu…
复杂版面阅读顺序推断新突破,聚焦中世纪手稿多栏交错文本,提升OCR数字化准确率。
arXiv:2607.01018v1 Announce Type: cross Abstract: Reading order inference remains a critical bottleneck in the digitization of complex historical manu…
多模态图RAG+长程视觉文档理解,突破复杂版面解析瓶颈,值得关注。
arXiv:2606.28780v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are widely applied to visual document understanding. Howeve…
长上下文视觉文档模型训练新方法,突破理解复杂文档的瓶颈,值得AI从业者关注。
arXiv:2602.15257v3 Announce Type: replace-cross Abstract: We present the first comprehensive, large-scale study of training long-context vision langua…
探索多模态文档集合中的多跳推理新基准DocHop-QA,推动视觉与文本信息联合理解。
arXiv:2508.15851v2 Announce Type: replace Abstract: Despite rapid progress in large language models (LLMs), current QA benchmarks still overlook the c…
布局导向的细粒度RAG方法,破解多模态文档理解中复杂结构信息的获取难题
arXiv:2605.22829v1 Announce Type: cross Abstract: Multimodal Retrieval-Augmented Generation (RAG) has emerged as an effective paradigm for enhancing L…
多模态大模型在真实收据文档理解上的基准测试与改进,揭示现有模型局限并推进从识别到推理的能力跃升。
arXiv:2605.22413v1 Announce Type: new Abstract: Extracting structured information from visual documents (Visual Information Extraction, VIE) is a corn…
AI因缺乏高质量PDF数据而“饥饿”,揭示大模型训练面临的关键瓶颈与解决思路。
Article URL: https://mkotlikov.substack.com/p/your-ai-is-starving-for-pdfs Comments URL: https://news.ycombinator.com/item?id=48173305 Points: 2 # Com…