DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Hannan, Tanveer, Mallios, Dimitrios, Pathak, Parth, Sardari, Faegheh, Seidl, Thomas, Bertasius, Gedas, Fayyaz, Mohsen, Sengupta, Sunando |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
by: Hannan, Tanveer, et al.
Published: (2024)
by: Hannan, Tanveer, et al.
Published: (2024)
RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
by: Hannan, Tanveer, et al.
Published: (2023)
by: Hannan, Tanveer, et al.
Published: (2023)
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
by: Bahrami, Emad, et al.
Published: (2026)
by: Bahrami, Emad, et al.
Published: (2026)
Siamese Vision Transformers are Scalable Audio-visual Learners
by: Lin, Yan-Bo, et al.
Published: (2024)
by: Lin, Yan-Bo, et al.
Published: (2024)
LoCoNet: Long-Short Context Network for Active Speaker Detection
by: Wang, Xizi, et al.
Published: (2023)
by: Wang, Xizi, et al.
Published: (2023)
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
by: Lan, Jian, et al.
Published: (2025)
by: Lan, Jian, et al.
Published: (2025)
BASKET: A Large-Scale Video Dataset for Fine-Grained Skill Estimation
by: Pan, Yulu, et al.
Published: (2025)
by: Pan, Yulu, et al.
Published: (2025)
AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
by: Zhang, Gengyuan, et al.
Published: (2025)
by: Zhang, Gengyuan, et al.
Published: (2025)
DeltaLLM: Compress LLMs with Low-Rank Deltas between Shared Weights
by: Mikaelyan, Liana, et al.
Published: (2025)
by: Mikaelyan, Liana, et al.
Published: (2025)
Unsupervised View-Invariant Human Posture Representation
by: Sardari, Faegheh, et al.
Published: (2021)
by: Sardari, Faegheh, et al.
Published: (2021)
BOSS: Benchmark for Observation Space Shift in Long-Horizon Task
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
by: Islam, Md Mohaiminul, et al.
Published: (2025)
by: Islam, Md Mohaiminul, et al.
Published: (2025)
Augmented Reality Demonstrations for Scalable Robot Imitation Learning
by: Yang, Yue, et al.
Published: (2024)
by: Yang, Yue, et al.
Published: (2024)
ARCADE: Scalable Demonstration Collection and Generation via Augmented Reality for Imitation Learning
by: Yang, Yue, et al.
Published: (2024)
by: Yang, Yue, et al.
Published: (2024)
SiLVR: A Simple Language-based Video Reasoning Framework
by: Zhang, Ce, et al.
Published: (2025)
by: Zhang, Ce, et al.
Published: (2025)
Video ReCap: Recursive Captioning of Hour-Long Videos
by: Islam, Md Mohaiminul, et al.
Published: (2024)
by: Islam, Md Mohaiminul, et al.
Published: (2024)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
by: Deng, Chao, et al.
Published: (2024)
by: Deng, Chao, et al.
Published: (2024)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
by: Wang, Ziyang, et al.
Published: (2026)
by: Wang, Ziyang, et al.
Published: (2026)
TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion
by: Tursynbek, Nurislam, et al.
Published: (2026)
by: Tursynbek, Nurislam, et al.
Published: (2026)
LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies
by: Yang, Yue, et al.
Published: (2026)
by: Yang, Yue, et al.
Published: (2026)
A Simple LLM Framework for Long-Range Video Question-Answering
by: Zhang, Ce, et al.
Published: (2023)
by: Zhang, Ce, et al.
Published: (2023)
An Effective-Efficient Approach for Dense Multi-Label Action Detection
by: Sardari, Faegheh, et al.
Published: (2024)
by: Sardari, Faegheh, et al.
Published: (2024)
Reframing Dense Action Detection (RefDense): A Paradigm Shift in Problem Solving & a Novel Optimization Strategy
by: Sardari, Faegheh, et al.
Published: (2025)
by: Sardari, Faegheh, et al.
Published: (2025)
CoLeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio-Visual Video Parsing
by: Sardari, Faegheh, et al.
Published: (2024)
by: Sardari, Faegheh, et al.
Published: (2024)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
by: Feng, Xiang, et al.
Published: (2026)
by: Feng, Xiang, et al.
Published: (2026)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
by: Ma, Yubo, et al.
Published: (2024)
by: Ma, Yubo, et al.
Published: (2024)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
by: Wang, Ziyang, et al.
Published: (2024)
by: Wang, Ziyang, et al.
Published: (2024)
VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos
by: Lin, Yan-Bo, et al.
Published: (2024)
by: Lin, Yan-Bo, et al.
Published: (2024)
ExAct: A Video-Language Benchmark for Expert Action Analysis
by: Yi, Han, et al.
Published: (2025)
by: Yi, Han, et al.
Published: (2025)
DocAtlas: Multilingual Document Understanding Across 80+ Languages
by: Heakl, Ahmed, et al.
Published: (2026)
by: Heakl, Ahmed, et al.
Published: (2026)
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
by: Zhou, Yiyang, et al.
Published: (2025)
by: Zhou, Yiyang, et al.
Published: (2025)
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
by: Rasekh, Ali, et al.
Published: (2025)
by: Rasekh, Ali, et al.
Published: (2025)
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
by: Yan, Hao, et al.
Published: (2026)
by: Yan, Hao, et al.
Published: (2026)
ContraDoc: Understanding Self-Contradictions in Documents with Large Language Models
by: Li, Jierui, et al.
Published: (2023)
by: Li, Jierui, et al.
Published: (2023)
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
by: Ni, Feng, et al.
Published: (2025)
by: Ni, Feng, et al.
Published: (2025)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
by: Huang, Kui, et al.
Published: (2025)
by: Huang, Kui, et al.
Published: (2025)
Do ESG funds engage in portfolio pumping to gain higher flows? An application of Benford's Law
by: Aineas Mallios, et al.
Published: (2024)
by: Aineas Mallios, et al.
Published: (2024)
SLM as Guardian: Pioneering AI Safety with Small Language Models
by: Kwon, Ohjoon, et al.
Published: (2024)
by: Kwon, Ohjoon, et al.
Published: (2024)
SLM-SQL: An Exploration of Small Language Models for Text-to-SQL
by: Sheng, Lei, et al.
Published: (2025)
by: Sheng, Lei, et al.
Published: (2025)
Similar Items
-
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
by: Hannan, Tanveer, et al.
Published: (2024) -
RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
by: Hannan, Tanveer, et al.
Published: (2023) -
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
by: Bahrami, Emad, et al.
Published: (2026) -
Siamese Vision Transformers are Scalable Audio-visual Learners
by: Lin, Yan-Bo, et al.
Published: (2024) -
LoCoNet: Long-Short Context Network for Active Speaker Detection
by: Wang, Xizi, et al.
Published: (2023)