Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hamed, Omar, Bakkali, Souhail, Moens, Marie-Francine, Blaschko, Matthew, Van Landeghem, Jordy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
von: Van Landeghem, Jordy, et al.
Veröffentlicht: (2024)
von: Van Landeghem, Jordy, et al.
Veröffentlicht: (2024)
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
von: Chau, Phan Phuong Mai, et al.
Veröffentlicht: (2024)
von: Chau, Phan Phuong Mai, et al.
Veröffentlicht: (2024)
Hybrid Retrieval-Augmented Generation for Robust Multilingual Document Question Answering
von: Mudet, Anthony, et al.
Veröffentlicht: (2025)
von: Mudet, Anthony, et al.
Veröffentlicht: (2025)
GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
von: Bakkali, Souhail, et al.
Veröffentlicht: (2023)
von: Bakkali, Souhail, et al.
Veröffentlicht: (2023)
Visually-Aware Context Modeling for News Image Captioning
von: Qu, Tingyu, et al.
Veröffentlicht: (2023)
von: Qu, Tingyu, et al.
Veröffentlicht: (2023)
DM-Align: Leveraging the Power of Natural Language Instructions to Make Changes to Images
von: Trusca, Maria Mihaela, et al.
Veröffentlicht: (2024)
von: Trusca, Maria Mihaela, et al.
Veröffentlicht: (2024)
Animate Your Motion: Turning Still Images into Dynamic Videos
von: Li, Mingxiao, et al.
Veröffentlicht: (2024)
von: Li, Mingxiao, et al.
Veröffentlicht: (2024)
Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias
von: Li, Mingxiao, et al.
Veröffentlicht: (2025)
von: Li, Mingxiao, et al.
Veröffentlicht: (2025)
Consistent Story Generation: Unlocking the Potential of Zigzag Sampling
von: Li, Mingxiao, et al.
Veröffentlicht: (2025)
von: Li, Mingxiao, et al.
Veröffentlicht: (2025)
Introducing Routing Functions to Vision-Language Parameter-Efficient Fine-Tuning with Low-Rank Bottlenecks
von: Qu, Tingyu, et al.
Veröffentlicht: (2024)
von: Qu, Tingyu, et al.
Veröffentlicht: (2024)
IDTrust: Deep Identity Document Quality Detection with Bandpass Filtering
von: Al-Ghadi, Musab, et al.
Veröffentlicht: (2024)
von: Al-Ghadi, Musab, et al.
Veröffentlicht: (2024)
Action-based image editing guided by human instructions
von: Trusca, Maria Mihaela, et al.
Veröffentlicht: (2024)
von: Trusca, Maria Mihaela, et al.
Veröffentlicht: (2024)
Evaluating the Impact of Khmer Font Types on Text Recognition
von: Nom, Vannkinh, et al.
Veröffentlicht: (2025)
von: Nom, Vannkinh, et al.
Veröffentlicht: (2025)
KhmerST: A Low-Resource Khmer Scene Text Detection and Recognition Benchmark
von: Nom, Vannkinh, et al.
Veröffentlicht: (2024)
von: Nom, Vannkinh, et al.
Veröffentlicht: (2024)
NeuroCine: Decoding Vivid Video Sequences from Human Brain Activties
von: Sun, Jingyuan, et al.
Veröffentlicht: (2024)
von: Sun, Jingyuan, et al.
Veröffentlicht: (2024)
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
von: Qu, Tingyu, et al.
Veröffentlicht: (2024)
von: Qu, Tingyu, et al.
Veröffentlicht: (2024)
NAB: Neural Adaptive Binning for Sparse-View CT reconstruction
von: Xie, Wangduo, et al.
Veröffentlicht: (2026)
von: Xie, Wangduo, et al.
Veröffentlicht: (2026)
WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
von: Jiang, Yiwen, et al.
Veröffentlicht: (2025)
von: Jiang, Yiwen, et al.
Veröffentlicht: (2025)
"Image, Tell me your story!" Predicting the original meta-context of visual misinformation
von: Tonglet, Jonathan, et al.
Veröffentlicht: (2024)
von: Tonglet, Jonathan, et al.
Veröffentlicht: (2024)
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
von: Meng, Rui, et al.
Veröffentlicht: (2025)
von: Meng, Rui, et al.
Veröffentlicht: (2025)
Alleviating Exposure Bias in Diffusion Models through Sampling with Shifted Time Steps
von: Li, Mingxiao, et al.
Veröffentlicht: (2023)
von: Li, Mingxiao, et al.
Veröffentlicht: (2023)
Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis
von: Jelaca, Aleksa, et al.
Veröffentlicht: (2025)
von: Jelaca, Aleksa, et al.
Veröffentlicht: (2025)
Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
von: Wang, Zifu, et al.
Veröffentlicht: (2025)
von: Wang, Zifu, et al.
Veröffentlicht: (2025)
AdaRevD: Adaptive Patch Exiting Reversible Decoder Pushes the Limit of Image Deblurring
von: Mao, Xintian, et al.
Veröffentlicht: (2024)
von: Mao, Xintian, et al.
Veröffentlicht: (2024)
A Simple Early Exiting Framework for Accelerated Sampling in Diffusion Models
von: Moon, Taehong, et al.
Veröffentlicht: (2024)
von: Moon, Taehong, et al.
Veröffentlicht: (2024)
Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control
von: Trusca, Maria Mihaela, et al.
Veröffentlicht: (2024)
von: Trusca, Maria Mihaela, et al.
Veröffentlicht: (2024)
Multimodal Language Models Cannot Spot Spatial Inconsistencies
von: Khangaonkar, Om, et al.
Veröffentlicht: (2026)
von: Khangaonkar, Om, et al.
Veröffentlicht: (2026)
Docopilot: Improving Multimodal Models for Document-Level Understanding
von: Duan, Yuchen, et al.
Veröffentlicht: (2025)
von: Duan, Yuchen, et al.
Veröffentlicht: (2025)
From Sight to Insight: Improving Visual Reasoning Capabilities of Multimodal Models via Reinforcement Learning
von: Sharif, Omar, et al.
Veröffentlicht: (2026)
von: Sharif, Omar, et al.
Veröffentlicht: (2026)
Efficient Information Extraction in Few-Shot Relation Classification through Contrastive Representation Learning
von: Borchert, Philipp, et al.
Veröffentlicht: (2024)
von: Borchert, Philipp, et al.
Veröffentlicht: (2024)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
von: Hennara, Khalil, et al.
Veröffentlicht: (2025)
von: Hennara, Khalil, et al.
Veröffentlicht: (2025)
MSNER: A Multilingual Speech Dataset for Named Entity Recognition
von: Meeus, Quentin, et al.
Veröffentlicht: (2024)
von: Meeus, Quentin, et al.
Veröffentlicht: (2024)
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
von: Constantinou, Christos, et al.
Veröffentlicht: (2024)
von: Constantinou, Christos, et al.
Veröffentlicht: (2024)
Multimodal Human-AI Synergy for Medical Imaging Quality Control: A Hybrid Intelligence Framework with Adaptive Dataset Curation and Closed-Loop Evaluation
von: Qin, Zhi, et al.
Veröffentlicht: (2025)
von: Qin, Zhi, et al.
Veröffentlicht: (2025)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
von: Peng, Xiangyu, et al.
Veröffentlicht: (2025)
von: Peng, Xiangyu, et al.
Veröffentlicht: (2025)
When normalization hallucinates: unseen risks in AI-powered whole slide image processing
von: Moens, Karel, et al.
Veröffentlicht: (2025)
von: Moens, Karel, et al.
Veröffentlicht: (2025)
Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
von: Xiao, Han, et al.
Veröffentlicht: (2025)
von: Xiao, Han, et al.
Veröffentlicht: (2025)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
von: Guo, Zichun, et al.
Veröffentlicht: (2026)
von: Guo, Zichun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
von: Van Landeghem, Jordy, et al.
Veröffentlicht: (2024) -
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization
von: Chau, Phan Phuong Mai, et al.
Veröffentlicht: (2024) -
Hybrid Retrieval-Augmented Generation for Robust Multilingual Document Question Answering
von: Mudet, Anthony, et al.
Veröffentlicht: (2025) -
GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
von: Bakkali, Souhail, et al.
Veröffentlicht: (2023) -
Visually-Aware Context Modeling for News Image Captioning
von: Qu, Tingyu, et al.
Veröffentlicht: (2023)