Improving MLLM Historical Record Extraction with Test-Time Image
Fuente:
arXiv
Saved in:
| Main Authors: | Archibald, Taylor, Martinez, Tony |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DELINE8K: A Synthetic Data Pipeline for the Semantic Segmentation of Historical Documents
by: Archibald, Taylor, et al.
Published: (2024)
by: Archibald, Taylor, et al.
Published: (2024)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
by: Xu, Binqian, et al.
Published: (2024)
by: Xu, Binqian, et al.
Published: (2024)
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2025)
by: Xiao, Han, et al.
Published: (2025)
Alt-Text with Context: Improving Accessibility for Images on Twitter
by: Srivatsan, Nikita, et al.
Published: (2023)
by: Srivatsan, Nikita, et al.
Published: (2023)
MLLM-as-a-Judge for Image Safety without Human Labeling
by: Wang, Zhenting, et al.
Published: (2024)
by: Wang, Zhenting, et al.
Published: (2024)
Leveraging Semantic Segmentation Masks with Embeddings for Fine-Grained Form Classification
by: Archibald, Taylor, et al.
Published: (2024)
by: Archibald, Taylor, et al.
Published: (2024)
Test-Time Training Done Right
by: Zhang, Tianyuan, et al.
Published: (2025)
by: Zhang, Tianyuan, et al.
Published: (2025)
Improve Academic Query Resolution through BERT-based Question Extraction from Images
by: Kamal, Nidhi, et al.
Published: (2024)
by: Kamal, Nidhi, et al.
Published: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
by: Yu, Eric Yang, et al.
Published: (2024)
by: Yu, Eric Yang, et al.
Published: (2024)
HATFormer: Historic Handwritten Arabic Text Recognition with Transformers
by: Chan, Adrian, et al.
Published: (2024)
by: Chan, Adrian, et al.
Published: (2024)
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
Model Interpretability and Rationale Extraction by Input Mask Optimization
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
Robustness of Structured Data Extraction from Perspectively Distorted Documents
by: Nakada, Hyakka, et al.
Published: (2025)
by: Nakada, Hyakka, et al.
Published: (2025)
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
by: Hu, Xueyu, et al.
Published: (2025)
by: Hu, Xueyu, et al.
Published: (2025)
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
by: Shen, Huawen, et al.
Published: (2024)
by: Shen, Huawen, et al.
Published: (2024)
RealKIE: Five Novel Datasets for Enterprise Key Information Extraction
by: Townsend, Benjamin, et al.
Published: (2024)
by: Townsend, Benjamin, et al.
Published: (2024)
MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
by: Park, Jonggwon, et al.
Published: (2025)
by: Park, Jonggwon, et al.
Published: (2025)
Improving Language Understanding from Screenshots
by: Gao, Tianyu, et al.
Published: (2024)
by: Gao, Tianyu, et al.
Published: (2024)
Optical Context Compression Is Just (Bad) Autoencoding
by: Lee, Ivan Yee, et al.
Published: (2025)
by: Lee, Ivan Yee, et al.
Published: (2025)
Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
by: Zhu, Yinglun, et al.
Published: (2025)
by: Zhu, Yinglun, et al.
Published: (2025)
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
by: Duan, Chengqi, et al.
Published: (2025)
by: Duan, Chengqi, et al.
Published: (2025)
Improving Resnet-9 Generalization Trained on Small Datasets
by: Awad, Omar Mohamed, et al.
Published: (2023)
by: Awad, Omar Mohamed, et al.
Published: (2023)
MOFI: Learning Image Representations from Noisy Entity Annotated Images
by: Wu, Wentao, et al.
Published: (2023)
by: Wu, Wentao, et al.
Published: (2023)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
by: Zhou, Yu, et al.
Published: (2025)
by: Zhou, Yu, et al.
Published: (2025)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
by: Maini, Pratyush, et al.
Published: (2023)
by: Maini, Pratyush, et al.
Published: (2023)
Improving Multimodal Large Language Models Using Continual Learning
by: Srivastava, Shikhar, et al.
Published: (2024)
by: Srivastava, Shikhar, et al.
Published: (2024)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
by: Chen, Yangyi, et al.
Published: (2023)
by: Chen, Yangyi, et al.
Published: (2023)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
by: Villegas, Danae Sánchez, et al.
Published: (2025)
by: Villegas, Danae Sánchez, et al.
Published: (2025)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
by: Lei, Jiayi, et al.
Published: (2025)
by: Lei, Jiayi, et al.
Published: (2025)
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
by: Ghosh, Adhiraj, et al.
Published: (2024)
by: Ghosh, Adhiraj, et al.
Published: (2024)
SimpleGPT: Improving GPT via A Simple Normalization Strategy
by: Chen, Marco, et al.
Published: (2026)
by: Chen, Marco, et al.
Published: (2026)
Dual-Process Image Generation
by: Luo, Grace, et al.
Published: (2025)
by: Luo, Grace, et al.
Published: (2025)
Deep Image-to-Recipe Translation
by: Ma, Jiangqin, et al.
Published: (2024)
by: Ma, Jiangqin, et al.
Published: (2024)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)
by: Heyward, Joseph, et al.
Published: (2024)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
by: Simoncini, Walter, et al.
Published: (2024)
by: Simoncini, Walter, et al.
Published: (2024)
GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models
by: Jin, Haibo, et al.
Published: (2024)
by: Jin, Haibo, et al.
Published: (2024)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
by: Gan, Woody Haosheng, et al.
Published: (2025)
by: Gan, Woody Haosheng, et al.
Published: (2025)
C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
by: Chen, Xiuwei, et al.
Published: (2025)
by: Chen, Xiuwei, et al.
Published: (2025)
Similar Items
-
DELINE8K: A Synthetic Data Pipeline for the Semantic Segmentation of Historical Documents
by: Archibald, Taylor, et al.
Published: (2024) -
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
by: Xu, Binqian, et al.
Published: (2024) -
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2025) -
Alt-Text with Context: Improving Accessibility for Images on Twitter
by: Srivatsan, Nikita, et al.
Published: (2023) -
MLLM-as-a-Judge for Image Safety without Human Labeling
by: Wang, Zhenting, et al.
Published: (2024)