Enhancing Vision Models for Text-Heavy Content Understanding and Interaction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | TG, Adithya, SK, Adithya, Bharadwaj, Abhinav R, HA, Abhiram, Narayan, Surabhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
IBiT: Utilizing Inductive Biases to Create a More Data Efficient Attention Mechanism
von: Giri, Adithya
Veröffentlicht: (2025)
von: Giri, Adithya
Veröffentlicht: (2025)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy
von: Huang, Yingbing, et al.
Veröffentlicht: (2026)
von: Huang, Yingbing, et al.
Veröffentlicht: (2026)
Can Vision-Language Models Solve Visual Math Equations?
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
von: Tartaglini, Alexa R., et al.
Veröffentlicht: (2025)
von: Tartaglini, Alexa R., et al.
Veröffentlicht: (2025)
Benchmarking Vision Language Models for Cultural Understanding
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
Partial Information Decomposition via Normalizing Flows in Latent Gaussian Distributions
von: Zhao, Wenyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Wenyuan, et al.
Veröffentlicht: (2025)
Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
von: Van, Minh-Hao, et al.
Veröffentlicht: (2025)
von: Van, Minh-Hao, et al.
Veröffentlicht: (2025)
Can Vision Language Models Understand Mimed Actions?
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
von: Yamabe, Shojiro, et al.
Veröffentlicht: (2025)
von: Yamabe, Shojiro, et al.
Veröffentlicht: (2025)
Enhancing Human-Computer Interaction in Chest X-ray Analysis using Vision and Language Model with Eye Gaze Patterns
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
von: Kim, Yunsoo, et al.
Veröffentlicht: (2024)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement
von: Jiao, Siwen, et al.
Veröffentlicht: (2024)
von: Jiao, Siwen, et al.
Veröffentlicht: (2024)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
MULTI: Multimodal Understanding Leaderboard with Text and Images
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models
von: Bhatia, Mehar, et al.
Veröffentlicht: (2024)
von: Bhatia, Mehar, et al.
Veröffentlicht: (2024)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
von: Kim, Mingyeong, et al.
Veröffentlicht: (2026)
von: Kim, Mingyeong, et al.
Veröffentlicht: (2026)
Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
von: Son, Jaemin, et al.
Veröffentlicht: (2025)
von: Son, Jaemin, et al.
Veröffentlicht: (2025)
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
von: Jiang, Jiachen, et al.
Veröffentlicht: (2025)
von: Jiang, Jiachen, et al.
Veröffentlicht: (2025)
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
Causal Graphical Models for Vision-Language Compositional Understanding
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2024)
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2024)
Improving Language Understanding from Screenshots
von: Gao, Tianyu, et al.
Veröffentlicht: (2024)
von: Gao, Tianyu, et al.
Veröffentlicht: (2024)
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries
von: Liu, Shudong, et al.
Veröffentlicht: (2025)
von: Liu, Shudong, et al.
Veröffentlicht: (2025)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
von: De, Anik, et al.
Veröffentlicht: (2025)
von: De, Anik, et al.
Veröffentlicht: (2025)
Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and Summarization
von: Luo, Richard, et al.
Veröffentlicht: (2024)
von: Luo, Richard, et al.
Veröffentlicht: (2024)
How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion Models
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
von: Chen, Chao, et al.
Veröffentlicht: (2025)
von: Chen, Chao, et al.
Veröffentlicht: (2025)
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
von: Zhang, Youliang, et al.
Veröffentlicht: (2026)
von: Zhang, Youliang, et al.
Veröffentlicht: (2026)
Vision-Language Agents for Interactive Forest Change Analysis
von: Brock, James, et al.
Veröffentlicht: (2026)
von: Brock, James, et al.
Veröffentlicht: (2026)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025) -
IBiT: Utilizing Inductive Biases to Create a More Data Efficient Attention Mechanism
von: Giri, Adithya
Veröffentlicht: (2025) -
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025) -
KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy
von: Huang, Yingbing, et al.
Veröffentlicht: (2026) -
Can Vision-Language Models Solve Visual Math Equations?
von: Choudhury, Monjoy Narayan, et al.
Veröffentlicht: (2025)