AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Masry, Ahmed, Rodriguez, Juan A., Zhang, Tianyu, Wang, Suyuchen, Wang, Chao, Feizi, Aarash, Suresh, Akshay Kalkunte, Puri, Abhay, Jian, Xiangru, Noël, Pierre-André, Madhusudhan, Sathwik Tejaswi, Pedersoli, Marco, Liu, Bang, Chapados, Nicolas, Bengio, Yoshua, Hoque, Enamul, Pal, Christopher, Laradji, Issam H., Vazquez, David, Taslakian, Perouz, Gella, Spandana, Rajeswar, Sai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026)
by: Gurung, Alexander, et al.
Published: (2026)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025)
by: Wang, Suyuchen, et al.
Published: (2025)
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
by: Rodriguez, Juan A., et al.
Published: (2025)
by: Rodriguez, Juan A., et al.
Published: (2025)
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
by: Awal, Rabiul, et al.
Published: (2025)
by: Awal, Rabiul, et al.
Published: (2025)
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
by: Jian, Xiangru, et al.
Published: (2026)
by: Jian, Xiangru, et al.
Published: (2026)
VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
by: Rodriguez, Juan, et al.
Published: (2024)
by: Rodriguez, Juan, et al.
Published: (2024)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
by: Maheshwary, Rishabh, et al.
Published: (2025)
by: Maheshwary, Rishabh, et al.
Published: (2025)
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
Grounding Computer Use Agents on Human Demonstrations
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025)
by: Bechard, Patrice, et al.
Published: (2025)
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
by: Rodriguez, Juan, et al.
Published: (2026)
by: Rodriguez, Juan, et al.
Published: (2026)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
StarVector: Generating Scalable Vector Graphics Code from Images and Text
by: Rodriguez, Juan A., et al.
Published: (2023)
by: Rodriguez, Juan A., et al.
Published: (2023)
DeepSRGM -- Sequence Classification and Ranking in Indian Classical Music with Deep Learning
by: Madhusudhan, Sathwik Tejaswi, et al.
Published: (2024)
by: Madhusudhan, Sathwik Tejaswi, et al.
Published: (2024)
Mem-$π$: Adaptive Memory through Learning When and What to Generate
by: Wang, Xiaoqiang, et al.
Published: (2026)
by: Wang, Xiaoqiang, et al.
Published: (2026)
Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models
by: Madhusudhan, Nishanth, et al.
Published: (2024)
by: Madhusudhan, Nishanth, et al.
Published: (2024)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
by: Monteiro, Joao, et al.
Published: (2024)
by: Monteiro, Joao, et al.
Published: (2024)
AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
by: Nguyen, Hoang, et al.
Published: (2025)
by: Nguyen, Hoang, et al.
Published: (2025)
Grammar Search for Multi-Agent Systems
by: Singh, Mayank, et al.
Published: (2025)
by: Singh, Mayank, et al.
Published: (2025)
AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery
by: Abaskohi, Amirhossein, et al.
Published: (2025)
by: Abaskohi, Amirhossein, et al.
Published: (2025)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
by: Pattnaik, Pulkit, et al.
Published: (2024)
by: Pattnaik, Pulkit, et al.
Published: (2024)
M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models
by: Maheshwary, Rishabh, et al.
Published: (2024)
by: Maheshwary, Rishabh, et al.
Published: (2024)
ChartInstruct: Instruction Tuning for Chart Comprehension and Reasoning
by: Masry, Ahmed, et al.
Published: (2024)
by: Masry, Ahmed, et al.
Published: (2024)
IntentGPT: Few-shot Intent Discovery with Large Language Models
by: Rodriguez, Juan A., et al.
Published: (2024)
by: Rodriguez, Juan A., et al.
Published: (2024)
Revitalizing Saturated Benchmarks: A Weighted Metric Approach for Differentiating Large Language Model Performance
by: Etzine, Bryan, et al.
Published: (2025)
by: Etzine, Bryan, et al.
Published: (2025)
DRBench: A Realistic Benchmark for Enterprise Deep Research
by: Abaskohi, Amirhossein, et al.
Published: (2025)
by: Abaskohi, Amirhossein, et al.
Published: (2025)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
by: Masry, Ahmed, et al.
Published: (2024)
by: Masry, Ahmed, et al.
Published: (2024)
LitLLM: A Toolkit for Scientific Literature Review
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
LitLLMs, LLMs for Literature Review: Are we there yet?
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
by: Tiwari, Aman, et al.
Published: (2024)
by: Tiwari, Aman, et al.
Published: (2024)
A Guide To Effectively Leveraging LLMs for Low-Resource Text Summarization: Data Augmentation and Semi-supervised Approaches
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
Prompt-based Pseudo-labeling Strategy for Sample-Efficient Semi-Supervised Extractive Summarization
by: Sahu, Gaurav, et al.
Published: (2023)
by: Sahu, Gaurav, et al.
Published: (2023)
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
Similar Items
-
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
by: Masry, Ahmed, et al.
Published: (2025) -
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
by: Masry, Ahmed, et al.
Published: (2025) -
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026) -
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025) -
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
by: Zhang, Tianyu, et al.
Published: (2025)