ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Masry, Ahmed, Thakkar, Megh, Bechard, Patrice, Madhusudhan, Sathwik Tejaswi, Awal, Rabiul, Mishra, Shambhavi, Suresh, Akshay Kalkunte, Daruru, Srivatsava, Hoque, Enamul, Gella, Spandana, Scholak, Torsten, Rajeswar, Sai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval
von: Sheikholeslami, Nima, et al.
Veröffentlicht: (2025)
von: Sheikholeslami, Nima, et al.
Veröffentlicht: (2025)
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
von: Maheshwary, Rishabh, et al.
Veröffentlicht: (2025)
von: Maheshwary, Rishabh, et al.
Veröffentlicht: (2025)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
von: Jian, Xiangru, et al.
Veröffentlicht: (2026)
von: Jian, Xiangru, et al.
Veröffentlicht: (2026)
StarFlow: Generating Structured Workflow Outputs From Sketch Images
von: Bechard, Patrice, et al.
Veröffentlicht: (2025)
von: Bechard, Patrice, et al.
Veröffentlicht: (2025)
DeepSRGM -- Sequence Classification and Ranking in Indian Classical Music with Deep Learning
von: Madhusudhan, Sathwik Tejaswi, et al.
Veröffentlicht: (2024)
von: Madhusudhan, Sathwik Tejaswi, et al.
Veröffentlicht: (2024)
Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models
von: Madhusudhan, Nishanth, et al.
Veröffentlicht: (2024)
von: Madhusudhan, Nishanth, et al.
Veröffentlicht: (2024)
AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
von: Nguyen, Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Hoang, et al.
Veröffentlicht: (2025)
R2V Agent: Teaching SLMs When to Ask for Help
von: Hemadri, Raghu Vamshi, et al.
Veröffentlicht: (2026)
von: Hemadri, Raghu Vamshi, et al.
Veröffentlicht: (2026)
Grammar Search for Multi-Agent Systems
von: Singh, Mayank, et al.
Veröffentlicht: (2025)
von: Singh, Mayank, et al.
Veröffentlicht: (2025)
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
von: Awal, Rabiul, et al.
Veröffentlicht: (2025)
von: Awal, Rabiul, et al.
Veröffentlicht: (2025)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
Apriel-H1: Towards Efficient Enterprise Reasoning Models
von: Ostapenko, Oleksiy, et al.
Veröffentlicht: (2025)
von: Ostapenko, Oleksiy, et al.
Veröffentlicht: (2025)
Curry-DPO: Enhancing Alignment using Curriculum Learning & Ranked Preferences
von: Pattnaik, Pulkit, et al.
Veröffentlicht: (2024)
von: Pattnaik, Pulkit, et al.
Veröffentlicht: (2024)
M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models
von: Maheshwary, Rishabh, et al.
Veröffentlicht: (2024)
von: Maheshwary, Rishabh, et al.
Veröffentlicht: (2024)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
von: Feizi, Aarash, et al.
Veröffentlicht: (2025)
von: Feizi, Aarash, et al.
Veröffentlicht: (2025)
ChartInstruct: Instruction Tuning for Chart Comprehension and Reasoning
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
Revitalizing Saturated Benchmarks: A Weighted Metric Approach for Differentiating Large Language Model Performance
von: Etzine, Bryan, et al.
Veröffentlicht: (2025)
von: Etzine, Bryan, et al.
Veröffentlicht: (2025)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
von: Nayak, Shravan, et al.
Veröffentlicht: (2025)
von: Nayak, Shravan, et al.
Veröffentlicht: (2025)
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
von: Rodriguez, Juan A., et al.
Veröffentlicht: (2025)
von: Rodriguez, Juan A., et al.
Veröffentlicht: (2025)
Reducing hallucination in structured outputs via Retrieval-Augmented Generation
von: Béchard, Patrice, et al.
Veröffentlicht: (2024)
von: Béchard, Patrice, et al.
Veröffentlicht: (2024)
Multi-task retriever fine-tuning for domain-specific and efficient RAG
von: Béchard, Patrice, et al.
Veröffentlicht: (2025)
von: Béchard, Patrice, et al.
Veröffentlicht: (2025)
Generating a Low-code Complete Workflow via Task Decomposition and RAG
von: Ayala, Orlando Marquez, et al.
Veröffentlicht: (2024)
von: Ayala, Orlando Marquez, et al.
Veröffentlicht: (2024)
Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
von: Tiwari, Aman, et al.
Veröffentlicht: (2024)
von: Tiwari, Aman, et al.
Veröffentlicht: (2024)
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
von: Malay, Shiva Krishna Reddy, et al.
Veröffentlicht: (2026)
von: Malay, Shiva Krishna Reddy, et al.
Veröffentlicht: (2026)
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
von: Rodriguez, Juan, et al.
Veröffentlicht: (2024)
von: Rodriguez, Juan, et al.
Veröffentlicht: (2024)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs
von: Hashemi, Masoud, et al.
Veröffentlicht: (2025)
von: Hashemi, Masoud, et al.
Veröffentlicht: (2025)
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
von: Zhang, Le, et al.
Veröffentlicht: (2023)
von: Zhang, Le, et al.
Veröffentlicht: (2023)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
Natural Language Generation for Visualizations: State of the Art, Challenges and Future Directions
von: Hoque, Enamul, et al.
Veröffentlicht: (2024)
von: Hoque, Enamul, et al.
Veröffentlicht: (2024)
Terminal Agents Suffice for Enterprise Automation
von: Bechard, Patrice, et al.
Veröffentlicht: (2026)
von: Bechard, Patrice, et al.
Veröffentlicht: (2026)
Grounding Computer Use Agents on Human Demonstrations
von: Feizi, Aarash, et al.
Veröffentlicht: (2025)
von: Feizi, Aarash, et al.
Veröffentlicht: (2025)
Are Large Vision Language Models up to the Challenge of Chart Comprehension and Reasoning? An Extensive Investigation into the Capabilities and Limitations of LVLMs
von: Islam, Mohammed Saidul, et al.
Veröffentlicht: (2024)
von: Islam, Mohammed Saidul, et al.
Veröffentlicht: (2024)
VisMin: Visual Minimal-Change Understanding
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
Apriel-1.5-15b-Thinker
von: Radhakrishna, Shruthan, et al.
Veröffentlicht: (2025)
von: Radhakrishna, Shruthan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval
von: Sheikholeslami, Nima, et al.
Veröffentlicht: (2025) -
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
von: Masry, Ahmed, et al.
Veröffentlicht: (2025) -
ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
von: Masry, Ahmed, et al.
Veröffentlicht: (2024) -
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
von: Maheshwary, Rishabh, et al.
Veröffentlicht: (2025) -
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)