Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Buettner, Kyle, Kovashka, Adriana |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling
by: Buettner, Kyle, et al.
Published: (2025)
by: Buettner, Kyle, et al.
Published: (2025)
Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition
by: Buettner, Kyle, et al.
Published: (2024)
by: Buettner, Kyle, et al.
Published: (2024)
Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
by: Wang, Yufei, et al.
Published: (2025)
by: Wang, Yufei, et al.
Published: (2025)
Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
by: Gungor, Cagri, et al.
Published: (2024)
by: Gungor, Cagri, et al.
Published: (2024)
Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativity
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
by: Ou, Siqu, et al.
Published: (2025)
by: Ou, Siqu, et al.
Published: (2025)
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion
by: Chen, Zhuokun, et al.
Published: (2024)
by: Chen, Zhuokun, et al.
Published: (2024)
Non-Contrastive Vision-Language Learning with Predictive Embedding Alignment
by: Kuhn, Lukas, et al.
Published: (2026)
by: Kuhn, Lukas, et al.
Published: (2026)
LVLM-Aided Alignment of Task-Specific Vision Models
by: Koebler, Alexander, et al.
Published: (2025)
by: Koebler, Alexander, et al.
Published: (2025)
mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval
by: Kim, Kyeong Seon, et al.
Published: (2026)
by: Kim, Kyeong Seon, et al.
Published: (2026)
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
by: Yu, Xiaomin, et al.
Published: (2026)
by: Yu, Xiaomin, et al.
Published: (2026)
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
by: Kolavi, Adithya S, et al.
Published: (2025)
by: Kolavi, Adithya S, et al.
Published: (2025)
One Model to Translate Them All: Universal Any-to-Any Translation for Heterogeneous Collaborative Perception
by: Li, Yang, et al.
Published: (2026)
by: Li, Yang, et al.
Published: (2026)
MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation
by: Li, Bo, et al.
Published: (2024)
by: Li, Bo, et al.
Published: (2024)
VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents
by: Zhang, Zhengbo, et al.
Published: (2026)
by: Zhang, Zhengbo, et al.
Published: (2026)
Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models
by: Bazi, Yakoub, et al.
Published: (2026)
by: Bazi, Yakoub, et al.
Published: (2026)
Bridging the Gap Between Multimodal Foundation Models and World Models
by: He, Xuehai
Published: (2025)
by: He, Xuehai
Published: (2025)
Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
by: Cao, Min, et al.
Published: (2025)
by: Cao, Min, et al.
Published: (2025)
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
by: Wen, Ziqi, et al.
Published: (2026)
by: Wen, Ziqi, et al.
Published: (2026)
Enhancing Weakly-Supervised Object Detection on Static Images through (Hallucinated) Motion
by: Gungor, Cagri, et al.
Published: (2024)
by: Gungor, Cagri, et al.
Published: (2024)
Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition
by: Malakouti, Sina, et al.
Published: (2025)
by: Malakouti, Sina, et al.
Published: (2025)
The Face of Persuasion: Analyzing Bias and Generating Culture-Aware Ads
by: Aghazadeh, Aysan, et al.
Published: (2025)
by: Aghazadeh, Aysan, et al.
Published: (2025)
VEIL: Vetting Extracted Image Labels from In-the-Wild Captions for Weakly-Supervised Object Detection
by: Rai, Arushi, et al.
Published: (2023)
by: Rai, Arushi, et al.
Published: (2023)
Generalizing Sports Feedback Generation by Watching Competitions and Reading Books: A Rock Climbing Case Study
by: Rai, Arushi, et al.
Published: (2026)
by: Rai, Arushi, et al.
Published: (2026)
Learning Consistent Temporal Grounding between Related Tasks in Sports Coaching
by: Rai, Arushi, et al.
Published: (2026)
by: Rai, Arushi, et al.
Published: (2026)
Grasping Partially Occluded Objects Using Autoencoder-Based Point Cloud Inpainting
by: Koebler, Alexander, et al.
Published: (2025)
by: Koebler, Alexander, et al.
Published: (2025)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2025)
by: Zhang, Jiarui, et al.
Published: (2025)
Edge Reliability Gap in Vision-Language Models: Quantifying Failure Modes of Compressed VLMs Under Visual Corruption
by: Erol, Mehmet Kaan
Published: (2026)
by: Erol, Mehmet Kaan
Published: (2026)
UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
by: Guo, Hongyu, et al.
Published: (2025)
by: Guo, Hongyu, et al.
Published: (2025)
Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval
by: Most, Alexander, et al.
Published: (2025)
by: Most, Alexander, et al.
Published: (2025)
Asynchronous Perception Machine For Efficient Test-Time-Training
by: Modi, Rajat, et al.
Published: (2024)
by: Modi, Rajat, et al.
Published: (2024)
Multimodal Contextualized Support for Enhancing Video Retrieval System
by: Nguyen-Le, Quoc-Bao, et al.
Published: (2024)
by: Nguyen-Le, Quoc-Bao, et al.
Published: (2024)
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
by: Kim, Jeonghwan, et al.
Published: (2026)
by: Kim, Jeonghwan, et al.
Published: (2026)
RALAD: Bridging the Real-to-Sim Domain Gap in Autonomous Driving with Retrieval-Augmented Learning
by: Zuo, Jiacheng, et al.
Published: (2025)
by: Zuo, Jiacheng, et al.
Published: (2025)
Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects
by: Lepori, Michael A., et al.
Published: (2024)
by: Lepori, Michael A., et al.
Published: (2024)
XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags
by: Shohan, Faisal Tareque, et al.
Published: (2024)
by: Shohan, Faisal Tareque, et al.
Published: (2024)
PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
Chitrakshara: A Large Multilingual Multimodal Dataset for Indian languages
by: Khan, Shaharukh, et al.
Published: (2026)
by: Khan, Shaharukh, et al.
Published: (2026)
On Train-Test Class Overlap and Detection for Image Retrieval
by: Song, Chull Hwan, et al.
Published: (2024)
by: Song, Chull Hwan, et al.
Published: (2024)
Similar Items
-
A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling
by: Buettner, Kyle, et al.
Published: (2025) -
Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition
by: Buettner, Kyle, et al.
Published: (2024) -
Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
by: Wang, Yufei, et al.
Published: (2025) -
Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
by: Gungor, Cagri, et al.
Published: (2024) -
Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativity
by: Hou, Zhaoyi Joey, et al.
Published: (2025)