Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yufei, Kovashka, Adriana, Fernández, Loretta, Coutanche, Marc N., Wiener, Seth |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval
by: Buettner, Kyle, et al.
Published: (2024)
by: Buettner, Kyle, et al.
Published: (2024)
A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling
by: Buettner, Kyle, et al.
Published: (2025)
by: Buettner, Kyle, et al.
Published: (2025)
Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
by: Gungor, Cagri, et al.
Published: (2024)
by: Gungor, Cagri, et al.
Published: (2024)
Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativity
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition
by: Buettner, Kyle, et al.
Published: (2024)
by: Buettner, Kyle, et al.
Published: (2024)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
by: Zhan, Yufei, et al.
Published: (2024)
by: Zhan, Yufei, et al.
Published: (2024)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
by: Peng, Tianhao, et al.
Published: (2025)
by: Peng, Tianhao, et al.
Published: (2025)
Using Multimodal Foundation Models and Clustering for Improved Style Ambiguity Loss
by: Baker, James
Published: (2024)
by: Baker, James
Published: (2024)
MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems
by: Yang, Peiru, et al.
Published: (2025)
by: Yang, Peiru, et al.
Published: (2025)
Weak to Strong: VLM-Based Pseudo-Labeling as a Weakly Supervised Training Strategy in Multimodal Video-based Hidden Emotion Understanding Tasks
by: Wang, Yufei, et al.
Published: (2026)
by: Wang, Yufei, et al.
Published: (2026)
PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding
by: Erregue, Iñaki, et al.
Published: (2026)
by: Erregue, Iñaki, et al.
Published: (2026)
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
by: Zheng, Shurong, et al.
Published: (2026)
by: Zheng, Shurong, et al.
Published: (2026)
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
by: Sun, Penglei, et al.
Published: (2025)
by: Sun, Penglei, et al.
Published: (2025)
DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
by: Wang, Zhu, et al.
Published: (2025)
by: Wang, Zhu, et al.
Published: (2025)
Breaking Scale Anchoring: Frequency Representation Learning for Accurate High-Resolution Inference from Low-Resolution Training
by: Wang, Wenshuo, et al.
Published: (2025)
by: Wang, Wenshuo, et al.
Published: (2025)
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Towards Holistic Surgical Scene Understanding
by: Valderrama, Natalia, et al.
Published: (2022)
by: Valderrama, Natalia, et al.
Published: (2022)
Towards Ambiguity-Free Spatial Foundation Model: Rethinking and Decoupling Depth Ambiguity
by: Xu, Xiaohao, et al.
Published: (2025)
by: Xu, Xiaohao, et al.
Published: (2025)
Towards Understanding Graphical Perception in Large Multimodal Models
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Plots Unlock Time-Series Understanding in Multimodal Models
by: Daswani, Mayank, et al.
Published: (2024)
by: Daswani, Mayank, et al.
Published: (2024)
GeoDecoder: Empowering Multimodal Map Understanding
by: Qi, Feng, et al.
Published: (2024)
by: Qi, Feng, et al.
Published: (2024)
Region-Level Context-Aware Multimodal Understanding
by: Wei, Hongliang, et al.
Published: (2025)
by: Wei, Hongliang, et al.
Published: (2025)
Towards Generalization of Tactile Image Generation: Reference-Free Evaluation in a Leakage-Free Setting
by: Gungor, Cagri, et al.
Published: (2025)
by: Gungor, Cagri, et al.
Published: (2025)
Diffusion Prior Interpolation for Flexibility Real-World Face Super-Resolution
by: Yang, Jiarui, et al.
Published: (2024)
by: Yang, Jiarui, et al.
Published: (2024)
Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition
by: Malakouti, Sina, et al.
Published: (2025)
by: Malakouti, Sina, et al.
Published: (2025)
The Face of Persuasion: Analyzing Bias and Generating Culture-Aware Ads
by: Aghazadeh, Aysan, et al.
Published: (2025)
by: Aghazadeh, Aysan, et al.
Published: (2025)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
by: Seth, Ashish, et al.
Published: (2024)
by: Seth, Ashish, et al.
Published: (2024)
VEIL: Vetting Extracted Image Labels from In-the-Wild Captions for Weakly-Supervised Object Detection
by: Rai, Arushi, et al.
Published: (2023)
by: Rai, Arushi, et al.
Published: (2023)
Enhancing Weakly-Supervised Object Detection on Static Images through (Hallucinated) Motion
by: Gungor, Cagri, et al.
Published: (2024)
by: Gungor, Cagri, et al.
Published: (2024)
Generalizing Sports Feedback Generation by Watching Competitions and Reading Books: A Rock Climbing Case Study
by: Rai, Arushi, et al.
Published: (2026)
by: Rai, Arushi, et al.
Published: (2026)
Learning Consistent Temporal Grounding between Related Tasks in Sports Coaching
by: Rai, Arushi, et al.
Published: (2026)
by: Rai, Arushi, et al.
Published: (2026)
The Demon is in Ambiguity: Revisiting Situation Recognition with Single Positive Multi-Label Learning
by: Lin, Yiming, et al.
Published: (2025)
by: Lin, Yiming, et al.
Published: (2025)
Apollo: An Exploration of Video Understanding in Large Multimodal Models
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
by: Yao, Louie Hong, et al.
Published: (2025)
by: Yao, Louie Hong, et al.
Published: (2025)
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
by: Xing, Shuo, et al.
Published: (2025)
by: Xing, Shuo, et al.
Published: (2025)
Evaluating Compositional Scene Understanding in Multimodal Generative Models
by: Fu, Shuhao, et al.
Published: (2025)
by: Fu, Shuhao, et al.
Published: (2025)
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
Similar Items
-
Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval
by: Buettner, Kyle, et al.
Published: (2024) -
A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling
by: Buettner, Kyle, et al.
Published: (2025) -
Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
by: Gungor, Cagri, et al.
Published: (2024) -
Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativity
by: Hou, Zhaoyi Joey, et al.
Published: (2025) -
Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition
by: Buettner, Kyle, et al.
Published: (2024)