Beyond Words: Multimodal LLM Knows When to Speak
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Zikai, Ouyang, Yi, Lee, Yi-Lun, Yu, Chen-Ping, Tsai, Yi-Hsuan, Yin, Zhaozheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
by: Lee, Yi-Lun, et al.
Published: (2024)
by: Lee, Yi-Lun, et al.
Published: (2024)
CCDNet: Learning to Detect Camouflage against Distractors in Infrared Small Target Detection
by: Liao, Zikai, et al.
Published: (2026)
by: Liao, Zikai, et al.
Published: (2026)
Exemplar Masking for Multimodal Incremental Learning
by: Lee, Yi-Lun, et al.
Published: (2024)
by: Lee, Yi-Lun, et al.
Published: (2024)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems
by: Madhusudhan, Nishanth, et al.
Published: (2026)
by: Madhusudhan, Nishanth, et al.
Published: (2026)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
by: Zheng, Ruobing, et al.
Published: (2026)
by: Zheng, Ruobing, et al.
Published: (2026)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
Toward Real-world BEV Perception: Depth Uncertainty Estimation via Gaussian Splatting
by: Lu, Shu-Wei, et al.
Published: (2025)
by: Lu, Shu-Wei, et al.
Published: (2025)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
by: Huang, Kui, et al.
Published: (2025)
by: Huang, Kui, et al.
Published: (2025)
Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cues
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
Weakly Supervised 3D Object Detection via Multi-Level Visual Guidance
by: Huang, Kuan-Chih, et al.
Published: (2023)
by: Huang, Kuan-Chih, et al.
Published: (2023)
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
by: Chen, Kai, et al.
Published: (2024)
by: Chen, Kai, et al.
Published: (2024)
Ranking-aware adapter for text-driven image ordering with CLIP
by: Yu, Wei-Hsiang, et al.
Published: (2024)
by: Yu, Wei-Hsiang, et al.
Published: (2024)
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
by: Zong, Yi, et al.
Published: (2024)
by: Zong, Yi, et al.
Published: (2024)
Generalizable Entity Grounding via Assistance of Large Language Model
by: Qi, Lu, et al.
Published: (2024)
by: Qi, Lu, et al.
Published: (2024)
LLM-driven Medical Report Generation via Communication-efficient Heterogeneous Federated Learning
by: Che, Haoxuan, et al.
Published: (2025)
by: Che, Haoxuan, et al.
Published: (2025)
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
by: Burapacheep, Jirayu, et al.
Published: (2024)
by: Burapacheep, Jirayu, et al.
Published: (2024)
Hypergraph-State Collaborative Reasoning for Multi-Object Tracking
by: Song, Zikai, et al.
Published: (2026)
by: Song, Zikai, et al.
Published: (2026)
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
by: Park, Junsung, et al.
Published: (2025)
by: Park, Junsung, et al.
Published: (2025)
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
by: Lawrence, Logan, et al.
Published: (2025)
by: Lawrence, Logan, et al.
Published: (2025)
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
by: Du, Dazhao, et al.
Published: (2026)
by: Du, Dazhao, et al.
Published: (2026)
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
by: Hou, Wenjun, et al.
Published: (2024)
by: Hou, Wenjun, et al.
Published: (2024)
Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering
by: Zou, Yuanhao, et al.
Published: (2025)
by: Zou, Yuanhao, et al.
Published: (2025)
SemiDAViL: Semi-supervised Domain Adaptation with Vision-Language Guidance for Semantic Segmentation
by: Basak, Hritam, et al.
Published: (2025)
by: Basak, Hritam, et al.
Published: (2025)
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
by: Lin, Junyan, et al.
Published: (2026)
by: Lin, Junyan, et al.
Published: (2026)
Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
by: Caffagni, Davide, et al.
Published: (2025)
by: Caffagni, Davide, et al.
Published: (2025)
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
by: He, Zoe Wanying, et al.
Published: (2025)
by: He, Zoe Wanying, et al.
Published: (2025)
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
by: Hsu, Ting-Yao E., et al.
Published: (2025)
by: Hsu, Ting-Yao E., et al.
Published: (2025)
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Imp: Highly Capable Large Multimodal Models for Mobile Devices
by: Shao, Zhenwei, et al.
Published: (2024)
by: Shao, Zhenwei, et al.
Published: (2024)
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning
by: Wang, Yeyuan, et al.
Published: (2025)
by: Wang, Yeyuan, et al.
Published: (2025)
DreamLLM: Synergistic Multimodal Comprehension and Creation
by: Dong, Runpei, et al.
Published: (2023)
by: Dong, Runpei, et al.
Published: (2023)
PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object Detection
by: Huang, Kuan-Chih, et al.
Published: (2023)
by: Huang, Kuan-Chih, et al.
Published: (2023)
VITA: Towards Open-Source Interactive Omni Multimodal LLM
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction
by: Lin, Zhi-Yi, et al.
Published: (2026)
by: Lin, Zhi-Yi, et al.
Published: (2026)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
by: Sun, Kaiser, et al.
Published: (2026)
by: Sun, Kaiser, et al.
Published: (2026)
Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts
by: Fang, Shuangkang, et al.
Published: (2024)
by: Fang, Shuangkang, et al.
Published: (2024)
SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
by: Shu, Fangxun, et al.
Published: (2025)
by: Shu, Fangxun, et al.
Published: (2025)
Similar Items
-
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
by: Lee, Yi-Lun, et al.
Published: (2024) -
CCDNet: Learning to Detect Camouflage against Distractors in Infrared Small Target Detection
by: Liao, Zikai, et al.
Published: (2026) -
Exemplar Masking for Multimodal Incremental Learning
by: Lee, Yi-Lun, et al.
Published: (2024) -
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
by: Wang, Yueqian, et al.
Published: (2024) -
Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems
by: Madhusudhan, Nishanth, et al.
Published: (2026)