MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Dazhao, Duan, Liao, Liu, Jian, Han, Tao, Zhang, Yujia, Liu, Eric, Chen, Xi, Guo, Song |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
by: Du, Dazhao, et al.
Published: (2026)
by: Du, Dazhao, et al.
Published: (2026)
Predicting the Future by Retrieving the Past
by: Du, Dazhao, et al.
Published: (2025)
by: Du, Dazhao, et al.
Published: (2025)
Beyond Words: Multimodal LLM Knows When to Speak
by: Liao, Zikai, et al.
Published: (2025)
by: Liao, Zikai, et al.
Published: (2025)
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
by: Cao, Fanpu, et al.
Published: (2026)
by: Cao, Fanpu, et al.
Published: (2026)
SafeGround: Know When to Trust GUI Grounding Models via Uncertainty Calibration
by: Wang, Qingni, et al.
Published: (2026)
by: Wang, Qingni, et al.
Published: (2026)
Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
From Visual Synthesis to Interactive Worlds: Toward Production-Ready 3D Asset Generation
by: Wu, Jiafeng, et al.
Published: (2026)
by: Wu, Jiafeng, et al.
Published: (2026)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues
by: Park, Beomchan, et al.
Published: (2026)
by: Park, Beomchan, et al.
Published: (2026)
Benchmarking Physics-Informed Time-Series Models for Operational Global Station Weather Forecasting
by: Han, Tao, et al.
Published: (2024)
by: Han, Tao, et al.
Published: (2024)
Large Language Models Know What To Say But Not When To Speak
by: Umair, Muhammad, et al.
Published: (2024)
by: Umair, Muhammad, et al.
Published: (2024)
3D Generation for Embodied AI and Robotic Simulation: A Survey
by: Ye, Tianwei, et al.
Published: (2026)
by: Ye, Tianwei, et al.
Published: (2026)
ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
by: Xu, Zhengzhuo, et al.
Published: (2025)
by: Xu, Zhengzhuo, et al.
Published: (2025)
VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation
by: Han, Qijun, et al.
Published: (2026)
by: Han, Qijun, et al.
Published: (2026)
TRACE: Temporal Grounding Video LLM via Causal Event Modeling
by: Guo, Yongxin, et al.
Published: (2024)
by: Guo, Yongxin, et al.
Published: (2024)
Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation
by: Luo, Jiani, et al.
Published: (2026)
by: Luo, Jiani, et al.
Published: (2026)
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
by: Zhang, Bob, et al.
Published: (2025)
by: Zhang, Bob, et al.
Published: (2025)
Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents
by: Zeng, Jingying, et al.
Published: (2025)
by: Zeng, Jingying, et al.
Published: (2025)
Prompt When the Animal is: Temporal Animal Behavior Grounding with Positional Recovery Training
by: Yan, Sheng, et al.
Published: (2024)
by: Yan, Sheng, et al.
Published: (2024)
Steering Language Models Before They Speak: Logit-Level Interventions
by: An, Hyeseon, et al.
Published: (2026)
by: An, Hyeseon, et al.
Published: (2026)
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
by: Gu, Bohai, et al.
Published: (2026)
by: Gu, Bohai, et al.
Published: (2026)
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
by: Wang, Yubin, et al.
Published: (2024)
by: Wang, Yubin, et al.
Published: (2024)
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
GUI Agents with Reinforcement Learning: Toward Digital Inhabitants
by: Hu, Junan, et al.
Published: (2026)
by: Hu, Junan, et al.
Published: (2026)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
by: Huang, Jen-Tse, et al.
Published: (2025)
by: Huang, Jen-Tse, et al.
Published: (2025)
A conjecture of Nadji, Ahmia and Ram\'ırez on congruences for biregular overpartitions
by: Tang, Dazhao
Published: (2025)
by: Tang, Dazhao
Published: (2025)
Simulation of lethal control and fertility control in a demographic model for Brandt's vole Microtus brandti. / Dazhao Shi
by: Shi, Dazhao
Published: (1995)
by: Shi, Dazhao
Published: (1995)
Diffusion Language Models Know the Answer Before Decoding
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models
by: Zheng, Ziwei, et al.
Published: (2025)
by: Zheng, Ziwei, et al.
Published: (2025)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
by: Ma, Minghui, et al.
Published: (2026)
by: Ma, Minghui, et al.
Published: (2026)
CaRT: Teaching LLM Agents to Know When They Know Enough
by: Liu, Grace, et al.
Published: (2025)
by: Liu, Grace, et al.
Published: (2025)
Micromobility Flow Prediction: A Bike Sharing Station-level Study via Multi-level Spatial-Temporal Attention Neural Network
by: Yang, Xi, et al.
Published: (2025)
by: Yang, Xi, et al.
Published: (2025)
Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning
by: Ding, Chenlu, et al.
Published: (2026)
by: Ding, Chenlu, et al.
Published: (2026)
Thinking Before You Speak: A Proactive Test-time Scaling Approach
by: Liu, Cong, et al.
Published: (2025)
by: Liu, Cong, et al.
Published: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
Can MLLMs "Read" What is Missing?
by: Guo, Jindi, et al.
Published: (2026)
by: Guo, Jindi, et al.
Published: (2026)
Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
by: Zhang, Qingjie, et al.
Published: (2025)
by: Zhang, Qingjie, et al.
Published: (2025)
Similar Items
-
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
by: Du, Dazhao, et al.
Published: (2026) -
Predicting the Future by Retrieving the Past
by: Du, Dazhao, et al.
Published: (2025) -
Beyond Words: Multimodal LLM Knows When to Speak
by: Liao, Zikai, et al.
Published: (2025) -
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
by: Cao, Fanpu, et al.
Published: (2026) -
SafeGround: Know When to Trust GUI Grounding Models via Uncertainty Calibration
by: Wang, Qingni, et al.
Published: (2026)