Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Jiafeng, Jiang, Shixin, Dong, Xuan, Wang, Ning, Chu, Zheng, Su, Hui, Fu, Jinlan, Liu, Ming, Ng, See-Kiong, Qin, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts
by: Zhu, Zhihao, et al.
Published: (2026)
by: Zhu, Zhihao, et al.
Published: (2026)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
CET2: Modelling Topic Transitions for Coherent and Engaging Knowledge-Grounded Conversations
by: Xu, Lin, et al.
Published: (2024)
by: Xu, Lin, et al.
Published: (2024)
Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs
by: Liang, Jiafeng, et al.
Published: (2026)
by: Liang, Jiafeng, et al.
Published: (2026)
Hint-before-Solving Prompting: Guiding LLMs to Effectively Utilize Encoded Knowledge
by: Fu, Jinlan, et al.
Published: (2024)
by: Fu, Jinlan, et al.
Published: (2024)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
by: Zhang, Haowei, et al.
Published: (2026)
by: Zhang, Haowei, et al.
Published: (2026)
Chain of Thought Explanation for Dialogue State Tracking
by: Xu, Lin, et al.
Published: (2024)
by: Xu, Lin, et al.
Published: (2024)
AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents
by: Liang, Jiafeng, et al.
Published: (2025)
by: Liang, Jiafeng, et al.
Published: (2025)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
by: Fu, Jinlan, et al.
Published: (2025)
by: Fu, Jinlan, et al.
Published: (2025)
Aligning Large Language Models for Faithful Integrity Against Opposing Argument
by: Zhao, Yong, et al.
Published: (2025)
by: Zhao, Yong, et al.
Published: (2025)
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities
by: Goodge, Adam, et al.
Published: (2025)
by: Goodge, Adam, et al.
Published: (2025)
MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
by: Fu, Jinlan, et al.
Published: (2025)
by: Fu, Jinlan, et al.
Published: (2025)
Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
by: Wu, Zhaomin, et al.
Published: (2025)
by: Wu, Zhaomin, et al.
Published: (2025)
Principled Multimodal Representation Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Continual Multimodal Contrastive Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Confidence Elicitation: A New Attack Vector for Large Language Models
by: Formento, Brian, et al.
Published: (2025)
by: Formento, Brian, et al.
Published: (2025)
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
by: Jiang, Shixin, et al.
Published: (2024)
by: Jiang, Shixin, et al.
Published: (2024)
D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance
by: Liu, Renyang, et al.
Published: (2024)
by: Liu, Renyang, et al.
Published: (2024)
GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
by: Liang, Jiafeng, et al.
Published: (2024)
by: Liang, Jiafeng, et al.
Published: (2024)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
by: Shi, Wenhao, et al.
Published: (2024)
by: Shi, Wenhao, et al.
Published: (2024)
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
by: Nguyen, Thong, et al.
Published: (2025)
by: Nguyen, Thong, et al.
Published: (2025)
MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration
by: Xu, Lin, et al.
Published: (2023)
by: Xu, Lin, et al.
Published: (2023)
MASim: Multilingual Agent-Based Simulation for Social Science
by: Zhang, Xuan, et al.
Published: (2025)
by: Zhang, Xuan, et al.
Published: (2025)
Calibrated Multimodal Representation Learning with Missing Modalities
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
by: Zhao, James Xu, et al.
Published: (2025)
by: Zhao, James Xu, et al.
Published: (2025)
Inference-Time Attribute Distribution Alignment for Unconditional Diffusion
by: Luan, Hao, et al.
Published: (2026)
by: Luan, Hao, et al.
Published: (2026)
DDPS: Discrete Diffusion Posterior Sampling for Paths in Layered Graphs
by: Luan, Hao, et al.
Published: (2025)
by: Luan, Hao, et al.
Published: (2025)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
by: Nguyen, Thong, et al.
Published: (2023)
by: Nguyen, Thong, et al.
Published: (2023)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
by: Bin, Yi, et al.
Published: (2024)
by: Bin, Yi, et al.
Published: (2024)
Ask-before-Plan: Proactive Language Agents for Real-World Planning
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
Multi-Modal One-Shot Federated Ensemble Learning for Medical Data with Vision Large Language Model
by: Wang, Naibo, et al.
Published: (2025)
by: Wang, Naibo, et al.
Published: (2025)
Multimodal Mathematical Reasoning with Diverse Solving Perspective
by: Shi, Wenhao, et al.
Published: (2025)
by: Shi, Wenhao, et al.
Published: (2025)
Mercury: A Code Efficiency Benchmark for Code Large Language Models
by: Du, Mingzhe, et al.
Published: (2024)
by: Du, Mingzhe, et al.
Published: (2024)
Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model Unlearning
by: Liu, Renyang, et al.
Published: (2025)
by: Liu, Renyang, et al.
Published: (2025)
Plug-and-Play Policy Planner for Large Language Model Powered Dialogue Agents
by: Deng, Yang, et al.
Published: (2023)
by: Deng, Yang, et al.
Published: (2023)
SemRoDe: Macro Adversarial Training to Learn Representations That are Robust to Word-Level Attacks
by: Formento, Brian, et al.
Published: (2024)
by: Formento, Brian, et al.
Published: (2024)
Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations
by: Deng, Yang, et al.
Published: (2024)
by: Deng, Yang, et al.
Published: (2024)
One-Shot Sequential Federated Learning for Non-IID Data by Enhancing Local Model Diversity
by: Wang, Naibo, et al.
Published: (2024)
by: Wang, Naibo, et al.
Published: (2024)
Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
by: Nguyen, Thong Thanh, et al.
Published: (2024)
by: Nguyen, Thong Thanh, et al.
Published: (2024)
On the Multi-turn Instruction Following for Conversational Web Agents
by: Deng, Yang, et al.
Published: (2024)
by: Deng, Yang, et al.
Published: (2024)
Similar Items
-
Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts
by: Zhu, Zhihao, et al.
Published: (2026) -
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
by: Chen, Qian, et al.
Published: (2026) -
CET2: Modelling Topic Transitions for Coherent and Engaging Knowledge-Grounded Conversations
by: Xu, Lin, et al.
Published: (2024) -
Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs
by: Liang, Jiafeng, et al.
Published: (2026) -
Hint-before-Solving Prompting: Guiding LLMs to Effectively Utilize Encoded Knowledge
by: Fu, Jinlan, et al.
Published: (2024)