Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xinbei, Wang, Yiting, Yao, Yao, Yuan, Tongxin, Zhang, Aston, Zhang, Zhuosheng, Zhao, Hai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation
by: Ma, Xinbei, et al.
Published: (2024)
by: Ma, Xinbei, et al.
Published: (2024)
You Only Look at Screens: Multimodal Chain-of-Action Agents
by: Zhang, Zhuosheng, et al.
Published: (2023)
by: Zhang, Zhuosheng, et al.
Published: (2023)
How Deep is Love in LLMs' Hearts? Exploring Semantic Size in Human-like Cognition
by: Yao, Yao, et al.
Published: (2025)
by: Yao, Yao, et al.
Published: (2025)
Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
by: Ju, Tianjie, et al.
Published: (2024)
by: Ju, Tianjie, et al.
Published: (2024)
Multimodal Chain-of-Thought Reasoning in Language Models
by: Zhang, Zhuosheng, et al.
Published: (2023)
by: Zhang, Zhuosheng, et al.
Published: (2023)
MEGen: Generative Backdoor into Large Language Models via Model Editing
by: Qiu, Jiyang, et al.
Published: (2024)
by: Qiu, Jiyang, et al.
Published: (2024)
On the Robustness of Editing Large Language Models
by: Ma, Xinbei, et al.
Published: (2024)
by: Ma, Xinbei, et al.
Published: (2024)
Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning
by: Wu, Zheng, et al.
Published: (2026)
by: Wu, Zheng, et al.
Published: (2026)
LESA: Learnable LLM Layer Scaling-Up
by: Yang, Yifei, et al.
Published: (2025)
by: Yang, Yifei, et al.
Published: (2025)
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints
by: Yang, Dongjie, et al.
Published: (2025)
by: Yang, Dongjie, et al.
Published: (2025)
Plan-over-Graph: Towards Parallelable LLM Agent Schedule
by: Zhang, Shiqi, et al.
Published: (2025)
by: Zhang, Shiqi, et al.
Published: (2025)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
by: Yuan, Tongxin, et al.
Published: (2024)
by: Yuan, Tongxin, et al.
Published: (2024)
GLaPE: Gold Label-agnostic Prompt Evaluation and Optimization for Large Language Model
by: Zhang, Xuanchang, et al.
Published: (2024)
by: Zhang, Xuanchang, et al.
Published: (2024)
Textual-to-Visual Iterative Self-Verification for Slide Generation
by: Xu, Yunqing, et al.
Published: (2025)
by: Xu, Yunqing, et al.
Published: (2025)
SirLLM: Streaming Infinite Retentive LLM
by: Yao, Yao, et al.
Published: (2024)
by: Yao, Yao, et al.
Published: (2024)
Mitigating Misleading Chain-of-Thought Reasoning with Selective Filtering
by: Wu, Yexin, et al.
Published: (2024)
by: Wu, Yexin, et al.
Published: (2024)
DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems
by: Zou, Anni, et al.
Published: (2024)
by: Zou, Anni, et al.
Published: (2024)
The Hunger Game Debate: On the Emergence of Over-Competition in Multi-Agent Systems
by: Ma, Xinbei, et al.
Published: (2025)
by: Ma, Xinbei, et al.
Published: (2025)
PROM: A Phrase-level Copying Mechanism with Pre-training for Abstractive Summarization
by: Ma, Xinbei, et al.
Published: (2023)
by: Ma, Xinbei, et al.
Published: (2023)
Self-Prompting Large Language Models for Zero-Shot Open-Domain QA
by: Li, Junlong, et al.
Published: (2022)
by: Li, Junlong, et al.
Published: (2022)
PGPO: Enhancing Agent Reasoning via Pseudocode-style Planning Guided Preference Optimization
by: Cao, Zouying, et al.
Published: (2025)
by: Cao, Zouying, et al.
Published: (2025)
GKT: A Novel Guidance-Based Knowledge Transfer Framework For Efficient Cloud-edge Collaboration LLM Deployment
by: Yao, Yao, et al.
Published: (2024)
by: Yao, Yao, et al.
Published: (2024)
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
by: Shi, Luohe, et al.
Published: (2024)
by: Shi, Luohe, et al.
Published: (2024)
Adaptive Milestone Reward for GUI Agents
by: Zheng, Congmin, et al.
Published: (2026)
by: Zheng, Congmin, et al.
Published: (2026)
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
by: Diao, Lingxiao, et al.
Published: (2025)
by: Diao, Lingxiao, et al.
Published: (2025)
Generalizable Chain-of-Thought Prompting in Mixed-task Scenarios with Large Language Models
by: Zou, Anni, et al.
Published: (2023)
by: Zou, Anni, et al.
Published: (2023)
Thinking in a Crowd: How Auxiliary Information Shapes LLM Reasoning
by: Zhao, Haodong, et al.
Published: (2025)
by: Zhao, Haodong, et al.
Published: (2025)
Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
by: Cheng, Pengzhou, et al.
Published: (2025)
by: Cheng, Pengzhou, et al.
Published: (2025)
Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Dynamic Planning for LLM-based Graphical User Interface Automation
by: Zhang, Shaoqing, et al.
Published: (2024)
by: Zhang, Shaoqing, et al.
Published: (2024)
Towards End-to-End Open Conversational Machine Reading
by: Zhou, Sizhe, et al.
Published: (2022)
by: Zhou, Sizhe, et al.
Published: (2022)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
by: Xu, Haolei, et al.
Published: (2026)
by: Xu, Haolei, et al.
Published: (2026)
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
by: Guo, Yuan, et al.
Published: (2025)
by: Guo, Yuan, et al.
Published: (2025)
Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit
by: Choi, Jiwoo, et al.
Published: (2026)
by: Choi, Jiwoo, et al.
Published: (2026)
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
by: Yao, Yao, et al.
Published: (2023)
by: Yao, Yao, et al.
Published: (2023)
Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models
by: Shi, Luohe, et al.
Published: (2024)
by: Shi, Luohe, et al.
Published: (2024)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
by: Xia, Fanzeng, et al.
Published: (2024)
by: Xia, Fanzeng, et al.
Published: (2024)
LITE: Modeling Environmental Ecosystems with Multimodal Large Language Models
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
by: Yang, Dongjie, et al.
Published: (2024)
by: Yang, Dongjie, et al.
Published: (2024)
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
by: Zhao, Beidi, et al.
Published: (2026)
by: Zhao, Beidi, et al.
Published: (2026)
Similar Items
-
CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation
by: Ma, Xinbei, et al.
Published: (2024) -
You Only Look at Screens: Multimodal Chain-of-Action Agents
by: Zhang, Zhuosheng, et al.
Published: (2023) -
How Deep is Love in LLMs' Hearts? Exploring Semantic Size in Human-like Cognition
by: Yao, Yao, et al.
Published: (2025) -
Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
by: Ju, Tianjie, et al.
Published: (2024) -
Multimodal Chain-of-Thought Reasoning in Language Models
by: Zhang, Zhuosheng, et al.
Published: (2023)