SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory
Fuente:
arXiv
Saved in:
| Main Authors: | Chai, Huacan, Wang, Yukai, Yang, Yingxuan, Peng, Dan, Song, Yuanyi, Fu, Zhihui, Liu, Weiwen, Lin, Jianghao, Wang, Jun, Zhang, Weinan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentNet: Decentralized Evolutionary Coordination for LLM-based Multi-Agent Systems
by: Yang, Yingxuan, et al.
Published: (2025)
by: Yang, Yingxuan, et al.
Published: (2025)
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
by: Wang, Jingxing, et al.
Published: (2026)
by: Wang, Jingxing, et al.
Published: (2026)
PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
by: Chai, Huacan, et al.
Published: (2025)
by: Chai, Huacan, et al.
Published: (2025)
SkillMAS: Skill Co-Evolution with LLM-based Multi-Agent System
by: Pan, Shuai, et al.
Published: (2026)
by: Pan, Shuai, et al.
Published: (2026)
A Survey of AI Agent Protocols
by: Yang, Yingxuan, et al.
Published: (2025)
by: Yang, Yingxuan, et al.
Published: (2025)
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
by: Zhou, Chenyu, et al.
Published: (2026)
by: Zhou, Chenyu, et al.
Published: (2026)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
by: Song, Yuanyi, et al.
Published: (2025)
by: Song, Yuanyi, et al.
Published: (2025)
Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
by: Wu, Zheng, et al.
Published: (2025)
by: Wu, Zheng, et al.
Published: (2025)
PocketLLM: Enabling On-Device Fine-Tuning for Personalized LLMs
by: Peng, Dan, et al.
Published: (2024)
by: Peng, Dan, et al.
Published: (2024)
P3: A Policy-Driven, Pace-Adaptive, and Diversity-Promoted Framework for data pruning in LLM Training
by: Yang, Yingxuan, et al.
Published: (2024)
by: Yang, Yingxuan, et al.
Published: (2024)
CodeApex: A Bilingual Programming Evaluation Benchmark for Large Language Models
by: Fu, Lingyue, et al.
Published: (2023)
by: Fu, Lingyue, et al.
Published: (2023)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
To Know is to Construct: Schema-Constrained Generation for Agent Memory
by: Zheng, Lei, et al.
Published: (2026)
by: Zheng, Lei, et al.
Published: (2026)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
by: Peng, Dan, et al.
Published: (2025)
by: Peng, Dan, et al.
Published: (2025)
Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning
by: Wu, Zheng, et al.
Published: (2026)
by: Wu, Zheng, et al.
Published: (2026)
DynaTree: Dynamic Agentic Retrieval Tree for Time-Sensitive News Retrieval
by: Qi, Siyuan, et al.
Published: (2026)
by: Qi, Siyuan, et al.
Published: (2026)
Agent Exchange: Shaping the Future of AI Agent Economics
by: Yang, Yingxuan, et al.
Published: (2025)
by: Yang, Yingxuan, et al.
Published: (2025)
Position: The Real Barrier to LLM Agent Usability is Agentic ROI
by: Liu, Weiwen, et al.
Published: (2025)
by: Liu, Weiwen, et al.
Published: (2025)
A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
by: Xi, Yunjia, et al.
Published: (2025)
by: Xi, Yunjia, et al.
Published: (2025)
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
by: Xi, Yunjia, et al.
Published: (2025)
by: Xi, Yunjia, et al.
Published: (2025)
OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents
by: Sun, Qiang, et al.
Published: (2024)
by: Sun, Qiang, et al.
Published: (2024)
From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory
by: Xia, Siyu, et al.
Published: (2025)
by: Xia, Siyu, et al.
Published: (2025)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
by: Zheng, Congmin, et al.
Published: (2025)
by: Zheng, Congmin, et al.
Published: (2025)
LLM-based Multi-Agent Systems: Techniques and Business Perspectives
by: Yang, Yingxuan, et al.
Published: (2024)
by: Yang, Yingxuan, et al.
Published: (2024)
ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
by: Zhu, Jiachen, et al.
Published: (2025)
by: Zhu, Jiachen, et al.
Published: (2025)
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
by: Wu, Zheng, et al.
Published: (2025)
by: Wu, Zheng, et al.
Published: (2025)
RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction
by: Bian, Haonan, et al.
Published: (2026)
by: Bian, Haonan, et al.
Published: (2026)
Contexting as Recommendation: Evolutionary Collaborative Filtering for Context Engineering
by: Zhu, Jiachen, et al.
Published: (2026)
by: Zhu, Jiachen, et al.
Published: (2026)
MMSkills: Towards Multimodal Skills for General Visual Agents
by: Zhang, Kangning, et al.
Published: (2026)
by: Zhang, Kangning, et al.
Published: (2026)
TRAD: Enhancing LLM Agents with Step-Wise Thought Retrieval and Aligned Decision
by: Zhou, Ruiwen, et al.
Published: (2024)
by: Zhou, Ruiwen, et al.
Published: (2024)
Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization
by: Zhu, Jiachen, et al.
Published: (2026)
by: Zhu, Jiachen, et al.
Published: (2026)
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
by: Bei, Yuanchen, et al.
Published: (2026)
by: Bei, Yuanchen, et al.
Published: (2026)
Adaptive Milestone Reward for GUI Agents
by: Zheng, Congmin, et al.
Published: (2026)
by: Zheng, Congmin, et al.
Published: (2026)
UEval: A Benchmark for Unified Multimodal Generation
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem
by: Wu, Fangwen, et al.
Published: (2025)
by: Wu, Fangwen, et al.
Published: (2025)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning
by: Yang, Bohao, et al.
Published: (2025)
by: Yang, Bohao, et al.
Published: (2025)
Decoupling Scores and Text: The Politeness Principle in Peer Review
by: Wen, Yingxuan
Published: (2026)
by: Wen, Yingxuan
Published: (2026)
Similar Items
-
AgentNet: Decentralized Evolutionary Coordination for LLM-based Multi-Agent Systems
by: Yang, Yingxuan, et al.
Published: (2025) -
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
by: Wang, Jingxing, et al.
Published: (2026) -
PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
by: Chai, Huacan, et al.
Published: (2025) -
SkillMAS: Skill Co-Evolution with LLM-based Multi-Agent System
by: Pan, Shuai, et al.
Published: (2026) -
A Survey of AI Agent Protocols
by: Yang, Yingxuan, et al.
Published: (2025)