MMA: Multimodal Memory Agent
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Yihao, Cheng, Wanru, Zhang, Zeyu, Tang, Hao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multimodal Data Storage and Retrieval for Embodied AI: A Survey
di: Lu, Yihao, et al.
Pubblicazione: (2025)
di: Lu, Yihao, et al.
Pubblicazione: (2025)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
di: Wu, Wei, et al.
Pubblicazione: (2026)
di: Wu, Wei, et al.
Pubblicazione: (2026)
GEMS: Agent-Native Multimodal Generation with Memory and Skills
di: He, Zefeng, et al.
Pubblicazione: (2026)
di: He, Zefeng, et al.
Pubblicazione: (2026)
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
di: Shen, Hao, et al.
Pubblicazione: (2024)
di: Shen, Hao, et al.
Pubblicazione: (2024)
PresentAgent: Multimodal Agent for Presentation Video Generation
di: Shi, Jingwei, et al.
Pubblicazione: (2025)
di: Shi, Jingwei, et al.
Pubblicazione: (2025)
DragMesh: Interactive 3D Generation Made Easy
di: Zhang, Tianshan, et al.
Pubblicazione: (2025)
di: Zhang, Tianshan, et al.
Pubblicazione: (2025)
PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing
di: Li, Peize, et al.
Pubblicazione: (2026)
di: Li, Peize, et al.
Pubblicazione: (2026)
3D CoCa v2: Contrastive Learners with Test-Time Search for Generalizable Spatial Intelligence
di: Tang, Hao, et al.
Pubblicazione: (2026)
di: Tang, Hao, et al.
Pubblicazione: (2026)
HSG: Hyperbolic Scene Graph
di: Wang, Liyang, et al.
Pubblicazione: (2026)
di: Wang, Liyang, et al.
Pubblicazione: (2026)
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
OCR-Agent: Agentic OCR with Capability and Memory Reflection
di: Wen, Shimin, et al.
Pubblicazione: (2026)
di: Wen, Shimin, et al.
Pubblicazione: (2026)
MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
di: Guo, Minghao, et al.
Pubblicazione: (2026)
di: Guo, Minghao, et al.
Pubblicazione: (2026)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
StereoAdapter-2: Globally Structure-Consistent Underwater Stereo Depth Estimation
di: Ren, Zeyu, et al.
Pubblicazione: (2026)
di: Ren, Zeyu, et al.
Pubblicazione: (2026)
AnyDepth: Depth Estimation Made Easy
di: Ren, Zeyu, et al.
Pubblicazione: (2026)
di: Ren, Zeyu, et al.
Pubblicazione: (2026)
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
di: Wang, Pan, et al.
Pubblicazione: (2026)
di: Wang, Pan, et al.
Pubblicazione: (2026)
MMA-Diffusion: MultiModal Attack on Diffusion Models
di: Yang, Yijun, et al.
Pubblicazione: (2023)
di: Yang, Yijun, et al.
Pubblicazione: (2023)
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
di: Zhang, Zeyu, et al.
Pubblicazione: (2024)
WebCryptoAgent: Agentic Crypto Trading with Web Informatics
di: Kurban, Ali, et al.
Pubblicazione: (2026)
di: Kurban, Ali, et al.
Pubblicazione: (2026)
Planning with Unified Multimodal Models
di: Sun, Yihao, et al.
Pubblicazione: (2025)
di: Sun, Yihao, et al.
Pubblicazione: (2025)
Code2Worlds: Empowering Coding LLMs for 4D World Generation
di: Zhang, Yi, et al.
Pubblicazione: (2026)
di: Zhang, Yi, et al.
Pubblicazione: (2026)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
di: Ge, Jinchao, et al.
Pubblicazione: (2025)
di: Ge, Jinchao, et al.
Pubblicazione: (2025)
MMA-UNet: A Multi-Modal Asymmetric UNet Architecture for Infrared and Visible Image Fusion
di: Huang, Jingxue, et al.
Pubblicazione: (2024)
di: Huang, Jingxue, et al.
Pubblicazione: (2024)
Light4D: Training-Free Extreme Viewpoint 4D Video Relighting
di: Wu, Zhenghuang, et al.
Pubblicazione: (2026)
di: Wu, Zhenghuang, et al.
Pubblicazione: (2026)
UniMesh: Unifying 3D Mesh Understanding and Generation
di: Huang, Peng, et al.
Pubblicazione: (2026)
di: Huang, Peng, et al.
Pubblicazione: (2026)
SafeMo: Linguistically Grounded Unlearning for Trustworthy Text-to-Motion Generation
di: Wang, Yiling, et al.
Pubblicazione: (2026)
di: Wang, Yiling, et al.
Pubblicazione: (2026)
MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
di: Wang, Hongpeng, et al.
Pubblicazione: (2026)
di: Wang, Hongpeng, et al.
Pubblicazione: (2026)
3D CoCa: Contrastive Learners are 3D Captioners
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
ReMoMask: Retrieval-Augmented Masked Motion Generation
di: Li, Zhengdao, et al.
Pubblicazione: (2025)
di: Li, Zhengdao, et al.
Pubblicazione: (2025)
EvoVLA: Self-Evolving Vision-Language-Action Model
di: Liu, Zeting, et al.
Pubblicazione: (2025)
di: Liu, Zeting, et al.
Pubblicazione: (2025)
Multimodal Alignment and Fusion: A Survey
di: Li, Songtao, et al.
Pubblicazione: (2024)
di: Li, Songtao, et al.
Pubblicazione: (2024)
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
di: Long, Lin, et al.
Pubblicazione: (2025)
di: Long, Lin, et al.
Pubblicazione: (2025)
Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
di: Zhang, Nonghai, et al.
Pubblicazione: (2025)
di: Zhang, Nonghai, et al.
Pubblicazione: (2025)
LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
di: Chen, Nan, et al.
Pubblicazione: (2025)
di: Chen, Nan, et al.
Pubblicazione: (2025)
Not all tokens contribute equally to diffusion learning
di: Zhang, Guoqing, et al.
Pubblicazione: (2026)
di: Zhang, Guoqing, et al.
Pubblicazione: (2026)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
di: Zeng, Ziyun, et al.
Pubblicazione: (2026)
di: Zeng, Ziyun, et al.
Pubblicazione: (2026)
LaMMA-P: Generalizable Multi-Agent Long-Horizon Task Allocation and Planning with LM-Driven PDDL Planner
di: Zhang, Xiaopan, et al.
Pubblicazione: (2024)
di: Zhang, Xiaopan, et al.
Pubblicazione: (2024)
MWM: Mobile World Models for Action-Conditioned Consistent Prediction
di: Yan, Han, et al.
Pubblicazione: (2026)
di: Yan, Han, et al.
Pubblicazione: (2026)
Nav-R1: Reasoning and Navigation in Embodied Scenes
di: Liu, Qingxiang, et al.
Pubblicazione: (2025)
di: Liu, Qingxiang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Multimodal Data Storage and Retrieval for Embodied AI: A Survey
di: Lu, Yihao, et al.
Pubblicazione: (2025) -
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
di: Wu, Wei, et al.
Pubblicazione: (2026) -
GEMS: Agent-Native Multimodal Generation with Memory and Skills
di: He, Zefeng, et al.
Pubblicazione: (2026) -
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
di: Shen, Hao, et al.
Pubblicazione: (2024) -
PresentAgent: Multimodal Agent for Presentation Video Generation
di: Shi, Jingwei, et al.
Pubblicazione: (2025)