Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
Fuente:
arXiv
Salvato in:
| Autori principali: | Bo, Weihao, Zhang, Shan, Sun, Yanpeng, Wu, Jingjing, Xie, Qunyi, Tan, Xiao, Chen, Kunbin, He, Wei, Li, Xiaofan, Zhao, Na, Wang, Jingdong, Li, Zechao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Artemis: Structured Visual Reasoning for Perception Policy Learning
di: Tang, Wei, et al.
Pubblicazione: (2025)
di: Tang, Wei, et al.
Pubblicazione: (2025)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
Exploring Effective Factors for Improving Visual In-Context Learning
di: Sun, Yanpeng, et al.
Pubblicazione: (2023)
di: Sun, Yanpeng, et al.
Pubblicazione: (2023)
FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
di: Bo, Weihao, et al.
Pubblicazione: (2025)
di: Bo, Weihao, et al.
Pubblicazione: (2025)
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
di: Tang, Wei, et al.
Pubblicazione: (2026)
di: Tang, Wei, et al.
Pubblicazione: (2026)
VRP-SAM: SAM with Visual Reference Prompt
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
Visual Position Prompt for MLLM based Visual Grounding
di: Tang, Wei, et al.
Pubblicazione: (2025)
di: Tang, Wei, et al.
Pubblicazione: (2025)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
di: Huang, Xiaohu, et al.
Pubblicazione: (2025)
di: Huang, Xiaohu, et al.
Pubblicazione: (2025)
Video4Edit: Viewing Image Editing as a Degenerate Temporal Process
di: Li, Xiaofan, et al.
Pubblicazione: (2025)
di: Li, Xiaofan, et al.
Pubblicazione: (2025)
AFANet: Adaptive Frequency-Aware Network for Weakly-Supervised Few-Shot Semantic Segmentation
di: Ma, Jiaqi, et al.
Pubblicazione: (2024)
di: Ma, Jiaqi, et al.
Pubblicazione: (2024)
FVAR: Visual Autoregressive Modeling via Next Focus Prediction
di: Li, Xiaofan, et al.
Pubblicazione: (2025)
di: Li, Xiaofan, et al.
Pubblicazione: (2025)
AgentSM: Semantic Memory for Agentic Text-to-SQL
di: Biswal, Asim, et al.
Pubblicazione: (2026)
di: Biswal, Asim, et al.
Pubblicazione: (2026)
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
di: Tan, Weihao, et al.
Pubblicazione: (2025)
di: Tan, Weihao, et al.
Pubblicazione: (2025)
RGD: Multi-LLM Based Agent Debugger via Refinement and Generation Guidance
di: Jin, Haolin, et al.
Pubblicazione: (2024)
di: Jin, Haolin, et al.
Pubblicazione: (2024)
TeleMem: Building Long-Term and Multimodal Memory for Agentic AI
di: Chen, Chunliang, et al.
Pubblicazione: (2025)
di: Chen, Chunliang, et al.
Pubblicazione: (2025)
Continual SFT Matches Multimodal RLHF with Negative Supervision
di: Zhu, Ke, et al.
Pubblicazione: (2024)
di: Zhu, Ke, et al.
Pubblicazione: (2024)
CaMML: Context-Aware Multimodal Learner for Large Models
di: Chen, Yixin, et al.
Pubblicazione: (2024)
di: Chen, Yixin, et al.
Pubblicazione: (2024)
Memory Caching: RNNs with Growing Memory
di: Behrouz, Ali, et al.
Pubblicazione: (2026)
di: Behrouz, Ali, et al.
Pubblicazione: (2026)
Agentic Reasoning and Refinement through Semantic Interaction
di: Tang, Xuxin, et al.
Pubblicazione: (2025)
di: Tang, Xuxin, et al.
Pubblicazione: (2025)
Generative Multimodal Models are In-Context Learners
di: Sun, Quan, et al.
Pubblicazione: (2023)
di: Sun, Quan, et al.
Pubblicazione: (2023)
Semantic Prompting: Agentic Incremental Narrative Refinement through Spatial Semantic Interaction
di: Tang, Xuxin, et al.
Pubblicazione: (2026)
di: Tang, Xuxin, et al.
Pubblicazione: (2026)
Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents
di: Yu, Yi, et al.
Pubblicazione: (2026)
di: Yu, Yi, et al.
Pubblicazione: (2026)
From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing
di: Zhang, Haoyuan, et al.
Pubblicazione: (2026)
di: Zhang, Haoyuan, et al.
Pubblicazione: (2026)
Through the Looking Glass: A Dual Perspective on Weakly-Supervised Few-Shot Segmentation
di: Ma, Jiaqi, et al.
Pubblicazione: (2025)
di: Ma, Jiaqi, et al.
Pubblicazione: (2025)
EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement
di: Xu, Zitong, et al.
Pubblicazione: (2026)
di: Xu, Zitong, et al.
Pubblicazione: (2026)
Wireless Agentic AI with Retrieval-Augmented Multimodal Semantic Perception
di: Liu, Guangyuan, et al.
Pubblicazione: (2025)
di: Liu, Guangyuan, et al.
Pubblicazione: (2025)
CSGO: Content-Style Composition in Text-to-Image Generation
di: Xing, Peng, et al.
Pubblicazione: (2024)
di: Xing, Peng, et al.
Pubblicazione: (2024)
Dropout Prompt Learning: Towards Robust and Adaptive Vision-Language Models
di: Chen, Biao, et al.
Pubblicazione: (2025)
di: Chen, Biao, et al.
Pubblicazione: (2025)
Semantic Generative Tuning for Unified Multimodal Models
di: Yu, Songsong, et al.
Pubblicazione: (2026)
di: Yu, Songsong, et al.
Pubblicazione: (2026)
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
di: Liu, Huan, et al.
Pubblicazione: (2024)
di: Liu, Huan, et al.
Pubblicazione: (2024)
ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation
di: Qian, Jingjing, et al.
Pubblicazione: (2026)
di: Qian, Jingjing, et al.
Pubblicazione: (2026)
Trustworthy Multimodal Fusion for Sentiment Analysis in Ordinal Sentiment Space
di: Xie, Zhuyang, et al.
Pubblicazione: (2024)
di: Xie, Zhuyang, et al.
Pubblicazione: (2024)
Stability of the 2D Boussinesq‐MHD System With Fractional Horizontal Dissipation
di: Yuan Sun, et al.
Pubblicazione: (2026)
di: Yuan Sun, et al.
Pubblicazione: (2026)
Provable Memory Efficient Self-Play Algorithm for Model-free Reinforcement Learning
di: Li, Na, et al.
Pubblicazione: (2025)
di: Li, Na, et al.
Pubblicazione: (2025)
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
di: Lyu, Pengyuan, et al.
Pubblicazione: (2024)
di: Lyu, Pengyuan, et al.
Pubblicazione: (2024)
PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis
di: Hua, Shengyi, et al.
Pubblicazione: (2025)
di: Hua, Shengyi, et al.
Pubblicazione: (2025)
Memory-Guided View Refinement for Dynamic Human-in-the-loop EQA
di: Lu, Xin, et al.
Pubblicazione: (2026)
di: Lu, Xin, et al.
Pubblicazione: (2026)
2D Memory Selectors with Giant Nonlinearity Enabled by Van der Waals Heterostructures
di: Xiaofan Wang, et al.
Pubblicazione: (2024)
di: Xiaofan Wang, et al.
Pubblicazione: (2024)
Controlling Shareholder Share Pledging and Corporate Social Security Contributions in China
di: Jingjing Jiang, et al.
Pubblicazione: (2025)
di: Jingjing Jiang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Artemis: Structured Visual Reasoning for Perception Policy Learning
di: Tang, Wei, et al.
Pubblicazione: (2025) -
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
di: Sun, Yanpeng, et al.
Pubblicazione: (2024) -
Exploring Effective Factors for Improving Visual In-Context Learning
di: Sun, Yanpeng, et al.
Pubblicazione: (2023) -
FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
di: Bo, Weihao, et al.
Pubblicazione: (2025) -
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
di: Tang, Wei, et al.
Pubblicazione: (2026)