Task-Focused Memorization for Multimodal Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zou, Tao, He, Yichen, Qiu, Tian, Lin, Yuan, Li, Hang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
von: Long, Lin, et al.
Veröffentlicht: (2025)
von: Long, Lin, et al.
Veröffentlicht: (2025)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
HybridDepth: Robust Metric Depth Fusion by Leveraging Depth from Focus and Single-Image Priors
von: Ganj, Ashkan, et al.
Veröffentlicht: (2024)
von: Ganj, Ashkan, et al.
Veröffentlicht: (2024)
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
von: Zeng, Ziyun, et al.
Veröffentlicht: (2025)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2025)
FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language Models
von: Wang, Hongyang, et al.
Veröffentlicht: (2025)
von: Wang, Hongyang, et al.
Veröffentlicht: (2025)
Agent Skills Should Go Beyond Text: The Case for Visual Skills
von: Xu, Binxiao, et al.
Veröffentlicht: (2026)
von: Xu, Binxiao, et al.
Veröffentlicht: (2026)
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
von: Jin, Er, et al.
Veröffentlicht: (2025)
von: Jin, Er, et al.
Veröffentlicht: (2025)
Modeling Visual Memorability Assessment with Autoencoders Reveals Characteristics of Memorable Images
von: Bagheri, Elham, et al.
Veröffentlicht: (2024)
von: Bagheri, Elham, et al.
Veröffentlicht: (2024)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
Memorizing SAM: 3D Medical Segment Anything Model with Memorizing Transformer
von: Shao, Xinyuan, et al.
Veröffentlicht: (2024)
von: Shao, Xinyuan, et al.
Veröffentlicht: (2024)
MMVP: A Multimodal MoCap Dataset with Vision and Pressure Sensors
von: Zhang, He, et al.
Veröffentlicht: (2024)
von: Zhang, He, et al.
Veröffentlicht: (2024)
Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models
von: Szu-Tu, Li-Zhong, et al.
Veröffentlicht: (2025)
von: Szu-Tu, Li-Zhong, et al.
Veröffentlicht: (2025)
Filtering Memorization from Parameter-Space in Diffusion Models
von: Zhe, Yu, et al.
Veröffentlicht: (2026)
von: Zhe, Yu, et al.
Veröffentlicht: (2026)
MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation
von: Zhang, Guohui, et al.
Veröffentlicht: (2025)
von: Zhang, Guohui, et al.
Veröffentlicht: (2025)
Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks
von: Jin, Jing, et al.
Veröffentlicht: (2026)
von: Jin, Jing, et al.
Veröffentlicht: (2026)
ChatterBox: Multi-round Multimodal Referring and Grounding
von: Tian, Yunjie, et al.
Veröffentlicht: (2024)
von: Tian, Yunjie, et al.
Veröffentlicht: (2024)
Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
Mitigating Memorization in Text-to-Image Diffusion via Region-Aware Prompt Augmentation and Multimodal Copy Detection
von: Chen, Yunzhuo, et al.
Veröffentlicht: (2026)
von: Chen, Yunzhuo, et al.
Veröffentlicht: (2026)
How Diffusion Models Memorize
von: Kim, Juyeop, et al.
Veröffentlicht: (2025)
von: Kim, Juyeop, et al.
Veröffentlicht: (2025)
Task-driven Image Fusion with Learnable Fusion Loss
von: Bai, Haowen, et al.
Veröffentlicht: (2024)
von: Bai, Haowen, et al.
Veröffentlicht: (2024)
One Model, Two Minds: Task-Conditioned Reasoning for Unified Image Quality and Aesthetic Assessment
von: Yin, Wen, et al.
Veröffentlicht: (2026)
von: Yin, Wen, et al.
Veröffentlicht: (2026)
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
von: Wu, Ziheng, et al.
Veröffentlicht: (2025)
von: Wu, Ziheng, et al.
Veröffentlicht: (2025)
Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
von: Song, Zijian, et al.
Veröffentlicht: (2025)
von: Song, Zijian, et al.
Veröffentlicht: (2025)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents
von: Wang, Shengqin, et al.
Veröffentlicht: (2026)
von: Wang, Shengqin, et al.
Veröffentlicht: (2026)
Data Processing Techniques for Modern Multimodal Models
von: Li, Yinheng, et al.
Veröffentlicht: (2024)
von: Li, Yinheng, et al.
Veröffentlicht: (2024)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
GEMS: Agent-Native Multimodal Generation with Memory and Skills
von: He, Zefeng, et al.
Veröffentlicht: (2026)
von: He, Zefeng, et al.
Veröffentlicht: (2026)
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
Investigating Memorization in Video Diffusion Models
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Towards Memorization-Free Diffusion Models
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
von: Zhang, Bofei, et al.
Veröffentlicht: (2025)
von: Zhang, Bofei, et al.
Veröffentlicht: (2025)
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
von: Jang, Lawrence, et al.
Veröffentlicht: (2024)
von: Jang, Lawrence, et al.
Veröffentlicht: (2024)
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
von: Luo, Katie, et al.
Veröffentlicht: (2025)
von: Luo, Katie, et al.
Veröffentlicht: (2025)
On Memorization in Diffusion Models
von: Gu, Xiangming, et al.
Veröffentlicht: (2023)
von: Gu, Xiangming, et al.
Veröffentlicht: (2023)
ObjMST: An Object-Focused Multimodal Style Transfer Framework
von: Kamra, Chanda Grover, et al.
Veröffentlicht: (2025)
von: Kamra, Chanda Grover, et al.
Veröffentlicht: (2025)
Evolving Without Ending: Unifying Multimodal Incremental Learning for Continual Panoptic Perception
von: Yuan, Bo, et al.
Veröffentlicht: (2026)
von: Yuan, Bo, et al.
Veröffentlicht: (2026)
An Inversion-based Measure of Memorization for Diffusion Models
von: Ma, Zhe, et al.
Veröffentlicht: (2024)
von: Ma, Zhe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
von: Long, Lin, et al.
Veröffentlicht: (2025) -
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026) -
HybridDepth: Robust Metric Depth Fusion by Leveraging Depth from Focus and Single-Image Priors
von: Ganj, Ashkan, et al.
Veröffentlicht: (2024) -
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
von: Zeng, Ziyun, et al.
Veröffentlicht: (2025) -
FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language Models
von: Wang, Hongyang, et al.
Veröffentlicht: (2025)