Saved in:
| Main Authors: | Wang, Pan, Hu, Yihao, Liu, Xiujin, Yang, Jingchu, Wang, Hang, Wen, Zhihao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.17933 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CFSR: Geometry-Conditioned Shadow Removal via Physical Disentanglement
by: Wang, Pan, et al.
Published: (2026)
by: Wang, Pan, et al.
Published: (2026)
UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2026)
by: Xiao, Han, et al.
Published: (2026)
A Multi-Strategy Framework for Enhancing Shatian Pomelo Detection in Real-World Orchards
by: Wang, Pan, et al.
Published: (2025)
by: Wang, Pan, et al.
Published: (2025)
GNC-Pose: Geometry-Aware GNC-PnP for Accurate 6D Pose Estimation
by: Liu, Xiujin
Published: (2025)
by: Liu, Xiujin
Published: (2025)
Agent Skills Should Go Beyond Text: The Case for Visual Skills
by: Xu, Binxiao, et al.
Published: (2026)
by: Xu, Binxiao, et al.
Published: (2026)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
MMA: Multimodal Memory Agent
by: Lu, Yihao, et al.
Published: (2026)
by: Lu, Yihao, et al.
Published: (2026)
Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling
by: Li, Zhihao, et al.
Published: (2025)
by: Li, Zhihao, et al.
Published: (2025)
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation
by: Chen, Sixiang, et al.
Published: (2026)
by: Chen, Sixiang, et al.
Published: (2026)
ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
by: Wang, Yichen, et al.
Published: (2025)
by: Wang, Yichen, et al.
Published: (2025)
GEMS: Agent-Native Multimodal Generation with Memory and Skills
by: He, Zefeng, et al.
Published: (2026)
by: He, Zefeng, et al.
Published: (2026)
A Novel Convolution and Attention Mechanism-based Model for 6D Object Pose Estimation
by: Du, Alexander, et al.
Published: (2024)
by: Du, Alexander, et al.
Published: (2024)
Training-Free Text-Guided Image Editing with Visual Autoregressive Model
by: Wang, Yufei, et al.
Published: (2025)
by: Wang, Yufei, et al.
Published: (2025)
SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents
by: Yang, Yu, et al.
Published: (2026)
by: Yang, Yu, et al.
Published: (2026)
GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
by: Wei, Tong, et al.
Published: (2025)
by: Wei, Tong, et al.
Published: (2025)
ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
by: Wang, Chujie, et al.
Published: (2025)
by: Wang, Chujie, et al.
Published: (2025)
MemRoPE: Training-Free Infinite Video Generation via Evolving Memory Tokens
by: Kim, Youngrae, et al.
Published: (2026)
by: Kim, Youngrae, et al.
Published: (2026)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
by: Huang, Brandon, et al.
Published: (2025)
by: Huang, Brandon, et al.
Published: (2025)
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
by: Li, Xinze, et al.
Published: (2026)
by: Li, Xinze, et al.
Published: (2026)
CogVLM: Visual Expert for Pretrained Language Models
by: Wang, Weihan, et al.
Published: (2023)
by: Wang, Weihan, et al.
Published: (2023)
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
by: Lin, Yuxin, et al.
Published: (2025)
by: Lin, Yuxin, et al.
Published: (2025)
Dual Latent Memory for Visual Multi-agent System
by: Yu, Xinlei, et al.
Published: (2026)
by: Yu, Xinlei, et al.
Published: (2026)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
MARS: Memory-Enhanced Agents with Reflective Self-improvement
by: Liang, Xuechen, et al.
Published: (2025)
by: Liang, Xuechen, et al.
Published: (2025)
BrainMem: Brain-Inspired Evolving Memory for Embodied Agent Task Planning
by: Ma, Xiaoyu, et al.
Published: (2026)
by: Ma, Xiaoyu, et al.
Published: (2026)
Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual Reasoners
by: Liu, Qingyang, et al.
Published: (2026)
by: Liu, Qingyang, et al.
Published: (2026)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
by: Zhao, Yi, et al.
Published: (2026)
by: Zhao, Yi, et al.
Published: (2026)
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
ShadowMamba: State-Space Model with Boundary-Region Selective Scan for Shadow Removal
by: Zhu, Xiujin, et al.
Published: (2024)
by: Zhu, Xiujin, et al.
Published: (2024)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
by: Wu, Hang, et al.
Published: (2026)
by: Wu, Hang, et al.
Published: (2026)
A VLM-based Method for Visual Anomaly Detection in Robotic Scientific Laboratories
by: Lin, Shiwei, et al.
Published: (2025)
by: Lin, Shiwei, et al.
Published: (2025)
Self-Evolving Visual Concept Library using Vision-Language Critics
by: Sehgal, Atharva, et al.
Published: (2025)
by: Sehgal, Atharva, et al.
Published: (2025)
Self-Improving VLM Judges Without Human Annotations
by: Lin, Inna Wanyin, et al.
Published: (2025)
by: Lin, Inna Wanyin, et al.
Published: (2025)
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
by: Zhang, Zaiwei, et al.
Published: (2024)
by: Zhang, Zaiwei, et al.
Published: (2024)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
CogVLM2: Visual Language Models for Image and Video Understanding
by: Hong, Wenyi, et al.
Published: (2024)
by: Hong, Wenyi, et al.
Published: (2024)
RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
by: Hu, Suhang, et al.
Published: (2025)
by: Hu, Suhang, et al.
Published: (2025)
See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
by: Tang, Tianci, et al.
Published: (2026)
by: Tang, Tianci, et al.
Published: (2026)
Similar Items
-
CFSR: Geometry-Conditioned Shadow Removal via Physical Disentanglement
by: Wang, Pan, et al.
Published: (2026) -
UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2026) -
A Multi-Strategy Framework for Enhancing Shatian Pomelo Detection in Real-World Orchards
by: Wang, Pan, et al.
Published: (2025) -
GNC-Pose: Geometry-Aware GNC-PnP for Accurate 6D Pose Estimation
by: Liu, Xiujin
Published: (2025) -
Agent Skills Should Go Beyond Text: The Case for Visual Skills
by: Xu, Binxiao, et al.
Published: (2026)