Agent Skills Should Go Beyond Text: The Case for Visual Skills
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Binxiao, An, Ruichuan, Zou, Bocheng, Hua, Hang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
von: Zhang, Qizhe, et al.
Veröffentlicht: (2023)
von: Zhang, Qizhe, et al.
Veröffentlicht: (2023)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model
von: Luo, Yulin, et al.
Veröffentlicht: (2024)
von: Luo, Yulin, et al.
Veröffentlicht: (2024)
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
von: Wang, Pan, et al.
Veröffentlicht: (2026)
von: Wang, Pan, et al.
Veröffentlicht: (2026)
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
von: Lin, Weifeng, et al.
Veröffentlicht: (2024)
von: Lin, Weifeng, et al.
Veröffentlicht: (2024)
AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents
von: Cui, Jieming, et al.
Veröffentlicht: (2024)
von: Cui, Jieming, et al.
Veröffentlicht: (2024)
Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2026)
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2026)
GEMS: Agent-Native Multimodal Generation with Memory and Skills
von: He, Zefeng, et al.
Veröffentlicht: (2026)
von: He, Zefeng, et al.
Veröffentlicht: (2026)
ProSkill: Segment-Level Skill Assessment in Procedural Videos
von: Mazzamuto, Michele, et al.
Veröffentlicht: (2026)
von: Mazzamuto, Michele, et al.
Veröffentlicht: (2026)
SkillSight: Efficient First-Person Skill Assessment with Gaze
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
von: Li, Puhao, et al.
Veröffentlicht: (2024)
von: Li, Puhao, et al.
Veröffentlicht: (2024)
SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
Decomposing Densification in Gaussian Splatting for Faster 3D Scene Reconstruction
von: Huang, Binxiao, et al.
Veröffentlicht: (2025)
von: Huang, Binxiao, et al.
Veröffentlicht: (2025)
SportSkills: Physical Skill Learning from Sports Instructional Videos
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
Affordance Agent Harness: Verification-Gated Skill Orchestration
von: Huang, Haojian, et al.
Veröffentlicht: (2026)
von: Huang, Haojian, et al.
Veröffentlicht: (2026)
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
ModSkill: Physical Character Skill Modularization
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
Skill-Conditioned Visual Geolocation for Vision-Language Models
von: Yang, Chenjie, et al.
Veröffentlicht: (2026)
von: Yang, Chenjie, et al.
Veröffentlicht: (2026)
SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
von: Dai, Gaole, et al.
Veröffentlicht: (2025)
von: Dai, Gaole, et al.
Veröffentlicht: (2025)
MME-CoF-Pro: Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints
von: Qi, Yu, et al.
Veröffentlicht: (2026)
von: Qi, Yu, et al.
Veröffentlicht: (2026)
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
von: Kim, Hanjung, et al.
Veröffentlicht: (2025)
von: Kim, Hanjung, et al.
Veröffentlicht: (2025)
Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
von: Chae, Hyunsik, et al.
Veröffentlicht: (2025)
von: Chae, Hyunsik, et al.
Veröffentlicht: (2025)
Token Coordinated Prompt Attention is Needed for Visual Prompting
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
Calisthenics Skills Temporal Video Segmentation
von: Finocchiaro, Antonio, et al.
Veröffentlicht: (2025)
von: Finocchiaro, Antonio, et al.
Veröffentlicht: (2025)
UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
Beyond Flat Text: Dual Self-inherited Guidance for Visual Text Generation
von: Luo, Minxing, et al.
Veröffentlicht: (2025)
von: Luo, Minxing, et al.
Veröffentlicht: (2025)
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
von: Zeng, Ziyun, et al.
Veröffentlicht: (2025)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2025)
Beyond Visual Field of View: Perceiving 3D Environment with Echoes and Vision
von: Zhu, Lingyu, et al.
Veröffentlicht: (2022)
von: Zhu, Lingyu, et al.
Veröffentlicht: (2022)
Learning Skill-Attributes for Transferable Assessment in Video
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2025)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2025)
Bridging the Skill Gap in Clinical CBCT Interpretation with CBCTRepD
von: Wu, Qinxin, et al.
Veröffentlicht: (2026)
von: Wu, Qinxin, et al.
Veröffentlicht: (2026)
AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots
von: Zhang, Likui, et al.
Veröffentlicht: (2026)
von: Zhang, Likui, et al.
Veröffentlicht: (2026)
SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation
von: Ren, Tianfei, et al.
Veröffentlicht: (2026)
von: Ren, Tianfei, et al.
Veröffentlicht: (2026)
Learning Human Skill Generators at Key-Step Levels
von: Wu, Yilu, et al.
Veröffentlicht: (2025)
von: Wu, Yilu, et al.
Veröffentlicht: (2025)
FunBench: Benchmarking Fundus Reading Skills of MLLMs
von: Wei, Qijie, et al.
Veröffentlicht: (2025)
von: Wei, Qijie, et al.
Veröffentlicht: (2025)
Beyond Text: Frozen Large Language Models in Visual Signal Comprehension
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
von: Bao, Fan, et al.
Veröffentlicht: (2024)
von: Bao, Fan, et al.
Veröffentlicht: (2024)
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
von: Li, Bozhou, et al.
Veröffentlicht: (2025)
von: Li, Bozhou, et al.
Veröffentlicht: (2025)
Quantum-enhanced Computer Vision: Going Beyond Classical Algorithms
von: Meli, Natacha Kuete, et al.
Veröffentlicht: (2025)
von: Meli, Natacha Kuete, et al.
Veröffentlicht: (2025)
MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models
von: Zou, Bocheng, et al.
Veröffentlicht: (2026)
von: Zou, Bocheng, et al.
Veröffentlicht: (2026)
AI-Driven Evaluation of Surgical Skill via Action Recognition
von: Meng, Yan, et al.
Veröffentlicht: (2025)
von: Meng, Yan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
von: Zhang, Qizhe, et al.
Veröffentlicht: (2023) -
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026) -
LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model
von: Luo, Yulin, et al.
Veröffentlicht: (2024) -
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
von: Wang, Pan, et al.
Veröffentlicht: (2026) -
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
von: Lin, Weifeng, et al.
Veröffentlicht: (2024)