Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Zhenlong, Qu, Xiangyan, Qian, Chengxuan, Chen, Rui, Tang, Jing, Sun, Lei, Chu, Xiangxiang, Zhang, Dapeng, Wang, Yiwei, Cai, Yujun, Li, Shuo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2026)
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2026)
AutoDrive-R$^2$: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2025)
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2025)
From Scale to Speed: Adaptive Test-Time Scaling for Image Editing
von: Qu, Xiangyan, et al.
Veröffentlicht: (2026)
von: Qu, Xiangyan, et al.
Veröffentlicht: (2026)
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
von: Su, Qile, et al.
Veröffentlicht: (2026)
von: Su, Qile, et al.
Veröffentlicht: (2026)
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
von: Tan, Rongbin, et al.
Veröffentlicht: (2026)
von: Tan, Rongbin, et al.
Veröffentlicht: (2026)
FingER: Content Aware Fine-grained Evaluation with Reasoning for AI-Generated Videos
von: Chen, Rui, et al.
Veröffentlicht: (2025)
von: Chen, Rui, et al.
Veröffentlicht: (2025)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
Pure Vision Language Action (VLA) Models: A Comprehensive Survey
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
DENOISER: Rethinking the Robustness for Open-Vocabulary Action Recognition
von: Cheng, Haozhe, et al.
Veröffentlicht: (2024)
von: Cheng, Haozhe, et al.
Veröffentlicht: (2024)
LLMs are Good Action Recognizers
von: Qu, Haoxuan, et al.
Veröffentlicht: (2024)
von: Qu, Haoxuan, et al.
Veröffentlicht: (2024)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
Learning to Generalize without Bias for Open-Vocabulary Action Recognition
von: Yu, Yating, et al.
Veröffentlicht: (2025)
von: Yu, Yating, et al.
Veröffentlicht: (2025)
HyLiFormer: Hyperbolic Linear Attention for Skeleton-based Human Action Recognition
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
Taylor Videos for Action Recognition
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
FactGuard: Agentic Video Misinformation Detection via Reinforcement Learning
von: Li, Zehao, et al.
Veröffentlicht: (2026)
von: Li, Zehao, et al.
Veröffentlicht: (2026)
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
von: So, Yerim, et al.
Veröffentlicht: (2026)
von: So, Yerim, et al.
Veröffentlicht: (2026)
Open-Vocabulary Video Anomaly Detection
von: Wu, Peng, et al.
Veröffentlicht: (2023)
von: Wu, Peng, et al.
Veröffentlicht: (2023)
Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition
von: Lian, Zheng, et al.
Veröffentlicht: (2026)
von: Lian, Zheng, et al.
Veröffentlicht: (2026)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
Understanding GUI Agent Localization Biases through Logit Sharpness
von: Tao, Xingjian, et al.
Veröffentlicht: (2025)
von: Tao, Xingjian, et al.
Veröffentlicht: (2025)
Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' Memory
von: Tao, Xingjian, et al.
Veröffentlicht: (2024)
von: Tao, Xingjian, et al.
Veröffentlicht: (2024)
FROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
AffectGPT-R1: Leveraging Reinforcement Learning for Open-Vocabulary Multimodal Emotion Recognition
von: Lian, Zheng, et al.
Veröffentlicht: (2025)
von: Lian, Zheng, et al.
Veröffentlicht: (2025)
AuTAgent: A Reinforcement Learning Framework for Tool-Augmented Audio Reasoning
von: Tong, Siqian, et al.
Veröffentlicht: (2026)
von: Tong, Siqian, et al.
Veröffentlicht: (2026)
GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation
von: Li, Sifan, et al.
Veröffentlicht: (2026)
von: Li, Sifan, et al.
Veröffentlicht: (2026)
Adaptive Label Correction for Robust Medical Image Segmentation with Noisy Labels
von: Qian, Chengxuan, et al.
Veröffentlicht: (2025)
von: Qian, Chengxuan, et al.
Veröffentlicht: (2025)
SED-MVS: Segmentation-Driven and Edge-Aligned Deformation Multi-View Stereo with Depth Restoration and Occlusion Constraint
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2025)
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2025)
Geometry-Guided Reinforcement Learning for Multi-view Consistent 3D Scene Editing
von: Wang, Jiyuan, et al.
Veröffentlicht: (2026)
von: Wang, Jiyuan, et al.
Veröffentlicht: (2026)
Scaling Open-Vocabulary Action Detection
von: Sia, Zhen Hao, et al.
Veröffentlicht: (2025)
von: Sia, Zhen Hao, et al.
Veröffentlicht: (2025)
DVP-MVS++: Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View Stereo
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2025)
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2025)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
von: Bai, Sule, et al.
Veröffentlicht: (2025)
von: Bai, Sule, et al.
Veröffentlicht: (2025)
Follow the Clues, Frame the Truth: Hybrid-evidential Deductive Reasoning in Open-Vocabulary Multimodal Emotion Recognition
von: Liu, Yu, et al.
Veröffentlicht: (2026)
von: Liu, Yu, et al.
Veröffentlicht: (2026)
EMPOWER: Evolutionary Medical Prompt Optimization With Reinforcement Learning
von: Chen, Yinda, et al.
Veröffentlicht: (2025)
von: Chen, Yinda, et al.
Veröffentlicht: (2025)
VOVTrack: Exploring the Potentiality in Videos for Open-Vocabulary Object Tracking
von: Qian, Zekun, et al.
Veröffentlicht: (2024)
von: Qian, Zekun, et al.
Veröffentlicht: (2024)
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
von: Cui, Yu, et al.
Veröffentlicht: (2025)
von: Cui, Yu, et al.
Veröffentlicht: (2025)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
von: Li, Sifan, et al.
Veröffentlicht: (2025)
von: Li, Sifan, et al.
Veröffentlicht: (2025)
Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs
von: Dan, Nifu, et al.
Veröffentlicht: (2025)
von: Dan, Nifu, et al.
Veröffentlicht: (2025)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
von: Wu, Yike, et al.
Veröffentlicht: (2025)
von: Wu, Yike, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2026) -
AutoDrive-R$^2$: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2025) -
From Scale to Speed: Adaptive Test-Time Scaling for Image Editing
von: Qu, Xiangyan, et al.
Veröffentlicht: (2026) -
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
von: Su, Qile, et al.
Veröffentlicht: (2026) -
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
von: Tan, Rongbin, et al.
Veröffentlicht: (2026)