ExAct: A Video-Language Benchmark for Expert Action Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yi, Han, Pan, Yulu, He, Feihong, Liu, Xinyu, Zhang, Benjamin, Oguntola, Oluwatumininu, Bertasius, Gedas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BASKET: A Large-Scale Video Dataset for Fine-Grained Skill Estimation
von: Pan, Yulu, et al.
Veröffentlicht: (2025)
von: Pan, Yulu, et al.
Veröffentlicht: (2025)
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
von: Pan, Yulu, et al.
Veröffentlicht: (2026)
von: Pan, Yulu, et al.
Veröffentlicht: (2026)
SiLVR: A Simple Language-based Video Reasoning Framework
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2023)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2023)
Siamese Vision Transformers are Scalable Audio-visual Learners
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion
von: Tursynbek, Nurislam, et al.
Veröffentlicht: (2026)
von: Tursynbek, Nurislam, et al.
Veröffentlicht: (2026)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
LoCoNet: Long-Short Context Network for Active Speaker Detection
von: Wang, Xizi, et al.
Veröffentlicht: (2023)
von: Wang, Xizi, et al.
Veröffentlicht: (2023)
Video ReCap: Recursive Captioning of Hour-Long Videos
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
von: Li, Baiqi, et al.
Veröffentlicht: (2026)
von: Li, Baiqi, et al.
Veröffentlicht: (2026)
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
BOSS: Benchmark for Observation Space Shift in Long-Horizon Task
von: Yang, Yue, et al.
Veröffentlicht: (2025)
von: Yang, Yue, et al.
Veröffentlicht: (2025)
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
A Simple LLM Framework for Long-Range Video Question-Answering
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2024)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
von: Dong, Shaoqi, et al.
Veröffentlicht: (2025)
von: Dong, Shaoqi, et al.
Veröffentlicht: (2025)
ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis
von: Fang, Yu, et al.
Veröffentlicht: (2025)
von: Fang, Yu, et al.
Veröffentlicht: (2025)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
DAM: Dynamic Adapter Merging for Continual Video QA Learning
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2026)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2026)
AesRM: Improving Video Aesthetics with Expert-Level Feedback
von: Han, Yujin, et al.
Veröffentlicht: (2026)
von: Han, Yujin, et al.
Veröffentlicht: (2026)
TimeRefine: Temporal Grounding with Time Refining Video LLM
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
von: Peng, Ziqiao, et al.
Veröffentlicht: (2025)
von: Peng, Ziqiao, et al.
Veröffentlicht: (2025)
PrototypeFormer: Learning to Explore Prototype Relationships for Few-shot Image Classification
von: Su, Meijuan, et al.
Veröffentlicht: (2023)
von: Su, Meijuan, et al.
Veröffentlicht: (2023)
FineDiffusion: Scaling up Diffusion Models for Fine-grained Image Generation with 10,000 Classes
von: Pan, Ziying, et al.
Veröffentlicht: (2024)
von: Pan, Ziying, et al.
Veröffentlicht: (2024)
COPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection
von: Jacob, Darryl Cherian, et al.
Veröffentlicht: (2026)
von: Jacob, Darryl Cherian, et al.
Veröffentlicht: (2026)
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
von: Wang, Yubin, et al.
Veröffentlicht: (2024)
von: Wang, Yubin, et al.
Veröffentlicht: (2024)
InstrAct: Towards Action-Centric Understanding in Instructional Videos
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2026)
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2026)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
von: Ling, Yiran, et al.
Veröffentlicht: (2026)
von: Ling, Yiran, et al.
Veröffentlicht: (2026)
ActAnywhere: Subject-Aware Video Background Generation
von: Pan, Boxiao, et al.
Veröffentlicht: (2024)
von: Pan, Boxiao, et al.
Veröffentlicht: (2024)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
Kronecker Mask and Interpretive Prompts are Language-Action Video Learners
von: Yang, Jingyi, et al.
Veröffentlicht: (2025)
von: Yang, Jingyi, et al.
Veröffentlicht: (2025)
LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies
von: Yang, Yue, et al.
Veröffentlicht: (2026)
von: Yang, Yue, et al.
Veröffentlicht: (2026)
Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models
von: Tan, Xudong, et al.
Veröffentlicht: (2025)
von: Tan, Xudong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BASKET: A Large-Scale Video Dataset for Fine-Grained Skill Estimation
von: Pan, Yulu, et al.
Veröffentlicht: (2025) -
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
von: Pan, Yulu, et al.
Veröffentlicht: (2026) -
SiLVR: A Simple Language-based Video Reasoning Framework
von: Zhang, Ce, et al.
Veröffentlicht: (2025) -
RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2023) -
Siamese Vision Transformers are Scalable Audio-visual Learners
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)