Enhancing Video Transformers for Action Understanding with VLM-aided Training
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Hui, Jian, Hu, Poppe, Ronald, Salah, Albert Ali |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Snakes and Ladders: Two Steps Up for VideoMamba
di: Lu, Hui, et al.
Pubblicazione: (2024)
di: Lu, Hui, et al.
Pubblicazione: (2024)
TCNet: Continuous Sign Language Recognition from Trajectories and Correlated Regions
di: Lu, Hui, et al.
Pubblicazione: (2024)
di: Lu, Hui, et al.
Pubblicazione: (2024)
About Time: Advances, Challenges, and Outlooks of Action Understanding
di: Stergiou, Alexandros, et al.
Pubblicazione: (2024)
di: Stergiou, Alexandros, et al.
Pubblicazione: (2024)
Representation Learning and Identity Adversarial Training for Facial Behavior Understanding
di: Ning, Mang, et al.
Pubblicazione: (2024)
di: Ning, Mang, et al.
Pubblicazione: (2024)
Self-Supervised Partial Cycle-Consistency for Multi-View Matching
di: Taggenbrock, Fedor, et al.
Pubblicazione: (2025)
di: Taggenbrock, Fedor, et al.
Pubblicazione: (2025)
Benchmarking and Enhancing VLM for Compressed Image Understanding
di: Zhang, Zifu, et al.
Pubblicazione: (2025)
di: Zhang, Zifu, et al.
Pubblicazione: (2025)
The Role of Video Generation in Enhancing Data-Limited Action Understanding
di: Li, Wei, et al.
Pubblicazione: (2025)
di: Li, Wei, et al.
Pubblicazione: (2025)
CogVLM2: Visual Language Models for Image and Video Understanding
di: Hong, Wenyi, et al.
Pubblicazione: (2024)
di: Hong, Wenyi, et al.
Pubblicazione: (2024)
Domain Adaptation of VLM for Soccer Video Understanding
di: Jiang, Tiancheng, et al.
Pubblicazione: (2025)
di: Jiang, Tiancheng, et al.
Pubblicazione: (2025)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
di: Xu, Ruyi, et al.
Pubblicazione: (2025)
di: Xu, Ruyi, et al.
Pubblicazione: (2025)
SVFormer: A Direct Training Spiking Transformer for Efficient Video Action Recognition
di: Yu, Liutao, et al.
Pubblicazione: (2024)
di: Yu, Liutao, et al.
Pubblicazione: (2024)
LongVLM: Efficient Long Video Understanding via Large Language Models
di: Weng, Yuetian, et al.
Pubblicazione: (2024)
di: Weng, Yuetian, et al.
Pubblicazione: (2024)
UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
di: Bai, Chengyu, et al.
Pubblicazione: (2025)
di: Bai, Chengyu, et al.
Pubblicazione: (2025)
EgoVLM: Policy Optimization for Egocentric Video Understanding
di: Vinod, Ashwin, et al.
Pubblicazione: (2025)
di: Vinod, Ashwin, et al.
Pubblicazione: (2025)
Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM
di: Yang, Sihan, et al.
Pubblicazione: (2025)
di: Yang, Sihan, et al.
Pubblicazione: (2025)
HieroAction: Hierarchically Guided VLM for Fine-Grained Action Analysis
di: Wu, Junhao, et al.
Pubblicazione: (2025)
di: Wu, Junhao, et al.
Pubblicazione: (2025)
ScVLM: Enhancing Vision-Language Model for Safety-Critical Event Understanding
di: Shi, Liang, et al.
Pubblicazione: (2024)
di: Shi, Liang, et al.
Pubblicazione: (2024)
Slot-VLM: SlowFast Slots for Video-Language Modeling
di: Xu, Jiaqi, et al.
Pubblicazione: (2024)
di: Xu, Jiaqi, et al.
Pubblicazione: (2024)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
di: Bae, Kyungho, et al.
Pubblicazione: (2025)
di: Bae, Kyungho, et al.
Pubblicazione: (2025)
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
di: Xue, Xizhe, et al.
Pubblicazione: (2024)
di: Xue, Xizhe, et al.
Pubblicazione: (2024)
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
di: Ning, Zhenyu, et al.
Pubblicazione: (2025)
di: Ning, Zhenyu, et al.
Pubblicazione: (2025)
ASTRA: An Action Spotting TRAnsformer for Soccer Videos
di: Xarles, Artur, et al.
Pubblicazione: (2024)
di: Xarles, Artur, et al.
Pubblicazione: (2024)
TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference
di: Hu, Junshan, et al.
Pubblicazione: (2025)
di: Hu, Junshan, et al.
Pubblicazione: (2025)
Dark Transformer: A Video Transformer for Action Recognition in the Dark
di: Ulhaq, Anwaar
Pubblicazione: (2024)
di: Ulhaq, Anwaar
Pubblicazione: (2024)
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
di: Wang, Ruoyu, et al.
Pubblicazione: (2024)
di: Wang, Ruoyu, et al.
Pubblicazione: (2024)
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
di: Huang, Zhipeng, et al.
Pubblicazione: (2024)
di: Huang, Zhipeng, et al.
Pubblicazione: (2024)
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
di: Fateh, Fawad Javed, et al.
Pubblicazione: (2024)
di: Fateh, Fawad Javed, et al.
Pubblicazione: (2024)
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
di: Bao, Wentao, et al.
Pubblicazione: (2024)
di: Bao, Wentao, et al.
Pubblicazione: (2024)
CoS: Chain-of-Shot Prompting for Long Video Understanding
di: Hu, Jian, et al.
Pubblicazione: (2025)
di: Hu, Jian, et al.
Pubblicazione: (2025)
Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$
di: Zhang, Jiangning, et al.
Pubblicazione: (2025)
di: Zhang, Jiangning, et al.
Pubblicazione: (2025)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
di: Tian, Rui, et al.
Pubblicazione: (2025)
di: Tian, Rui, et al.
Pubblicazione: (2025)
TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation
di: Liu, Jiaxing, et al.
Pubblicazione: (2026)
di: Liu, Jiaxing, et al.
Pubblicazione: (2026)
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
di: Lin, Hangyu, et al.
Pubblicazione: (2026)
di: Lin, Hangyu, et al.
Pubblicazione: (2026)
Enhanced Partially Relevant Video Retrieval through Inter- and Intra-Sample Analysis with Coherence Prediction
di: Ren, Junlong, et al.
Pubblicazione: (2025)
di: Ren, Junlong, et al.
Pubblicazione: (2025)
EarlyTom: Early Token Compression Completes Fast Video Understanding
di: Wang, Hesong, et al.
Pubblicazione: (2026)
di: Wang, Hesong, et al.
Pubblicazione: (2026)
Versatile Editing of Video Content, Actions, and Dynamics without Training
di: Kulikov, Vladimir, et al.
Pubblicazione: (2026)
di: Kulikov, Vladimir, et al.
Pubblicazione: (2026)
E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs
di: Liu, Xianjie, et al.
Pubblicazione: (2026)
di: Liu, Xianjie, et al.
Pubblicazione: (2026)
ROOT: VLM based System for Indoor Scene Understanding and Beyond
di: Wang, Yonghui, et al.
Pubblicazione: (2024)
di: Wang, Yonghui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Snakes and Ladders: Two Steps Up for VideoMamba
di: Lu, Hui, et al.
Pubblicazione: (2024) -
TCNet: Continuous Sign Language Recognition from Trajectories and Correlated Regions
di: Lu, Hui, et al.
Pubblicazione: (2024) -
About Time: Advances, Challenges, and Outlooks of Action Understanding
di: Stergiou, Alexandros, et al.
Pubblicazione: (2024) -
Representation Learning and Identity Adversarial Training for Facial Behavior Understanding
di: Ning, Mang, et al.
Pubblicazione: (2024) -
Self-Supervised Partial Cycle-Consistency for Multi-View Matching
di: Taggenbrock, Fedor, et al.
Pubblicazione: (2025)