Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Liyang, Zhu, Sihan, Guo, Yunjie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023)
by: Yang, Min, et al.
Published: (2023)
MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos
by: Tang, Yin, et al.
Published: (2024)
by: Tang, Yin, et al.
Published: (2024)
Multi-Stage Boundary-Aware Transformer Network for Action Segmentation in Untrimmed Surgical Videos
by: Shuvo, Rezowan, et al.
Published: (2025)
by: Shuvo, Rezowan, et al.
Published: (2025)
MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
by: Feng, Yue, et al.
Published: (2025)
by: Feng, Yue, et al.
Published: (2025)
Localizing Moments of Actions in Untrimmed Videos of Infants with Autism Spectrum Disorder
by: Helvaci, Halil Ismail, et al.
Published: (2024)
by: Helvaci, Halil Ismail, et al.
Published: (2024)
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
by: Chen, Brian, et al.
Published: (2023)
by: Chen, Brian, et al.
Published: (2023)
MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
by: Feng, Yue, et al.
Published: (2026)
by: Feng, Yue, et al.
Published: (2026)
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos
by: Li, Kaining, et al.
Published: (2025)
by: Li, Kaining, et al.
Published: (2025)
LVLM-empowered Multi-modal Representation Learning for Visual Place Recognition
by: Wang, Teng, et al.
Published: (2024)
by: Wang, Teng, et al.
Published: (2024)
ContextGuard-LVLM: Enhancing News Veracity through Fine-grained Cross-modal Contextual Consistency Verification
by: Ma, Sihan, et al.
Published: (2025)
by: Ma, Sihan, et al.
Published: (2025)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
by: Tang, Yolo Yunlong, et al.
Published: (2024)
by: Tang, Yolo Yunlong, et al.
Published: (2024)
Cutup and Detect: Human Fall Detection on Cutup Untrimmed Videos Using a Large Foundational Video Understanding Model
by: Grutschus, Till, et al.
Published: (2024)
by: Grutschus, Till, et al.
Published: (2024)
Mamba-OTR: a Mamba-based Solution for Online Take and Release Detection from Untrimmed Egocentric Video
by: Catinello, Alessandro Sebastiano, et al.
Published: (2025)
by: Catinello, Alessandro Sebastiano, et al.
Published: (2025)
LAVID: An Agentic LVLM Framework for Diffusion-Generated Video Detection
by: Liu, Qingyuan, et al.
Published: (2025)
by: Liu, Qingyuan, et al.
Published: (2025)
AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
by: Ahn, Sunghyun, et al.
Published: (2025)
by: Ahn, Sunghyun, et al.
Published: (2025)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Frequency Guidance Matters: Skeletal Action Recognition by Frequency-Aware Mixed Transformer
by: Wu, Wenhan, et al.
Published: (2024)
by: Wu, Wenhan, et al.
Published: (2024)
SBF: An Effective Representation to Augment Skeleton for Video-based Human Action Recognition
by: Peng, Zhuoxuan, et al.
Published: (2026)
by: Peng, Zhuoxuan, et al.
Published: (2026)
Video Domain Incremental Learning for Human Action Recognition in Home Environments
by: Hu, Yuanda, et al.
Published: (2024)
by: Hu, Yuanda, et al.
Published: (2024)
MVP-Shot: Multi-Velocity Progressive-Alignment Framework for Few-Shot Action Recognition
by: Qu, Hongyu, et al.
Published: (2024)
by: Qu, Hongyu, et al.
Published: (2024)
Selective Volume Mixup for Video Action Recognition
by: Tan, Yi, et al.
Published: (2023)
by: Tan, Yi, et al.
Published: (2023)
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
VALD: Multi-Stage Vision Attack Detection for Efficient LVLM Defense
by: Kadvil, Nadav, et al.
Published: (2026)
by: Kadvil, Nadav, et al.
Published: (2026)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
by: Peng, Jingwei, et al.
Published: (2025)
by: Peng, Jingwei, et al.
Published: (2025)
Action Selection Learning for Multi-label Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Pose-Aware Multi-Level Motion Parsing for Action Quality Assessment
by: Zhu, Shuaikang, et al.
Published: (2025)
by: Zhu, Shuaikang, et al.
Published: (2025)
Taylor Videos for Action Recognition
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Multi-view Distillation based on Multi-modal Fusion for Few-shot Action Recognition(CLIP-$\mathrm{M^2}$DF)
by: Guo, Fei, et al.
Published: (2024)
by: Guo, Fei, et al.
Published: (2024)
Video-to-Task Learning via Motion-Guided Attention for Few-Shot Action Recognition
by: Guo, Hanyu, et al.
Published: (2024)
by: Guo, Hanyu, et al.
Published: (2024)
M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition
by: Wang, Mengmeng, et al.
Published: (2024)
by: Wang, Mengmeng, et al.
Published: (2024)
MMAD: Multi-label Micro-Action Detection in Videos
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers
by: Ma, Ji, et al.
Published: (2025)
by: Ma, Ji, et al.
Published: (2025)
Zero-Shot Action Recognition in Surveillance Videos
by: Pereira, Joao, et al.
Published: (2024)
by: Pereira, Joao, et al.
Published: (2024)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
MedM-VL: What Makes a Good Medical LVLM?
by: Shi, Yiming, et al.
Published: (2025)
by: Shi, Yiming, et al.
Published: (2025)
Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
by: Sun, Baoli, et al.
Published: (2025)
by: Sun, Baoli, et al.
Published: (2025)
LVLM-Composer's Explicit Planning for Image Generation
by: Ramsey, Spencer, et al.
Published: (2025)
by: Ramsey, Spencer, et al.
Published: (2025)
Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning
by: Chen, Cheng, et al.
Published: (2025)
by: Chen, Cheng, et al.
Published: (2025)
Similar Items
-
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023) -
MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos
by: Tang, Yin, et al.
Published: (2024) -
Multi-Stage Boundary-Aware Transformer Network for Action Segmentation in Untrimmed Surgical Videos
by: Shuvo, Rezowan, et al.
Published: (2025) -
MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
by: Feng, Yue, et al.
Published: (2025) -
Localizing Moments of Actions in Untrimmed Videos of Infants with Autism Spectrum Disorder
by: Helvaci, Halil Ismail, et al.
Published: (2024)