Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Yehna, Kim, Young-Eun, Lee, Seong-Whan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ID-EA: Identity-driven Text Enhancement and Adaptation with Textual Inversion for Personalized Text-to-Image Generation
di: Jin, Hyun-Jun, et al.
Pubblicazione: (2025)
di: Jin, Hyun-Jun, et al.
Pubblicazione: (2025)
FAR-Net: Multi-Stage Fusion Network with Enhanced Semantic Alignment and Adaptive Reconciliation for Composed Image Retrieval
di: Park, Jeong-Woo, et al.
Pubblicazione: (2025)
di: Park, Jeong-Woo, et al.
Pubblicazione: (2025)
DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer
di: Kim, Ho-Joong, et al.
Pubblicazione: (2025)
di: Kim, Ho-Joong, et al.
Pubblicazione: (2025)
TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression
di: Kim, Ho-Joong, et al.
Pubblicazione: (2024)
di: Kim, Ho-Joong, et al.
Pubblicazione: (2024)
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition
di: Hu, Xiaodan, et al.
Pubblicazione: (2025)
di: Hu, Xiaodan, et al.
Pubblicazione: (2025)
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
di: So, Yerim, et al.
Pubblicazione: (2026)
di: So, Yerim, et al.
Pubblicazione: (2026)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
di: Lee, SuBeen, et al.
Pubblicazione: (2025)
di: Lee, SuBeen, et al.
Pubblicazione: (2025)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
di: Kim, Ji-Hyeon, et al.
Pubblicazione: (2026)
di: Kim, Ji-Hyeon, et al.
Pubblicazione: (2026)
Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
di: Hong, Jung-Ho, et al.
Pubblicazione: (2025)
di: Hong, Jung-Ho, et al.
Pubblicazione: (2025)
Illuminating Salient Contributions in Neuron Activation with Attribution Equilibrium
di: Nam, Woo-Jeoung, et al.
Pubblicazione: (2022)
di: Nam, Woo-Jeoung, et al.
Pubblicazione: (2022)
AM-SORT: Adaptable Motion Predictor with Historical Trajectory Embedding for Multi-Object Tracking
di: Kim, Vitaliy, et al.
Pubblicazione: (2024)
di: Kim, Vitaliy, et al.
Pubblicazione: (2024)
D$^2$ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action Recognition
di: Pei, Wenjie, et al.
Pubblicazione: (2023)
di: Pei, Wenjie, et al.
Pubblicazione: (2023)
MCoT-RE: Multi-Faceted Chain-of-Thought and Re-Ranking for Training-Free Zero-Shot Composed Image Retrieval
di: Park, Jeong-Woo, et al.
Pubblicazione: (2025)
di: Park, Jeong-Woo, et al.
Pubblicazione: (2025)
Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition
di: Do, Jeonghyeok, et al.
Pubblicazione: (2024)
di: Do, Jeonghyeok, et al.
Pubblicazione: (2024)
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation
di: Woo, Young Beom, et al.
Pubblicazione: (2025)
di: Woo, Young Beom, et al.
Pubblicazione: (2025)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
di: Oh, Ju-Young, et al.
Pubblicazione: (2025)
di: Oh, Ju-Young, et al.
Pubblicazione: (2025)
Leveraging Temporal Contextualization for Video Action Recognition
di: Kim, Minji, et al.
Pubblicazione: (2024)
di: Kim, Minji, et al.
Pubblicazione: (2024)
Spatio-Temporal Proximity-Aware Dual-Path Model for Panoramic Activity Recognition
di: Lee, Sumin, et al.
Pubblicazione: (2024)
di: Lee, Sumin, et al.
Pubblicazione: (2024)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
di: Rajendiran, Ramanathan, et al.
Pubblicazione: (2023)
di: Rajendiran, Ramanathan, et al.
Pubblicazione: (2023)
Part-aware Unified Representation of Language and Skeleton for Zero-shot Action Recognition
di: Zhu, Anqi, et al.
Pubblicazione: (2024)
di: Zhu, Anqi, et al.
Pubblicazione: (2024)
OAD-Promoter: Enhancing Zero-shot VQA using Large Language Models with Object Attribute Description
di: Xu, Quanxing, et al.
Pubblicazione: (2025)
di: Xu, Quanxing, et al.
Pubblicazione: (2025)
Appearance Debiased Gaze Estimation via Stochastic Subject-Wise Adversarial Learning
di: Kim, Suneung, et al.
Pubblicazione: (2024)
di: Kim, Suneung, et al.
Pubblicazione: (2024)
PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition
di: Lee, Jongseo, et al.
Pubblicazione: (2025)
di: Lee, Jongseo, et al.
Pubblicazione: (2025)
PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition
di: Lee, Sanghyeon, et al.
Pubblicazione: (2026)
di: Lee, Sanghyeon, et al.
Pubblicazione: (2026)
Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation
di: Choi, Sun-Hyuk, et al.
Pubblicazione: (2025)
di: Choi, Sun-Hyuk, et al.
Pubblicazione: (2025)
Zero-shot Compositional Action Recognition with Neural Logic Constraints
di: Ye, Gefan, et al.
Pubblicazione: (2025)
di: Ye, Gefan, et al.
Pubblicazione: (2025)
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
di: Jon, Hyo Jin, et al.
Pubblicazione: (2026)
di: Jon, Hyo Jin, et al.
Pubblicazione: (2026)
SurgX: Neuron-Concept Association for Explainable Surgical Phase Recognition
di: Kim, Ka Young, et al.
Pubblicazione: (2025)
di: Kim, Ka Young, et al.
Pubblicazione: (2025)
Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition
di: Jung, Gunho, et al.
Pubblicazione: (2025)
di: Jung, Gunho, et al.
Pubblicazione: (2025)
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
di: Lee, Jongseo, et al.
Pubblicazione: (2025)
di: Lee, Jongseo, et al.
Pubblicazione: (2025)
SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition
di: Do, Jeonghyeok, et al.
Pubblicazione: (2024)
di: Do, Jeonghyeok, et al.
Pubblicazione: (2024)
TIFu: Tri-directional Implicit Function for High-Fidelity 3D Character Reconstruction
di: Lim, Byoungsung, et al.
Pubblicazione: (2024)
di: Lim, Byoungsung, et al.
Pubblicazione: (2024)
Adaptive Prototype Model for Attribute-based Multi-label Few-shot Action Recognition
di: Xiao, Juefeng, et al.
Pubblicazione: (2025)
di: Xiao, Juefeng, et al.
Pubblicazione: (2025)
Unified Language-driven Zero-shot Domain Adaptation
di: Yang, Senqiao, et al.
Pubblicazione: (2024)
di: Yang, Senqiao, et al.
Pubblicazione: (2024)
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
di: Yamane, Taiga, et al.
Pubblicazione: (2025)
di: Yamane, Taiga, et al.
Pubblicazione: (2025)
Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition
di: Gunasekara, Shanaka Ramesh, et al.
Pubblicazione: (2025)
di: Gunasekara, Shanaka Ramesh, et al.
Pubblicazione: (2025)
Local Representative Token Guided Merging for Text-to-Image Generation
di: Lee, Min-Jeong, et al.
Pubblicazione: (2025)
di: Lee, Min-Jeong, et al.
Pubblicazione: (2025)
Flashback: Memory-Driven Zero-shot, Real-time Video Anomaly Detection
di: Lee, Hyogun, et al.
Pubblicazione: (2025)
di: Lee, Hyogun, et al.
Pubblicazione: (2025)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
di: Huang, Wei-Jhe, et al.
Pubblicazione: (2024)
di: Huang, Wei-Jhe, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ID-EA: Identity-driven Text Enhancement and Adaptation with Textual Inversion for Personalized Text-to-Image Generation
di: Jin, Hyun-Jun, et al.
Pubblicazione: (2025) -
FAR-Net: Multi-Stage Fusion Network with Enhanced Semantic Alignment and Adaptive Reconciliation for Composed Image Retrieval
di: Park, Jeong-Woo, et al.
Pubblicazione: (2025) -
DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer
di: Kim, Ho-Joong, et al.
Pubblicazione: (2025) -
TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression
di: Kim, Ho-Joong, et al.
Pubblicazione: (2024) -
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition
di: Hu, Xiaodan, et al.
Pubblicazione: (2025)