Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Xiaodan, Zou, Chuhang, Wang, Suchen, Kim, Jaechul, Ahuja, Narendra |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
por: Kim, Yehna, et al.
Publicado: (2025)
por: Kim, Yehna, et al.
Publicado: (2025)
Self-Calibrating 4D Novel View Synthesis from Monocular Videos Using Gaussian Splatting
por: Li, Fang, et al.
Publicado: (2024)
por: Li, Fang, et al.
Publicado: (2024)
Telling Stories for Common Sense Zero-Shot Action Recognition
por: Gowda, Shreyank N, et al.
Publicado: (2023)
por: Gowda, Shreyank N, et al.
Publicado: (2023)
Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
por: Zhang, Hao, et al.
Publicado: (2025)
por: Zhang, Hao, et al.
Publicado: (2025)
S3O: A Dual-Phase Approach for Reconstructing Dynamic Shape and Skeleton of Articulated Objects from Single Monocular Video
por: Zhang, Hao, et al.
Publicado: (2024)
por: Zhang, Hao, et al.
Publicado: (2024)
RGB-Only Supervised Camera Parameter Optimization in Dynamic Scenes
por: Li, Fang, et al.
Publicado: (2025)
por: Li, Fang, et al.
Publicado: (2025)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
por: Zhou, Jiaming, et al.
Publicado: (2024)
por: Zhou, Jiaming, et al.
Publicado: (2024)
Top-Down Framework for Weakly-supervised Grounded Image Captioning
por: Cai, Chen, et al.
Publicado: (2023)
por: Cai, Chen, et al.
Publicado: (2023)
Rethinking Prompting Strategies for Multi-Label Recognition with Partial Annotations
por: Rawlekar, Samyak, et al.
Publicado: (2024)
por: Rawlekar, Samyak, et al.
Publicado: (2024)
Leveraging Temporal Contextualization for Video Action Recognition
por: Kim, Minji, et al.
Publicado: (2024)
por: Kim, Minji, et al.
Publicado: (2024)
HOI-aware Adaptive Network for Weakly-supervised Action Segmentation
por: Zhang, Runzhong, et al.
Publicado: (2026)
por: Zhang, Runzhong, et al.
Publicado: (2026)
Measuring the (Un)Faithfulness of Concept-Based Explanations
por: Kumar, Shubham, et al.
Publicado: (2025)
por: Kumar, Shubham, et al.
Publicado: (2025)
Language Model Guided Interpretable Video Action Reasoning
por: Wang, Ning, et al.
Publicado: (2024)
por: Wang, Ning, et al.
Publicado: (2024)
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
por: Liu, Chuhang, et al.
Publicado: (2026)
por: Liu, Chuhang, et al.
Publicado: (2026)
City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs
por: Dalal, Dwip, et al.
Publicado: (2025)
por: Dalal, Dwip, et al.
Publicado: (2025)
Learning Implicit Representation for Reconstructing Articulated Objects
por: Zhang, Hao, et al.
Publicado: (2024)
por: Zhang, Hao, et al.
Publicado: (2024)
Common Sense Reasoning for Deepfake Detection
por: Zhang, Yue, et al.
Publicado: (2024)
por: Zhang, Yue, et al.
Publicado: (2024)
Collaboratively Self-supervised Video Representation Learning for Action Recognition
por: Zhang, Jie, et al.
Publicado: (2024)
por: Zhang, Jie, et al.
Publicado: (2024)
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
por: Kim, Joowon, et al.
Publicado: (2026)
por: Kim, Joowon, et al.
Publicado: (2026)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
por: Wang, Cong, et al.
Publicado: (2023)
por: Wang, Cong, et al.
Publicado: (2023)
Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
por: Sun, Baoli, et al.
Publicado: (2025)
por: Sun, Baoli, et al.
Publicado: (2025)
Efficiently Disentangling CLIP for Multi-Object Perception
por: Rawlekar, Samyak, et al.
Publicado: (2025)
por: Rawlekar, Samyak, et al.
Publicado: (2025)
MonoPatchNeRF: Improving Neural Radiance Fields with Patch-based Monocular Guidance
por: Wu, Yuqun, et al.
Publicado: (2024)
por: Wu, Yuqun, et al.
Publicado: (2024)
Plenoptic PNG: Real-Time Neural Radiance Fields in 150 KB
por: Lee, Jae Yong, et al.
Publicado: (2024)
por: Lee, Jae Yong, et al.
Publicado: (2024)
Taylor Videos for Action Recognition
por: Wang, Lei, et al.
Publicado: (2024)
por: Wang, Lei, et al.
Publicado: (2024)
PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling
por: Zhang, Hao, et al.
Publicado: (2025)
por: Zhang, Hao, et al.
Publicado: (2025)
MagicPose4D: Crafting Articulated Models with Appearance and Motion Control
por: Zhang, Hao, et al.
Publicado: (2024)
por: Zhang, Hao, et al.
Publicado: (2024)
Video Domain Incremental Learning for Human Action Recognition in Home Environments
por: Hu, Yuanda, et al.
Publicado: (2024)
por: Hu, Yuanda, et al.
Publicado: (2024)
Selective Volume Mixup for Video Action Recognition
por: Tan, Yi, et al.
Publicado: (2023)
por: Tan, Yi, et al.
Publicado: (2023)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
por: Hu, Pengfei, et al.
Publicado: (2025)
por: Hu, Pengfei, et al.
Publicado: (2025)
Piecewise-Linear Manifolds for Deep Metric Learning
por: Bhatnagar, Shubhang, et al.
Publicado: (2024)
por: Bhatnagar, Shubhang, et al.
Publicado: (2024)
ExACT: Language-guided Conceptual Reasoning and Uncertainty Estimation for Event-based Action Recognition and More
por: Zhou, Jiazhou, et al.
Publicado: (2024)
por: Zhou, Jiazhou, et al.
Publicado: (2024)
Exploring Ordinal Bias in Action Recognition for Instructional Videos
por: Kim, Joochan, et al.
Publicado: (2025)
por: Kim, Joochan, et al.
Publicado: (2025)
BridgeIV: Bridging Customized Image and Video Generation through Test-Time Autoregressive Identity Propagation
por: Hu, Panwen, et al.
Publicado: (2025)
por: Hu, Panwen, et al.
Publicado: (2025)
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
por: Lee, Jongseo, et al.
Publicado: (2025)
por: Lee, Jongseo, et al.
Publicado: (2025)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
por: Peng, Jingwei, et al.
Publicado: (2025)
por: Peng, Jingwei, et al.
Publicado: (2025)
CSL: Class-Agnostic Structure-Constrained Learning for Segmentation Including the Unseen
por: Zhang, Hao, et al.
Publicado: (2023)
por: Zhang, Hao, et al.
Publicado: (2023)
Direct and Explicit 3D Generation from a Single Image
por: Wu, Haoyu, et al.
Publicado: (2024)
por: Wu, Haoyu, et al.
Publicado: (2024)
BEAR: A Video Dataset For Fine-grained Behaviors Recognition Oriented with Action and Environment Factors
por: Hu, Chengyang, et al.
Publicado: (2025)
por: Hu, Chengyang, et al.
Publicado: (2025)
HabitAction: A Video Dataset for Human Habitual Behavior Recognition
por: Li, Hongwu, et al.
Publicado: (2024)
por: Li, Hongwu, et al.
Publicado: (2024)
Ejemplares similares
-
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
por: Kim, Yehna, et al.
Publicado: (2025) -
Self-Calibrating 4D Novel View Synthesis from Monocular Videos Using Gaussian Splatting
por: Li, Fang, et al.
Publicado: (2024) -
Telling Stories for Common Sense Zero-Shot Action Recognition
por: Gowda, Shreyank N, et al.
Publicado: (2023) -
Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
por: Zhang, Hao, et al.
Publicado: (2025) -
S3O: A Dual-Phase Approach for Reconstructing Dynamic Shape and Skeleton of Articulated Objects from Single Monocular Video
por: Zhang, Hao, et al.
Publicado: (2024)