AM Flow: Adapters for Temporal Processing in Action Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Agrawal, Tanay, Ali, Abid, Dantcheva, Antitza, Bremond, Francois |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LAC: Latent Action Composition for Skeleton-based Action Segmentation
von: Yang, Di, et al.
Veröffentlicht: (2023)
von: Yang, Di, et al.
Veröffentlicht: (2023)
LIA-X: Interpretable Latent Portrait Animator
von: Wang, Yaohui, et al.
Veröffentlicht: (2025)
von: Wang, Yaohui, et al.
Veröffentlicht: (2025)
Beyond the Visible: A Survey on Cross-spectral Face Recognition
von: Anghelone, David, et al.
Veröffentlicht: (2022)
von: Anghelone, David, et al.
Veröffentlicht: (2022)
Now You See Me, Now You Don't: A Unified Framework for Expression Consistent Anonymization in Talking Head Videos
von: Egin, Anil, et al.
Veröffentlicht: (2026)
von: Egin, Anil, et al.
Veröffentlicht: (2026)
DenVisCoM: Dense Vision Correspondence Mamba for Efficient and Real-time Optical Flow and Stereo Estimation
von: Anand, Tushar, et al.
Veröffentlicht: (2026)
von: Anand, Tushar, et al.
Veröffentlicht: (2026)
Are Visual-Language Models Effective in Action Recognition? A Comparative Study
von: Ali, Mahmoud, et al.
Veröffentlicht: (2024)
von: Ali, Mahmoud, et al.
Veröffentlicht: (2024)
THEval. Evaluation Framework for Talking Head Video Generation
von: Quignon, Nabyl, et al.
Veröffentlicht: (2025)
von: Quignon, Nabyl, et al.
Veröffentlicht: (2025)
HFNeRF: Learning Human Biomechanic Features with Neural Radiance Fields
von: Dey, Arnab, et al.
Veröffentlicht: (2024)
von: Dey, Arnab, et al.
Veröffentlicht: (2024)
Just Dance with $π$! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection
von: Majhi, Snehashis, et al.
Veröffentlicht: (2025)
von: Majhi, Snehashis, et al.
Veröffentlicht: (2025)
Beyond Real versus Fake Towards Intent-Aware Video Analysis
von: Atreya, Saurabh, et al.
Veröffentlicht: (2025)
von: Atreya, Saurabh, et al.
Veröffentlicht: (2025)
AI killed the video star. Audio-driven diffusion model for expressive talking head generation
von: Chopin, Baptiste, et al.
Veröffentlicht: (2025)
von: Chopin, Baptiste, et al.
Veröffentlicht: (2025)
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
von: Chopin, Baptiste, et al.
Veröffentlicht: (2025)
von: Chopin, Baptiste, et al.
Veröffentlicht: (2025)
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
von: Bora, Maheswar, et al.
Veröffentlicht: (2025)
von: Bora, Maheswar, et al.
Veröffentlicht: (2025)
CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets
von: Agrawal, Tanay, et al.
Veröffentlicht: (2025)
von: Agrawal, Tanay, et al.
Veröffentlicht: (2025)
Introducing Gating and Context into Temporal Action Detection
von: Reka, Aglind, et al.
Veröffentlicht: (2024)
von: Reka, Aglind, et al.
Veröffentlicht: (2024)
Temporally Propagated Masks and Bounding Boxes: Combining the Best of Both Worlds for Multi-Object Tracking
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2024)
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2024)
Loose Social-Interaction Recognition in Real-world Therapy Scenarios
von: Ali, Abid, et al.
Veröffentlicht: (2024)
von: Ali, Abid, et al.
Veröffentlicht: (2024)
D$^2$ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action Recognition
von: Pei, Wenjie, et al.
Veröffentlicht: (2023)
von: Pei, Wenjie, et al.
Veröffentlicht: (2023)
MVP: Multimodal Emotion Recognition based on Video and Physiological Signals
von: Strizhkova, Valeriya, et al.
Veröffentlicht: (2025)
von: Strizhkova, Valeriya, et al.
Veröffentlicht: (2025)
LEO: Generative Latent Image Animator for Human Video Synthesis
von: Wang, Yaohui, et al.
Veröffentlicht: (2023)
von: Wang, Yaohui, et al.
Veröffentlicht: (2023)
Weakly-supervised Autism Severity Assessment in Long Videos
von: Ali, Abid, et al.
Veröffentlicht: (2024)
von: Ali, Abid, et al.
Veröffentlicht: (2024)
TALON: Token-Aligned Lightweight Adapters for 6-DoF Spacecraft Pose Estimation
von: Ali, Abid, et al.
Veröffentlicht: (2026)
von: Ali, Abid, et al.
Veröffentlicht: (2026)
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection
von: Qiu, Yicheng, et al.
Veröffentlicht: (2026)
von: Qiu, Yicheng, et al.
Veröffentlicht: (2026)
B-MoE: A Body-Part-Aware Mixture-of-Experts "All Parts Matter" Approach to Micro-Action Recognition
von: Poddar, Nishit, et al.
Veröffentlicht: (2026)
von: Poddar, Nishit, et al.
Veröffentlicht: (2026)
GHNeRF: Learning Generalizable Human Features with Efficient Neural Radiance Fields
von: Dey, Arnab, et al.
Veröffentlicht: (2024)
von: Dey, Arnab, et al.
Veröffentlicht: (2024)
No Train Yet Gain: Towards Generic Multi-Object Tracking in Sports and Beyond
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2025)
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2025)
Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action Recognition
von: Cao, Congqi, et al.
Veröffentlicht: (2024)
von: Cao, Congqi, et al.
Veröffentlicht: (2024)
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos
von: Li, Kaining, et al.
Veröffentlicht: (2025)
von: Li, Kaining, et al.
Veröffentlicht: (2025)
MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training
von: Zeeshan, Muhammad Osama, et al.
Veröffentlicht: (2025)
von: Zeeshan, Muhammad Osama, et al.
Veröffentlicht: (2025)
What Matters in Autonomous Driving Anomaly Detection: A Weakly Supervised Horizon
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2024)
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2024)
Task-Adapter++: Task-specific Adaptation with Order-aware Alignment for Few-shot Action Recognition
von: Cao, Congqi, et al.
Veröffentlicht: (2025)
von: Cao, Congqi, et al.
Veröffentlicht: (2025)
Leveraging Temporal Contextualization for Video Action Recognition
von: Kim, Minji, et al.
Veröffentlicht: (2024)
von: Kim, Minji, et al.
Veröffentlicht: (2024)
LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
Effective Adapter for Face Recognition in the Wild
von: Liu, Yunhao, et al.
Veröffentlicht: (2023)
von: Liu, Yunhao, et al.
Veröffentlicht: (2023)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
von: Rajendiran, Ramanathan, et al.
Veröffentlicht: (2023)
von: Rajendiran, Ramanathan, et al.
Veröffentlicht: (2023)
T-Gated Adapter: A Lightweight Temporal Adapter for Vision-Language Medical Segmentation
von: Khadka, Pranjal
Veröffentlicht: (2026)
von: Khadka, Pranjal
Veröffentlicht: (2026)
Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
von: Jiang, Yue, et al.
Veröffentlicht: (2026)
von: Jiang, Yue, et al.
Veröffentlicht: (2026)
DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition
von: Ullah, Hayat, et al.
Veröffentlicht: (2025)
von: Ullah, Hayat, et al.
Veröffentlicht: (2025)
SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition
von: Do, Jeonghyeok, et al.
Veröffentlicht: (2024)
von: Do, Jeonghyeok, et al.
Veröffentlicht: (2024)
Joint Temporal Pooling for Improving Skeleton-based Action Recognition
von: Gunasekara, Shanaka Ramesh, et al.
Veröffentlicht: (2024)
von: Gunasekara, Shanaka Ramesh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LAC: Latent Action Composition for Skeleton-based Action Segmentation
von: Yang, Di, et al.
Veröffentlicht: (2023) -
LIA-X: Interpretable Latent Portrait Animator
von: Wang, Yaohui, et al.
Veröffentlicht: (2025) -
Beyond the Visible: A Survey on Cross-spectral Face Recognition
von: Anghelone, David, et al.
Veröffentlicht: (2022) -
Now You See Me, Now You Don't: A Unified Framework for Expression Consistent Anonymization in Talking Head Videos
von: Egin, Anil, et al.
Veröffentlicht: (2026) -
DenVisCoM: Dense Vision Correspondence Mamba for Efficient and Real-time Optical Flow and Stereo Estimation
von: Anand, Tushar, et al.
Veröffentlicht: (2026)