YOWOv3: An Efficient and Generalized Framework for Human Action Detection and Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dang, Duc Manh Nguyen, Duong, Viet Hang, Wang, Jia Ching, Duc, Nhan Bui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hierarchical Neural Collapse Detection Transformer for Class Incremental Object Detection
von: Pham, Duc Thanh, et al.
Veröffentlicht: (2025)
von: Pham, Duc Thanh, et al.
Veröffentlicht: (2025)
LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models
von: Debnath, Soumyaratna, et al.
Veröffentlicht: (2026)
von: Debnath, Soumyaratna, et al.
Veröffentlicht: (2026)
MMAP: A Multi-Magnification and Prototype-Aware Architecture for Predicting Spatial Gene Expression
von: Nguyen, Hai Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Hai Dang, et al.
Veröffentlicht: (2025)
CLIPping the Deception: Adapting Vision-Language Models for Universal Deepfake Detection
von: Khan, Sohail Ahmed, et al.
Veröffentlicht: (2024)
von: Khan, Sohail Ahmed, et al.
Veröffentlicht: (2024)
MonoVQD: Monocular 3D Object Detection with Variational Query Denoising and Self-Distillation
von: Vu, Kiet Dang, et al.
Veröffentlicht: (2025)
von: Vu, Kiet Dang, et al.
Veröffentlicht: (2025)
N-EIoU-YOLOv9: A Signal-Aware Bounding Box Regression Loss for Lightweight Mobile Detection of Rice Leaf Diseases
von: Duc, Dung Ta Nguyen, et al.
Veröffentlicht: (2026)
von: Duc, Dung Ta Nguyen, et al.
Veröffentlicht: (2026)
Representation Learning with Semantic-aware Instance and Sparse Token Alignments
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2026)
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2026)
Collaborative Perceiver: Elevating Vision-based 3D Object Detection via Local Density-Aware Spatial Occupancy
von: Yuan, Jicheng, et al.
Veröffentlicht: (2025)
von: Yuan, Jicheng, et al.
Veröffentlicht: (2025)
Mono3DV: Monocular 3D Object Detection with 3D-Aware Bipartite Matching and Variational Query DeNoising
von: Vu, Kiet Dang, et al.
Veröffentlicht: (2026)
von: Vu, Kiet Dang, et al.
Veröffentlicht: (2026)
QCFace: Image Quality Control for boosting Face Representation & Recognition
von: Doan-Ngo, Duc-Phuong, et al.
Veröffentlicht: (2025)
von: Doan-Ngo, Duc-Phuong, et al.
Veröffentlicht: (2025)
Domain Generalization through Spatial Relation Induction over Visual Primitives
von: Nguyen, Dat, et al.
Veröffentlicht: (2026)
von: Nguyen, Dat, et al.
Veröffentlicht: (2026)
AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
von: Jiang, Yichen, et al.
Veröffentlicht: (2025)
von: Jiang, Yichen, et al.
Veröffentlicht: (2025)
Frequency Adapter with SAM for Generalized Medical Image Segmentation
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2026)
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2026)
Adaptive Cache Enhancement for Test-Time Adaptation of Vision-Language Models
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2025)
Detection Fire in Camera RGB-NIR
von: Khai, Nguyen Truong, et al.
Veröffentlicht: (2025)
von: Khai, Nguyen Truong, et al.
Veröffentlicht: (2025)
Exploring the Practicality of Federated Learning: A Survey Towards the Communication Perspective
von: Le, Khiem, et al.
Veröffentlicht: (2024)
von: Le, Khiem, et al.
Veröffentlicht: (2024)
Count What You Want: Exemplar Identification and Few-shot Counting of Human Actions in the Wild
von: Huang, Yifeng, et al.
Veröffentlicht: (2023)
von: Huang, Yifeng, et al.
Veröffentlicht: (2023)
Clinical Graph-Mediated Distillation for Unpaired MRI-to-CFI Hypertension Prediction
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
Unsupervised Domain Adaptation with SAM-RefiSeR for Enhanced Brain Tumor Segmentation
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
von: Imans, Dillan, et al.
Veröffentlicht: (2026)
Object Detection in Thermal Images Using Deep Learning for Unmanned Aerial Vehicles
von: Tu, Minh Dang, et al.
Veröffentlicht: (2024)
von: Tu, Minh Dang, et al.
Veröffentlicht: (2024)
Bridging Classification and Segmentation in Osteosarcoma Assessment via Foundation and Discrete Diffusion Models
von: Nguyen, Manh Duong, et al.
Veröffentlicht: (2025)
von: Nguyen, Manh Duong, et al.
Veröffentlicht: (2025)
Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model Guidance
von: Pham, Duc-Hai, et al.
Veröffentlicht: (2024)
von: Pham, Duc-Hai, et al.
Veröffentlicht: (2024)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
von: Le, Minh Khoa, et al.
Veröffentlicht: (2026)
von: Le, Minh Khoa, et al.
Veröffentlicht: (2026)
Sampling Foundational Transformer: A Theoretical Perspective
von: Nguyen, Viet Anh, et al.
Veröffentlicht: (2024)
von: Nguyen, Viet Anh, et al.
Veröffentlicht: (2024)
Physics-informed Ground Reaction Dynamics from Human Motion Capture
von: Le, Cuong, et al.
Veröffentlicht: (2025)
von: Le, Cuong, et al.
Veröffentlicht: (2025)
Bidirectional Diffusion Bridge Models
von: Kieu, Duc, et al.
Veröffentlicht: (2025)
von: Kieu, Duc, et al.
Veröffentlicht: (2025)
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023)
Improved Training Technique for Shortcut Models
von: Nguyen, Anh, et al.
Veröffentlicht: (2025)
von: Nguyen, Anh, et al.
Veröffentlicht: (2025)
Enhanced Generative Data Augmentation for Semantic Segmentation via Stronger Guidance
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
von: Che, Quang-Huy, et al.
Veröffentlicht: (2024)
A statistical method for crack pre-detection in 3D concrete images
von: Makogin, Vitalii, et al.
Veröffentlicht: (2024)
von: Makogin, Vitalii, et al.
Veröffentlicht: (2024)
Active Generation Network of Human Skeleton for Action Recognition
von: Liu, Long, et al.
Veröffentlicht: (2024)
von: Liu, Long, et al.
Veröffentlicht: (2024)
A model-agnostic active learning approach for animal detection from camera traps
von: Nguyen, Thi Thu Thuy, et al.
Veröffentlicht: (2025)
von: Nguyen, Thi Thu Thuy, et al.
Veröffentlicht: (2025)
PDIWS: Thermal Imaging Dataset for Person Detection in Intrusion Warning Systems
von: Thuan, Nguyen Duc, et al.
Veröffentlicht: (2023)
von: Thuan, Nguyen Duc, et al.
Veröffentlicht: (2023)
Action Recognition Using Temporal Shift Module and Ensemble Learning
von: Duong, Anh-Kiet, et al.
Veröffentlicht: (2025)
von: Duong, Anh-Kiet, et al.
Veröffentlicht: (2025)
IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
Variational Contrastive Learning for Skeleton-based Action Recognition
von: Nguyen, Dang Dinh, et al.
Veröffentlicht: (2026)
von: Nguyen, Dang Dinh, et al.
Veröffentlicht: (2026)
Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text Recognition
von: Le, Kha Nhat, et al.
Veröffentlicht: (2024)
von: Le, Kha Nhat, et al.
Veröffentlicht: (2024)
h-Edit: Effective and Flexible Diffusion-Based Editing via Doob's h-Transform
von: Nguyen, Toan, et al.
Veröffentlicht: (2025)
von: Nguyen, Toan, et al.
Veröffentlicht: (2025)
Multi-scale Feature Enhancement in Multi-task Learning for Medical Image Analysis
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2024)
von: Bui, Phuoc-Nguyen, et al.
Veröffentlicht: (2024)
FedBlock: A Blockchain Approach to Federated Learning against Backdoor Attacks
von: Nguyen, Duong H., et al.
Veröffentlicht: (2024)
von: Nguyen, Duong H., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hierarchical Neural Collapse Detection Transformer for Class Incremental Object Detection
von: Pham, Duc Thanh, et al.
Veröffentlicht: (2025) -
LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models
von: Debnath, Soumyaratna, et al.
Veröffentlicht: (2026) -
MMAP: A Multi-Magnification and Prototype-Aware Architecture for Predicting Spatial Gene Expression
von: Nguyen, Hai Dang, et al.
Veröffentlicht: (2025) -
CLIPping the Deception: Adapting Vision-Language Models for Universal Deepfake Detection
von: Khan, Sohail Ahmed, et al.
Veröffentlicht: (2024) -
MonoVQD: Monocular 3D Object Detection with Variational Query Denoising and Self-Distillation
von: Vu, Kiet Dang, et al.
Veröffentlicht: (2025)