CAKE: Real-time Action Detection via Motion Distillation and Background-aware Contrastive Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Hoang, Hieu, Tran, Dung Trung, Nguyen, Hong, Nguyen, Nam-Phong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
by: Nguyen, Hong, et al.
Published: (2025)
by: Nguyen, Hong, et al.
Published: (2025)
VRAE: Vertical Residual Autoencoder for License Plate Denoising and Deblurring
by: Nguyen, Cuong, et al.
Published: (2025)
by: Nguyen, Cuong, et al.
Published: (2025)
Semise: Semi-supervised learning for severity representation in medical image
by: Tran, Dung T., et al.
Published: (2025)
by: Tran, Dung T., et al.
Published: (2025)
MonoVQD: Monocular 3D Object Detection with Variational Query Denoising and Self-Distillation
by: Vu, Kiet Dang, et al.
Published: (2025)
by: Vu, Kiet Dang, et al.
Published: (2025)
Emotion-Aware Classroom Quality Assessment Leveraging IoT-Based Real-Time Student Monitoring
by: Nguyen, Hai, et al.
Published: (2026)
by: Nguyen, Hai, et al.
Published: (2026)
Mono3DV: Monocular 3D Object Detection with 3D-Aware Bipartite Matching and Variational Query DeNoising
by: Vu, Kiet Dang, et al.
Published: (2026)
by: Vu, Kiet Dang, et al.
Published: (2026)
View-aware Cross-modal Distillation for Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
A Novel Tuning Method for Real-time Multiple-Object Tracking Utilizing Thermal Sensor with Complexity Motion Pattern
by: Tran, Duong Nguyen-Ngoc, et al.
Published: (2025)
by: Tran, Duong Nguyen-Ngoc, et al.
Published: (2025)
PixelRush: Ultra-Fast, Training-Free High-Resolution Image Generation via One-step Diffusion
by: Lai, Hong-Phuc, et al.
Published: (2026)
by: Lai, Hong-Phuc, et al.
Published: (2026)
Evaluating Precise Geolocation Inference Capabilities of Vision Language Models
by: Jay, Neel, et al.
Published: (2025)
by: Jay, Neel, et al.
Published: (2025)
ConPro: Learning Severity Representation for Medical Images using Contrastive Learning and Preference Optimization
by: Nguyen, Hong, et al.
Published: (2024)
by: Nguyen, Hong, et al.
Published: (2024)
PADM: A Physics-aware Diffusion Model for Attenuation Correction
by: Pham, Trung Kien, et al.
Published: (2025)
by: Pham, Trung Kien, et al.
Published: (2025)
Image-level Regression for Uncertainty-aware Retinal Image Segmentation
by: Dang, Trung, et al.
Published: (2024)
by: Dang, Trung, et al.
Published: (2024)
Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
by: Nguyen, Thong Thanh, et al.
Published: (2024)
by: Nguyen, Thong Thanh, et al.
Published: (2024)
SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score Distillation
by: Nguyen, Thuan Hoang, et al.
Published: (2023)
by: Nguyen, Thuan Hoang, et al.
Published: (2023)
Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
by: Nguyen, Hieu, et al.
Published: (2024)
by: Nguyen, Hieu, et al.
Published: (2024)
SepFormer: Coarse-to-fine Separator Regression Network for Table Structure Recognition
by: Nguyen, Nam Quan, et al.
Published: (2025)
by: Nguyen, Nam Quan, et al.
Published: (2025)
MatSSL: Robust Self-Supervised Representation Learning for Metallographic Image Segmentation
by: Nguyen, Hoang Hai Nam, et al.
Published: (2025)
by: Nguyen, Hoang Hai Nam, et al.
Published: (2025)
ModeDreamer: Mode Guiding Score Distillation for Text-to-3D Generation using Reference Image Prompts
by: Tran, Uy Dieu, et al.
Published: (2024)
by: Tran, Uy Dieu, et al.
Published: (2024)
Guiding Noisy Label Conditional Diffusion Models with Score-based Discriminator Correction
by: Cong, Dat Nguyen, et al.
Published: (2025)
by: Cong, Dat Nguyen, et al.
Published: (2025)
Retrospective Feature Estimation for Continual Learning
by: Nguyen, Nghia D., et al.
Published: (2024)
by: Nguyen, Nghia D., et al.
Published: (2024)
Cycle Training with Semi-Supervised Domain Adaptation: Bridging Accuracy and Efficiency for Real-Time Mobile Scene Detection
by: Phan-Nguyen, Huu-Phong, et al.
Published: (2025)
by: Phan-Nguyen, Huu-Phong, et al.
Published: (2025)
Fourier-Attentive Representation Learning: A Fourier-Guided Framework for Few-Shot Generalization in Vision-Language Models
by: Pham, Hieu Dinh Trung, et al.
Published: (2025)
by: Pham, Hieu Dinh Trung, et al.
Published: (2025)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
by: Nguyen, Hieu, et al.
Published: (2025)
by: Nguyen, Hieu, et al.
Published: (2025)
Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification
by: Nguyen, Y Hop, et al.
Published: (2025)
by: Nguyen, Y Hop, et al.
Published: (2025)
SiNGR: Brain Tumor Segmentation via Signed Normalized Geodesic Transform Regression
by: Dang, Trung, et al.
Published: (2024)
by: Dang, Trung, et al.
Published: (2024)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
by: Vu, Sinh Trong, et al.
Published: (2025)
by: Vu, Sinh Trong, et al.
Published: (2025)
SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation
by: Pham, Duc-Hai, et al.
Published: (2024)
by: Pham, Duc-Hai, et al.
Published: (2024)
Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model Guidance
by: Pham, Duc-Hai, et al.
Published: (2024)
by: Pham, Duc-Hai, et al.
Published: (2024)
Bridging Classification and Segmentation in Osteosarcoma Assessment via Foundation and Discrete Diffusion Models
by: Nguyen, Manh Duong, et al.
Published: (2025)
by: Nguyen, Manh Duong, et al.
Published: (2025)
Persistent Test-time Adaptation in Recurring Testing Scenarios
by: Hoang, Trung-Hieu, et al.
Published: (2023)
by: Hoang, Trung-Hieu, et al.
Published: (2023)
Action Selection Learning for Multi-label Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
EVL-ECG: Efficient ECG Interpretation With Multi-Aspect Heterogeneous Knowledge Distillation
by: Hong, Dang Nguyen, et al.
Published: (2026)
by: Hong, Dang Nguyen, et al.
Published: (2026)
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
Learning to Stop Overthinking at Test Time
by: Bao, Hieu Tran, et al.
Published: (2025)
by: Bao, Hieu Tran, et al.
Published: (2025)
ConstStyle: Robust Domain Generalization with Unified Style Transformation
by: Tran, Nam Duong, et al.
Published: (2025)
by: Tran, Nam Duong, et al.
Published: (2025)
Handling Supervision Scarcity in Chest X-ray Classification: Long-Tailed and Zero-Shot Learning
by: Pham, Ha-Hieu, et al.
Published: (2026)
by: Pham, Ha-Hieu, et al.
Published: (2026)
Alice v1: Distillation-Enhanced Video Generation Surpassing Closed-Source Models
by: Xiaoyu, Wang, et al.
Published: (2026)
by: Xiaoyu, Wang, et al.
Published: (2026)
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
by: Nguyen, Viet, et al.
Published: (2024)
by: Nguyen, Viet, et al.
Published: (2024)
Automated Image Recognition Framework
by: Nguyen, Quang-Binh, et al.
Published: (2025)
by: Nguyen, Quang-Binh, et al.
Published: (2025)
Similar Items
-
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
by: Nguyen, Hong, et al.
Published: (2025) -
VRAE: Vertical Residual Autoencoder for License Plate Denoising and Deblurring
by: Nguyen, Cuong, et al.
Published: (2025) -
Semise: Semi-supervised learning for severity representation in medical image
by: Tran, Dung T., et al.
Published: (2025) -
MonoVQD: Monocular 3D Object Detection with Variational Query Denoising and Self-Distillation
by: Vu, Kiet Dang, et al.
Published: (2025) -
Emotion-Aware Classroom Quality Assessment Leveraging IoT-Based Real-Time Student Monitoring
by: Nguyen, Hai, et al.
Published: (2026)