View-aware Cross-modal Distillation for Multi-view Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Trung Thanh, Kawanishi, Yasutomo, John, Vijay, Komamizu, Takahiro, Ide, Ichiro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
Action Selection Learning for Multi-label Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
Tracking Small Birds by Detection Candidate Region Filtering and Detection History-aware Association
by: Liu, Tingwei, et al.
Published: (2024)
by: Liu, Tingwei, et al.
Published: (2024)
ForestMamba: Sparse Mamba with Geometry-guided Queries for 3D Forest Point Cloud Segmentation
by: Nguyen, Trung Thanh, et al.
Published: (2026)
by: Nguyen, Trung Thanh, et al.
Published: (2026)
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
by: Chen, Junan, et al.
Published: (2025)
by: Chen, Junan, et al.
Published: (2025)
Multi-View Video-Based Learning: Leveraging Weak Labels for Frame-Level Perception
by: John, Vijay, et al.
Published: (2024)
by: John, Vijay, et al.
Published: (2024)
Small Object Detection for Birds with Swin Transformer
by: Huo, Da, et al.
Published: (2025)
by: Huo, Da, et al.
Published: (2025)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
by: Zhang, Bolin, et al.
Published: (2026)
by: Zhang, Bolin, et al.
Published: (2026)
Multi-view Action Recognition via Directed Gromov-Wasserstein Discrepancy
by: Nguyen, Hoang-Quan, et al.
Published: (2024)
by: Nguyen, Hoang-Quan, et al.
Published: (2024)
Cross-view Action Recognition Understanding From Exocentric to Egocentric Perspective
by: Truong, Thanh-Dat, et al.
Published: (2023)
by: Truong, Thanh-Dat, et al.
Published: (2023)
Static and Dynamic Graph Alignment Network for Temporal Video Grounding
by: Hu, Zhanjie, et al.
Published: (2026)
by: Hu, Zhanjie, et al.
Published: (2026)
CAKE: Real-time Action Detection via Motion Distillation and Background-aware Contrastive Learning
by: Hoang, Hieu, et al.
Published: (2026)
by: Hoang, Hieu, et al.
Published: (2026)
Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements
by: Kinoshita, Genki, et al.
Published: (2026)
by: Kinoshita, Genki, et al.
Published: (2026)
Multi-view Distillation based on Multi-modal Fusion for Few-shot Action Recognition(CLIP-$\mathrm{M^2}$DF)
by: Guo, Fei, et al.
Published: (2024)
by: Guo, Fei, et al.
Published: (2024)
ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking
by: Ge, Jiawei, et al.
Published: (2026)
by: Ge, Jiawei, et al.
Published: (2026)
PADM: A Physics-aware Diffusion Model for Attenuation Correction
by: Pham, Trung Kien, et al.
Published: (2025)
by: Pham, Trung Kien, et al.
Published: (2025)
FROSS: Faster-than-Real-Time Online 3D Semantic Scene Graph Generation from RGB-D Images
by: Hou, Hao-Yu, et al.
Published: (2025)
by: Hou, Hao-Yu, et al.
Published: (2025)
MAVR-Net: Robust Multi-View Learning for MAV Action Recognition with Cross-View Attention
by: Zhang, Nengbo, et al.
Published: (2025)
by: Zhang, Nengbo, et al.
Published: (2025)
A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions
by: Inadumi, Shun, et al.
Published: (2024)
by: Inadumi, Shun, et al.
Published: (2024)
Optimized View and Geometry Distillation from Multi-view Diffuser
by: Zhang, Youjia, et al.
Published: (2023)
by: Zhang, Youjia, et al.
Published: (2023)
Trunk-branch Contrastive Network with Multi-view Deformable Aggregation for Multi-view Action Recognition
by: Yang, Yingyuan, et al.
Published: (2025)
by: Yang, Yingyuan, et al.
Published: (2025)
Class-agnostic 3D Segmentation by Granularity-Consistent Automatic 2D Mask Tracking
by: Wang, Juan, et al.
Published: (2025)
by: Wang, Juan, et al.
Published: (2025)
Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
by: He, Wen-Jue, et al.
Published: (2025)
by: He, Wen-Jue, et al.
Published: (2025)
MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition
by: Xu, Zhengyi, et al.
Published: (2026)
by: Xu, Zhengyi, et al.
Published: (2026)
Frequency Attention for Knowledge Distillation
by: Pham, Cuong, et al.
Published: (2024)
by: Pham, Cuong, et al.
Published: (2024)
Self-supervised Video Object Segmentation with Distillation Learning of Deformable Attention
by: Truong, Quang-Trung, et al.
Published: (2024)
by: Truong, Quang-Trung, et al.
Published: (2024)
Skarimva: Skeleton-based Action Recognition is a Multi-view Application
by: Bermuth, Daniel, et al.
Published: (2026)
by: Bermuth, Daniel, et al.
Published: (2026)
REACH: Hand Pose Estimation from Room Corners
by: Nakamura, Shu, et al.
Published: (2026)
by: Nakamura, Shu, et al.
Published: (2026)
ICG-MVSNet: Learning Intra-view and Cross-view Relationships for Guidance in Multi-View Stereo
by: Hu, Yuxi, et al.
Published: (2025)
by: Hu, Yuxi, et al.
Published: (2025)
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching
by: Tian, Mengxiao, et al.
Published: (2025)
by: Tian, Mengxiao, et al.
Published: (2025)
MV-GMN: State Space Model for Multi-View Action Recognition
by: Lin, Yuhui, et al.
Published: (2025)
by: Lin, Yuhui, et al.
Published: (2025)
Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
by: Wang, Hu, et al.
Published: (2023)
by: Wang, Hu, et al.
Published: (2023)
Cross-view Masked Diffusion Transformers for Person Image Synthesis
by: Pham, Trung X., et al.
Published: (2024)
by: Pham, Trung X., et al.
Published: (2024)
DiffAugment: Diffusion based Long-Tailed Visual Relationship Recognition
by: Gupta, Parul, et al.
Published: (2024)
by: Gupta, Parul, et al.
Published: (2024)
Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP
by: Yu, Yating, et al.
Published: (2024)
by: Yu, Yating, et al.
Published: (2024)
TinyBEV: Cross Modal Knowledge Distillation for Efficient Multi Task Bird's Eye View Perception and Planning
by: Khan, Reeshad, et al.
Published: (2025)
by: Khan, Reeshad, et al.
Published: (2025)
GaMO: Geometry-aware Multi-view Diffusion Outpainting for Sparse-View 3D Reconstruction
by: Huang, Yi-Chuan, et al.
Published: (2025)
by: Huang, Yi-Chuan, et al.
Published: (2025)
Thermal Polarimetric Multi-view Stereo
by: Kushida, Takahiro, et al.
Published: (2025)
by: Kushida, Takahiro, et al.
Published: (2025)
Similar Items
-
MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025) -
Action Selection Learning for Multi-label Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2024) -
MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
by: Nguyen, Trung Thanh, et al.
Published: (2025) -
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
by: Nguyen, Trung Thanh, et al.
Published: (2024) -
Tracking Small Birds by Detection Candidate Region Filtering and Detection History-aware Association
by: Liu, Tingwei, et al.
Published: (2024)