ActivityCLIP: Enhancing Group Activity Recognition by Mining Complementary Information from Text to Supplement Image Modality
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Guoliang, Yin, Jianqin, Zhou, Feng, Dang, Yonghao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023)
by: Zhang, Shaojie, et al.
Published: (2023)
L2HCount:Generalizing Crowd Counting from Low to High Crowd Density via Density Simulation
by: Xu, Guoliang, et al.
Published: (2025)
by: Xu, Guoliang, et al.
Published: (2025)
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023)
by: Zhang, Shaojie, et al.
Published: (2023)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
by: Li, Huilai, et al.
Published: (2026)
by: Li, Huilai, et al.
Published: (2026)
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
by: Jiang, Yuanyuan, et al.
Published: (2022)
by: Jiang, Yuanyuan, et al.
Published: (2022)
ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
by: Ma, Yiyang, et al.
Published: (2025)
by: Ma, Yiyang, et al.
Published: (2025)
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
by: Wei, Wei, et al.
Published: (2025)
by: Wei, Wei, et al.
Published: (2025)
Physics-constrained Attack against Convolution-based Human Motion Prediction
by: Duan, Chengxu, et al.
Published: (2023)
by: Duan, Chengxu, et al.
Published: (2023)
MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection
by: Dang, Yonghao, et al.
Published: (2024)
by: Dang, Yonghao, et al.
Published: (2024)
Kinematics Modeling Network for Video-based Human Pose Estimation
by: Dang, Yonghao, et al.
Published: (2022)
by: Dang, Yonghao, et al.
Published: (2022)
Pixels or Positions? Benchmarking Modalities in Group Activity Recognition
by: Karki, Drishya, et al.
Published: (2025)
by: Karki, Drishya, et al.
Published: (2025)
3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
by: Zhang, Yixun, et al.
Published: (2025)
by: Zhang, Yixun, et al.
Published: (2025)
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering
by: Jiang, Yuanyuan, et al.
Published: (2024)
by: Jiang, Yuanyuan, et al.
Published: (2024)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
by: Li, Huilai, et al.
Published: (2025)
by: Li, Huilai, et al.
Published: (2025)
Towards Physically Realizable Adversarial Attacks in Embodied Vision Navigation
by: Chen, Meng, et al.
Published: (2024)
by: Chen, Meng, et al.
Published: (2024)
CMIP-CIL: A Cross-Modal Benchmark for Image-Point Class Incremental Learning
by: Qi, Chao, et al.
Published: (2025)
by: Qi, Chao, et al.
Published: (2025)
DHRNet: A Dual-Path Hierarchical Relation Network for Multi-Person Pose Estimation
by: Dang, Yonghao, et al.
Published: (2024)
by: Dang, Yonghao, et al.
Published: (2024)
DEFT-LLM: Disentangled Expert Feature Tuning for Micro-Expression Recognition
by: Zhang, Ren, et al.
Published: (2025)
by: Zhang, Ren, et al.
Published: (2025)
PromptGAR: Flexible Promptive Group Activity Recognition
by: Jin, Zhangyu, et al.
Published: (2025)
by: Jin, Zhangyu, et al.
Published: (2025)
A Two-stream Hybrid CNN-Transformer Network for Skeleton-based Human Interaction Recognition
by: Yin, Ruoqi, et al.
Published: (2023)
by: Yin, Ruoqi, et al.
Published: (2023)
BiTAA: A Bi-Task Adversarial Attack for Object Detection and Depth Estimation via 3D Gaussian Splatting
by: Zhang, Yixun, et al.
Published: (2025)
by: Zhang, Yixun, et al.
Published: (2025)
LiGAR: LiDAR-Guided Hierarchical Transformer for Multi-Modal Group Activity Recognition
by: Chappa, Naga Venkata Sai Raviteja, et al.
Published: (2024)
by: Chappa, Naga Venkata Sai Raviteja, et al.
Published: (2024)
ARIC: An Activity Recognition Dataset in Classroom Surveillance Images
by: Xu, Linfeng, et al.
Published: (2024)
by: Xu, Linfeng, et al.
Published: (2024)
Mining Open Semantics from CLIP: A Relation Transition Perspective for Few-Shot Learning
by: Yan, Cilin, et al.
Published: (2024)
by: Yan, Cilin, et al.
Published: (2024)
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
by: Wang, Jingyun, et al.
Published: (2024)
by: Wang, Jingyun, et al.
Published: (2024)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
by: Csizmadia, Daniel, et al.
Published: (2025)
by: Csizmadia, Daniel, et al.
Published: (2025)
Exploring Position Encoding in Diffusion U-Net for Training-free High-resolution Image Generation
by: Zhou, Feng, et al.
Published: (2025)
by: Zhou, Feng, et al.
Published: (2025)
Unified Framework with Consistency across Modalities for Human Activity Recognition
by: Tran, Tuyen, et al.
Published: (2024)
by: Tran, Tuyen, et al.
Published: (2024)
Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion Models
by: Xu, Ganxi, et al.
Published: (2025)
by: Xu, Ganxi, et al.
Published: (2025)
VicTR: Video-conditioned Text Representations for Activity Recognition
by: Kahatapitiya, Kumara, et al.
Published: (2023)
by: Kahatapitiya, Kumara, et al.
Published: (2023)
AsyncDiff: Asynchronous Timestep Conditioning for Enhanced Text-to-Image Diffusion Inference
by: Xu, Longhuan, et al.
Published: (2025)
by: Xu, Longhuan, et al.
Published: (2025)
Towards more realistic human motion prediction with attention to motion coordination
by: Ding, Pengxiang, et al.
Published: (2024)
by: Ding, Pengxiang, et al.
Published: (2024)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
Few-Shot Continual Learning for Activity Recognition in Classroom Surveillance Images
by: Qian, Yilei, et al.
Published: (2024)
by: Qian, Yilei, et al.
Published: (2024)
CLIP in Medical Imaging: A Survey
by: Zhao, Zihao, et al.
Published: (2023)
by: Zhao, Zihao, et al.
Published: (2023)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
by: Zhang, Zilun, et al.
Published: (2022)
by: Zhang, Zilun, et al.
Published: (2022)
Skeleton-based Group Activity Recognition via Spatial-Temporal Panoramic Graph
by: Li, Zhengcen, et al.
Published: (2024)
by: Li, Zhengcen, et al.
Published: (2024)
A Comprehensive Methodological Survey of Human Activity Recognition Across Divers Data Modalities
by: Shin, Jungpil, et al.
Published: (2024)
by: Shin, Jungpil, et al.
Published: (2024)
Bilateral Event Mining and Complementary for Event Stream Super-Resolution
by: Huang, Zhilin, et al.
Published: (2024)
by: Huang, Zhilin, et al.
Published: (2024)
VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling
by: Zhang, Qian, et al.
Published: (2024)
by: Zhang, Qian, et al.
Published: (2024)
Similar Items
-
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023) -
L2HCount:Generalizing Crowd Counting from Low to High Crowd Density via Density Simulation
by: Xu, Guoliang, et al.
Published: (2025) -
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023) -
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
by: Li, Huilai, et al.
Published: (2026) -
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
by: Jiang, Yuanyuan, et al.
Published: (2022)