DENOISER: Rethinking the Robustness for Open-Vocabulary Action Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Cheng, Haozhe, Ju, Cheng, Wang, Haicheng, Liu, Jinxiang, Chen, Mengting, Hu, Qiang, Zhang, Xiaoyun, Wang, Yanfeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
por: Lin, Kun-Yu, et al.
Publicado: (2024)
por: Lin, Kun-Yu, et al.
Publicado: (2024)
Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
por: Yuan, Zhenlong, et al.
Publicado: (2025)
por: Yuan, Zhenlong, et al.
Publicado: (2025)
MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
por: Wang, Kuo, et al.
Publicado: (2024)
por: Wang, Kuo, et al.
Publicado: (2024)
Learning to Generalize without Bias for Open-Vocabulary Action Recognition
por: Yu, Yating, et al.
Publicado: (2025)
por: Yu, Yating, et al.
Publicado: (2025)
AttrSeg: Open-Vocabulary Semantic Segmentation via Attribute Decomposition-Aggregation
por: Ma, Chaofan, et al.
Publicado: (2023)
por: Ma, Chaofan, et al.
Publicado: (2023)
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
por: So, Yerim, et al.
Publicado: (2026)
por: So, Yerim, et al.
Publicado: (2026)
Local Spatiotemporal Convolutional Network for Robust Gait Recognition
por: Wang, Xiaoyun, et al.
Publicado: (2026)
por: Wang, Xiaoyun, et al.
Publicado: (2026)
MegaFusion: Extend Diffusion Models towards Higher-resolution Image Generation without Further Tuning
por: Wu, Haoning, et al.
Publicado: (2024)
por: Wu, Haoning, et al.
Publicado: (2024)
Contrast-Unity for Partially-Supervised Temporal Sentence Grounding
por: Wang, Haicheng, et al.
Publicado: (2025)
por: Wang, Haicheng, et al.
Publicado: (2025)
Open-Vocabulary Spatio-Temporal Action Detection
por: Wu, Tao, et al.
Publicado: (2024)
por: Wu, Tao, et al.
Publicado: (2024)
Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models
por: Ju, Chen, et al.
Publicado: (2024)
por: Ju, Chen, et al.
Publicado: (2024)
Open Vocabulary Monocular 3D Object Detection
por: Yao, Jin, et al.
Publicado: (2024)
por: Yao, Jin, et al.
Publicado: (2024)
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
por: Li, Yongkang, et al.
Publicado: (2024)
por: Li, Yongkang, et al.
Publicado: (2024)
Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training
por: Wang, Haicheng, et al.
Publicado: (2024)
por: Wang, Haicheng, et al.
Publicado: (2024)
Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models
por: Liu, Chang, et al.
Publicado: (2023)
por: Liu, Chang, et al.
Publicado: (2023)
Scaling Open-Vocabulary Action Detection
por: Sia, Zhen Hao, et al.
Publicado: (2025)
por: Sia, Zhen Hao, et al.
Publicado: (2025)
OwlSight: A Robust Illumination Adaptation Framework for Dark Video Human Action Recognition
por: Cheng, Shihao, et al.
Publicado: (2025)
por: Cheng, Shihao, et al.
Publicado: (2025)
YOLO-World: Real-Time Open-Vocabulary Object Detection
por: Cheng, Tianheng, et al.
Publicado: (2024)
por: Cheng, Tianheng, et al.
Publicado: (2024)
OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer
por: Wang, Yu, et al.
Publicado: (2024)
por: Wang, Yu, et al.
Publicado: (2024)
Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection
por: Zhu, Sa, et al.
Publicado: (2026)
por: Zhu, Sa, et al.
Publicado: (2026)
Open-Vocabulary SAM3D: Towards Training-free Open-Vocabulary 3D Scene Understanding
por: Tai, Hanchen, et al.
Publicado: (2024)
por: Tai, Hanchen, et al.
Publicado: (2024)
DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label Recognition
por: Liu, Haijing, et al.
Publicado: (2025)
por: Liu, Haijing, et al.
Publicado: (2025)
OVMR: Open-Vocabulary Recognition with Multi-Modal References
por: Ma, Zehong, et al.
Publicado: (2024)
por: Ma, Zehong, et al.
Publicado: (2024)
One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution
por: Fang, Yushun, et al.
Publicado: (2025)
por: Fang, Yushun, et al.
Publicado: (2025)
Exploring Open-Vocabulary Object Recognition in Images using CLIP
por: Chen, Wei Yu, et al.
Publicado: (2026)
por: Chen, Wei Yu, et al.
Publicado: (2026)
Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
por: Zhu, Sa, et al.
Publicado: (2026)
por: Zhu, Sa, et al.
Publicado: (2026)
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
por: Bao, Wentao, et al.
Publicado: (2024)
por: Bao, Wentao, et al.
Publicado: (2024)
Anomize: Better Open Vocabulary Video Anomaly Detection
por: Li, Fei, et al.
Publicado: (2025)
por: Li, Fei, et al.
Publicado: (2025)
Align then Adapt: Rethinking Parameter-Efficient Transfer Learning in 4D Perception
por: Sun, Yiding, et al.
Publicado: (2026)
por: Sun, Yiding, et al.
Publicado: (2026)
FROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition
por: Huang, Xiaohu, et al.
Publicado: (2024)
por: Huang, Xiaohu, et al.
Publicado: (2024)
Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
por: Li, Yunheng, et al.
Publicado: (2024)
por: Li, Yunheng, et al.
Publicado: (2024)
Causal Prompt Calibration Guided Segment Anything Model for Open-Vocabulary Multi-Entity Segmentation
por: Wang, Jingyao, et al.
Publicado: (2025)
por: Wang, Jingyao, et al.
Publicado: (2025)
Category-Adaptive Cross-Modal Semantic Refinement and Transfer for Open-Vocabulary Multi-Label Recognition
por: Liu, Haijing, et al.
Publicado: (2024)
por: Liu, Haijing, et al.
Publicado: (2024)
Audio-Visual Segmentation via Unlabeled Frame Exploitation
por: Liu, Jinxiang, et al.
Publicado: (2024)
por: Liu, Jinxiang, et al.
Publicado: (2024)
OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
por: Huang, Sheng-Yu, et al.
Publicado: (2026)
por: Huang, Sheng-Yu, et al.
Publicado: (2026)
Improving Human Image Animation via Semantic Representation Alignment
por: Liu, Chang, et al.
Publicado: (2026)
por: Liu, Chang, et al.
Publicado: (2026)
VRVVC: Variable-Rate NeRF-Based Volumetric Video Compression
por: Hu, Qiang, et al.
Publicado: (2024)
por: Hu, Qiang, et al.
Publicado: (2024)
Rethinking Open-Vocabulary Segmentation of Radiance Fields in 3D Space
por: Lee, Hyunjee, et al.
Publicado: (2024)
por: Lee, Hyunjee, et al.
Publicado: (2024)
Multi-Modal Prototypes for Open-World Semantic Segmentation
por: Yang, Yuhuan, et al.
Publicado: (2023)
por: Yang, Yuhuan, et al.
Publicado: (2023)
Open-Vocabulary Video Anomaly Detection
por: Wu, Peng, et al.
Publicado: (2023)
por: Wu, Peng, et al.
Publicado: (2023)
Ejemplares similares
-
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
por: Lin, Kun-Yu, et al.
Publicado: (2024) -
Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
por: Yuan, Zhenlong, et al.
Publicado: (2025) -
MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
por: Wang, Kuo, et al.
Publicado: (2024) -
Learning to Generalize without Bias for Open-Vocabulary Action Recognition
por: Yu, Yating, et al.
Publicado: (2025) -
AttrSeg: Open-Vocabulary Semantic Segmentation via Attribute Decomposition-Aggregation
por: Ma, Chaofan, et al.
Publicado: (2023)