Gespeichert in:
| Hauptverfasser: | Sia, Zhen Hao, Rawat, Yogesh Singh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2504.03096 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stable Mean Teacher for Semi-supervised Video Action Detection
von: Kumar, Akash, et al.
Veröffentlicht: (2024)
von: Kumar, Akash, et al.
Veröffentlicht: (2024)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
von: Modi, Rajat, et al.
Veröffentlicht: (2024)
von: Modi, Rajat, et al.
Veröffentlicht: (2024)
Semi-supervised Active Learning for Video Action Detection
von: Singh, Ayush, et al.
Veröffentlicht: (2023)
von: Singh, Ayush, et al.
Veröffentlicht: (2023)
Activity-Biometrics: Person Identification from Daily Activities
von: Azad, Shehreen, et al.
Veröffentlicht: (2024)
von: Azad, Shehreen, et al.
Veröffentlicht: (2024)
Open-Vocabulary Spatio-Temporal Action Detection
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
von: Ahmad, Shahzad, et al.
Veröffentlicht: (2023)
von: Ahmad, Shahzad, et al.
Veröffentlicht: (2023)
Asynchronous Perception Machine For Efficient Test-Time-Training
von: Modi, Rajat, et al.
Veröffentlicht: (2024)
von: Modi, Rajat, et al.
Veröffentlicht: (2024)
Scaling Open-Vocabulary Object Detection
von: Minderer, Matthias, et al.
Veröffentlicht: (2023)
von: Minderer, Matthias, et al.
Veröffentlicht: (2023)
MolVision: Molecular Property Prediction with Vision Language Models
von: Adak, Deepan, et al.
Veröffentlicht: (2025)
von: Adak, Deepan, et al.
Veröffentlicht: (2025)
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
von: Azad, Shehreen, et al.
Veröffentlicht: (2025)
von: Azad, Shehreen, et al.
Veröffentlicht: (2025)
iSafetyBench: A video-language benchmark for safety in industrial environment
von: Abdullah, Raiyaan, et al.
Veröffentlicht: (2025)
von: Abdullah, Raiyaan, et al.
Veröffentlicht: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
StreamReady: Learning What to Answer and When in Long Streaming Videos
von: Azad, Shehreen, et al.
Veröffentlicht: (2026)
von: Azad, Shehreen, et al.
Veröffentlicht: (2026)
Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-ID
von: Liang, Xin, et al.
Veröffentlicht: (2025)
von: Liang, Xin, et al.
Veröffentlicht: (2025)
DisenQ: Disentangling Q-Former for Activity-Biometrics
von: Azad, Shehreen, et al.
Veröffentlicht: (2025)
von: Azad, Shehreen, et al.
Veröffentlicht: (2025)
Coarse Attribute Prediction with Task Agnostic Distillation for Real World Clothes Changing ReID
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction
von: Mitra, Sirshapan, et al.
Veröffentlicht: (2026)
von: Mitra, Sirshapan, et al.
Veröffentlicht: (2026)
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
von: Bao, Wentao, et al.
Veröffentlicht: (2024)
von: Bao, Wentao, et al.
Veröffentlicht: (2024)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
GaitCrafter: Diffusion Model for Biometric Preserving Gait Synthesis
von: Mitra, Sirshapan, et al.
Veröffentlicht: (2025)
von: Mitra, Sirshapan, et al.
Veröffentlicht: (2025)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
von: Garg, Aaryan, et al.
Veröffentlicht: (2025)
von: Garg, Aaryan, et al.
Veröffentlicht: (2025)
Navigating Hallucinations for Reasoning of Unintentional Activities
von: Grover, Shresth, et al.
Veröffentlicht: (2024)
von: Grover, Shresth, et al.
Veröffentlicht: (2024)
Advancing Automatic Photovoltaic Defect Detection using Semi-Supervised Semantic Segmentation of Electroluminescence Images
von: Jha, Abhishek, et al.
Veröffentlicht: (2024)
von: Jha, Abhishek, et al.
Veröffentlicht: (2024)
Open Vocabulary Monocular 3D Object Detection
von: Yao, Jin, et al.
Veröffentlicht: (2024)
von: Yao, Jin, et al.
Veröffentlicht: (2024)
DENOISER: Rethinking the Robustness for Open-Vocabulary Action Recognition
von: Cheng, Haozhe, et al.
Veröffentlicht: (2024)
von: Cheng, Haozhe, et al.
Veröffentlicht: (2024)
Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
von: Ni, TsaiChing, et al.
Veröffentlicht: (2025)
von: Ni, TsaiChing, et al.
Veröffentlicht: (2025)
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection
von: Zhu, Sa, et al.
Veröffentlicht: (2026)
von: Zhu, Sa, et al.
Veröffentlicht: (2026)
Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2025)
von: Yuan, Zhenlong, et al.
Veröffentlicht: (2025)
Learning to Generalize without Bias for Open-Vocabulary Action Recognition
von: Yu, Yating, et al.
Veröffentlicht: (2025)
von: Yu, Yating, et al.
Veröffentlicht: (2025)
Open-Vocabulary Video Anomaly Detection
von: Wu, Peng, et al.
Veröffentlicht: (2023)
von: Wu, Peng, et al.
Veröffentlicht: (2023)
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
von: Zhu, Sa, et al.
Veröffentlicht: (2026)
von: Zhu, Sa, et al.
Veröffentlicht: (2026)
FROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection
von: Chen, Fangyi, et al.
Veröffentlicht: (2024)
von: Chen, Fangyi, et al.
Veröffentlicht: (2024)
Learning to Detect and Segment for Open Vocabulary Object Detection
von: Wang, Tao, et al.
Veröffentlicht: (2022)
von: Wang, Tao, et al.
Veröffentlicht: (2022)
LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
von: Pathak, Priyank, et al.
Veröffentlicht: (2025)
OmViD: Omni-supervised active learning for video action detection
von: Rana, Aayush, et al.
Veröffentlicht: (2025)
von: Rana, Aayush, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Stable Mean Teacher for Semi-supervised Video Action Detection
von: Kumar, Akash, et al.
Veröffentlicht: (2024) -
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
von: Modi, Rajat, et al.
Veröffentlicht: (2024) -
Semi-supervised Active Learning for Video Action Detection
von: Singh, Ayush, et al.
Veröffentlicht: (2023) -
Activity-Biometrics: Person Identification from Daily Activities
von: Azad, Shehreen, et al.
Veröffentlicht: (2024) -
Open-Vocabulary Spatio-Temporal Action Detection
von: Wu, Tao, et al.
Veröffentlicht: (2024)