DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025
Fuente:
arXiv
Saved in:
| Main Authors: | Kamoto, Umihiro, Ishibashi, Tatsuya, Kugo, Noriyuki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VDMA: Video Question Answering with Dynamically Generated Multi-Agents
by: Kugo, Noriyuki, et al.
Published: (2024)
by: Kugo, Noriyuki, et al.
Published: (2024)
Multimodal Structured Generation: CVPR's 2nd MMFM Challenge Technical Report
by: Cesista, Franz Louis
Published: (2024)
by: Cesista, Franz Louis
Published: (2024)
Technique Report of CVPR 2024 PBDL Challenges
by: Fu, Ying, et al.
Published: (2024)
by: Fu, Ying, et al.
Published: (2024)
Technical Report for CVPR 2024 WeatherProof Dataset Challenge: Semantic Segmentation on Paired Real Data
by: Cao, Guojin, et al.
Published: (2024)
by: Cao, Guojin, et al.
Published: (2024)
Technical Report of NICE Challenge at CVPR 2024: Caption Re-ranking Evaluation Using Ensembled CLIP and Consensus Scores
by: Jeong, Kiyoon, et al.
Published: (2024)
by: Jeong, Kiyoon, et al.
Published: (2024)
Multi-Modal UAV Detection, Classification and Tracking Algorithm -- Technical Report for CVPR 2024 UG2 Challenge
by: Deng, Tianchen, et al.
Published: (2024)
by: Deng, Tianchen, et al.
Published: (2024)
Separating Drone Point Clouds From Complex Backgrounds by Cluster Filter -- Technical Report for CVPR 2024 UG2 Challenge
by: Liang, Hanfang, et al.
Published: (2024)
by: Liang, Hanfang, et al.
Published: (2024)
Technical Report for the 5th CLVision Challenge at CVPR: Addressing the Class-Incremental with Repetition using Unlabeled Data -- 4th Place Solution
by: Moraiti, Panagiota, et al.
Published: (2025)
by: Moraiti, Panagiota, et al.
Published: (2025)
DIVE: Taming DINO for Subject-Driven Video Editing
by: Huang, Yi, et al.
Published: (2024)
by: Huang, Yi, et al.
Published: (2024)
MapVision: CVPR 2024 Autonomous Grand Challenge Mapless Driving Tech Report
by: Yang, Zhongyu, et al.
Published: (2024)
by: Yang, Zhongyu, et al.
Published: (2024)
The Solution for the CVPR2024 NICE Image Captioning Challenge
by: Huang, Longfei, et al.
Published: (2024)
by: Huang, Longfei, et al.
Published: (2024)
WiCV at CVPR 2025: The Women in Computer Vision Workshop
by: Talavera, Estefania, et al.
Published: (2025)
by: Talavera, Estefania, et al.
Published: (2025)
ReferDINO-Plus: 2nd Solution for 4th PVUW MeViS Challenge at CVPR 2025
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
Geospatial Foundational Embedder: Top-1 Winning Solution on EarthVision Embed2Scale Challenge (CVPR 2025)
by: Xu, Zirui, et al.
Published: (2025)
by: Xu, Zirui, et al.
Published: (2025)
The Solution for CVPR2024 Foundational Few-Shot Object Detection Challenge
by: Pan, Hongpeng, et al.
Published: (2024)
by: Pan, Hongpeng, et al.
Published: (2024)
DEAP DIVE: Dataset Investigation with Vision transformers for EEG evaluation
by: Hoffsommer, Annemarie, et al.
Published: (2025)
by: Hoffsommer, Annemarie, et al.
Published: (2025)
SIS-Challenge: Event-based Spatio-temporal Instance Segmentation Challenge at the CVPR 2025 Event-based Vision Workshop
by: Hamann, Friedhelm, et al.
Published: (2025)
by: Hamann, Friedhelm, et al.
Published: (2025)
UnDIVE: Generalized Underwater Video Enhancement Using Generative Priors
by: Srinath, Suhas, et al.
Published: (2024)
by: Srinath, Suhas, et al.
Published: (2024)
WiCV@CVPR2024: The Thirteenth Women In Computer Vision Workshop at the Annual CVPR Conference
by: Aslam, Asra, et al.
Published: (2024)
by: Aslam, Asra, et al.
Published: (2024)
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
by: Li, Yinqi, et al.
Published: (2025)
by: Li, Yinqi, et al.
Published: (2025)
Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
by: Kugo, Noriyuki, et al.
Published: (2025)
by: Kugo, Noriyuki, et al.
Published: (2025)
1st Place Winner of the 2024 Pixel-level Video Understanding in the Wild (CVPR'24 PVUW) Challenge in Video Panoptic Segmentation and Best Long Video Consistency of Video Semantic Segmentation
by: Liu, Qingfeng, et al.
Published: (2024)
by: Liu, Qingfeng, et al.
Published: (2024)
The Solution for the CVPR2023 NICE Image Captioning Challenge
by: Wu, Xiangyu, et al.
Published: (2023)
by: Wu, Xiangyu, et al.
Published: (2023)
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
by: Ma, Guoqing, et al.
Published: (2025)
by: Ma, Guoqing, et al.
Published: (2025)
SuperAD: A Training-free Anomaly Classification and Segmentation Method for CVPR 2025 VAND 3.0 Workshop Challenge Track 1: Adapt & Detect
by: Zhang, Huaiyuan, et al.
Published: (2025)
by: Zhang, Huaiyuan, et al.
Published: (2025)
LongCat-Video Technical Report
by: Meituan LongCat Team, et al.
Published: (2025)
by: Meituan LongCat Team, et al.
Published: (2025)
DIVE: Towards Descriptive and Diverse Visual Commonsense Generation
by: Park, Jun-Hyung, et al.
Published: (2024)
by: Park, Jun-Hyung, et al.
Published: (2024)
A Robust Semantic Segmentation Pipeline for the CVPR 2026 8th UG2+ Challenge Track 2
by: Chai, Jinming, et al.
Published: (2026)
by: Chai, Jinming, et al.
Published: (2026)
EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge
by: Chen, Zhiwei, et al.
Published: (2026)
by: Chen, Zhiwei, et al.
Published: (2026)
3rd Place at CVPR 2026 CASTLE Challenge: Agentic Multi-View Long-Context Video Understanding via Hierarchical Knowledge Graph Retrieval
by: Albusayes, Raghad, et al.
Published: (2026)
by: Albusayes, Raghad, et al.
Published: (2026)
Driving with InternVL: Oustanding Champion in the Track on Driving with Language of the Autonomous Grand Challenge at CVPR 2024
by: Li, Jiahan, et al.
Published: (2024)
by: Li, Jiahan, et al.
Published: (2024)
MetaFood CVPR 2024 Challenge on Physically Informed 3D Food Reconstruction: Methods and Results
by: He, Jiangpeng, et al.
Published: (2024)
by: He, Jiangpeng, et al.
Published: (2024)
HunyuanVideo 1.5 Technical Report
by: Wu, Bing, et al.
Published: (2025)
by: Wu, Bing, et al.
Published: (2025)
LSVOS 2025 Challenge Report: Recent Advances in Complex Video Object Segmentation
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Technical Report for Argoverse2 Scenario Mining Challenges on Iterative Error Correction and Spatially-Aware Prompting
by: Chen, Yifei, et al.
Published: (2025)
by: Chen, Yifei, et al.
Published: (2025)
1st Place Solution for MOSE Track in CVPR 2024 PVUW Workshop: Complex Video Object Segmentation
by: Miao, Deshui, et al.
Published: (2024)
by: Miao, Deshui, et al.
Published: (2024)
2nd Place Solution for MOSE Track in CVPR 2024 PVUW workshop: Complex Video Object Segmentation
by: Xu, Zhensong, et al.
Published: (2024)
by: Xu, Zhensong, et al.
Published: (2024)
A Two-Stage Adverse Weather Semantic Segmentation Method for WeatherProof Challenge CVPR 2024 Workshop UG2+
by: Wang, Jianzhao, et al.
Published: (2024)
by: Wang, Jianzhao, et al.
Published: (2024)
PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild
by: Ding, Henghui, et al.
Published: (2025)
by: Ding, Henghui, et al.
Published: (2025)
Similar Items
-
VDMA: Video Question Answering with Dynamically Generated Multi-Agents
by: Kugo, Noriyuki, et al.
Published: (2024) -
Multimodal Structured Generation: CVPR's 2nd MMFM Challenge Technical Report
by: Cesista, Franz Louis
Published: (2024) -
Technique Report of CVPR 2024 PBDL Challenges
by: Fu, Ying, et al.
Published: (2024) -
Technical Report for CVPR 2024 WeatherProof Dataset Challenge: Semantic Segmentation on Paired Real Data
by: Cao, Guojin, et al.
Published: (2024) -
Technical Report of NICE Challenge at CVPR 2024: Caption Re-ranking Evaluation Using Ensembled CLIP and Consensus Scores
by: Jeong, Kiyoon, et al.
Published: (2024)