The Solution for the ICCV 2023 Perception Test Challenge 2023 -- Task 6 -- Grounded videoQA
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Hailiang, Chao, Dian, Guan, Zhihao, Yang, Yang |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Solution for Point Tracking Task of ICCV 1st Perception Test Challenge 2023
par: Pan, Hongpeng, et autres
Publié: (2024)
par: Pan, Hongpeng, et autres
Publié: (2024)
The Solution for Temporal Sound Localisation Task of ICCV 1st Perception Test Challenge 2023
par: Huang, Yurui, et autres
Publié: (2024)
par: Huang, Yurui, et autres
Publié: (2024)
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge
par: Chao, Dian, et autres
Publié: (2024)
par: Chao, Dian, et autres
Publié: (2024)
Solution for OOD-CV Workshop SSB Challenge 2024 (Open-Set Recognition Track)
par: Feng, Mingxu, et autres
Publié: (2024)
par: Feng, Mingxu, et autres
Publié: (2024)
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge
par: Peng, Yingzhe, et autres
Publié: (2024)
par: Peng, Yingzhe, et autres
Publié: (2024)
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
par: Huang, Longfei, et autres
Publié: (2024)
par: Huang, Longfei, et autres
Publié: (2024)
The Solution for Language-Enhanced Image New Category Discovery
par: Xu, Haonan, et autres
Publié: (2024)
par: Xu, Haonan, et autres
Publié: (2024)
1st Place Solution for ICCV 2023 OmniObject3D Challenge: Sparse-View Reconstruction
par: Du, Hang, et autres
Publié: (2024)
par: Du, Hang, et autres
Publié: (2024)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
par: Heyward, Joseph, et autres
Publié: (2024)
par: Heyward, Joseph, et autres
Publié: (2024)
The Solution for the CVPR2023 NICE Image Captioning Challenge
par: Wu, Xiangyu, et autres
Publié: (2023)
par: Wu, Xiangyu, et autres
Publié: (2023)
The Solution for Single Object Tracking Task of Perception Test Challenge 2024
par: Zhong, Zhiqiang, et autres
Publié: (2024)
par: Zhong, Zhiqiang, et autres
Publié: (2024)
Solution for Point Tracking Task of ECCV 2nd Perception Test Challenge 2024
par: Zhang, Yuxuan, et autres
Publié: (2024)
par: Zhang, Yuxuan, et autres
Publié: (2024)
The Solution for the CVPR 2023 1st foundation model challenge-Track2
par: Xu, Haonan, et autres
Publié: (2024)
par: Xu, Haonan, et autres
Publié: (2024)
The Solution for the GAIIC2024 RGB-TIR object detection Challenge
par: Wu, Xiangyu, et autres
Publié: (2024)
par: Wu, Xiangyu, et autres
Publié: (2024)
Neural Material Adaptor for Visual Grounding of Intrinsic Dynamics
par: Cao, Junyi, et autres
Publié: (2024)
par: Cao, Junyi, et autres
Publié: (2024)
3rd Place Solution to ICCV LargeFineFoodAI Retrieval
par: Zhong, Yang, et autres
Publié: (2025)
par: Zhong, Yang, et autres
Publié: (2025)
Universal Lesion Segmentation Challenge 2023: A Comparative Research of Different Algorithms
par: Shi, Kaiwen, et autres
Publié: (2025)
par: Shi, Kaiwen, et autres
Publié: (2025)
The Solution for Temporal Action Localisation Task of Perception Test Challenge 2024
par: Han, Yinan, et autres
Publié: (2024)
par: Han, Yinan, et autres
Publié: (2024)
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
par: Wang, Xiyao, et autres
Publié: (2025)
par: Wang, Xiyao, et autres
Publié: (2025)
WoodScape Motion Segmentation for Autonomous Driving -- CVPR 2023 OmniCV Workshop Challenge
par: Ramachandran, Saravanabalagi, et autres
Publié: (2023)
par: Ramachandran, Saravanabalagi, et autres
Publié: (2023)
Look, Remember and Reason: Grounded reasoning in videos with language models
par: Bhattacharyya, Apratim, et autres
Publié: (2023)
par: Bhattacharyya, Apratim, et autres
Publié: (2023)
MVP: Winning Solution to SMP Challenge 2025 Video Track
par: Ye, Liliang, et autres
Publié: (2025)
par: Ye, Liliang, et autres
Publié: (2025)
Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge
par: Wu, Xiangyu, et autres
Publié: (2024)
par: Wu, Xiangyu, et autres
Publié: (2024)
TempCore: Are Video QA Benchmarks Temporally Grounded? A Frame Selection Sensitivity Analysis and Benchmark
par: Ok, Hyunjong, et autres
Publié: (2025)
par: Ok, Hyunjong, et autres
Publié: (2025)
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
par: Safwan, Itbaan, et autres
Publié: (2025)
par: Safwan, Itbaan, et autres
Publié: (2025)
Analysis of the BraTS 2023 Intracranial Meningioma Segmentation Challenge
par: LaBella, Dominic, et autres
Publié: (2024)
par: LaBella, Dominic, et autres
Publié: (2024)
Many Perception Tasks are Highly Redundant Functions of their Input Data
par: Ramesh, Rahul, et autres
Publié: (2024)
par: Ramesh, Rahul, et autres
Publié: (2024)
SGW-based Multi-Task Learning in Vision Tasks
par: Zhang, Ruiyuan, et autres
Publié: (2024)
par: Zhang, Ruiyuan, et autres
Publié: (2024)
Test-Time Canonicalization by Foundation Models for Robust Perception
par: Singhal, Utkarsh, et autres
Publié: (2025)
par: Singhal, Utkarsh, et autres
Publié: (2025)
Made to Order: Discovering monotonic temporal changes via self-supervised video ordering
par: Yang, Charig, et autres
Publié: (2024)
par: Yang, Charig, et autres
Publié: (2024)
Zero-Shot Neural Architecture Search: Challenges, Solutions, and Opportunities
par: Li, Guihong, et autres
Publié: (2023)
par: Li, Guihong, et autres
Publié: (2023)
Towards Real-world Debiasing: Rethinking Evaluation, Challenge, and Solution
par: Kuang, Peng, et autres
Publié: (2024)
par: Kuang, Peng, et autres
Publié: (2024)
Deep Learning Meets OBIA: Tasks, Challenges, Strategies, and Perspectives
par: Ma, Lei, et autres
Publié: (2024)
par: Ma, Lei, et autres
Publié: (2024)
Salient Temporal Encoding for Dynamic Scene Graph Generation
par: Zhu, Zhihao
Publié: (2025)
par: Zhu, Zhihao
Publié: (2025)
Self-Improving Small Object Grounding in LVLMs
par: Yang, Tianze, et autres
Publié: (2026)
par: Yang, Tianze, et autres
Publié: (2026)
MVAR: MultiVariate AutoRegressive Air Pollutants Forecasting Model
par: Fan, Xu, et autres
Publié: (2025)
par: Fan, Xu, et autres
Publié: (2025)
PCoTTA: Continual Test-Time Adaptation for Multi-Task Point Cloud Understanding
par: Jiang, Jincen, et autres
Publié: (2024)
par: Jiang, Jincen, et autres
Publié: (2024)
Hierarchical Invariance for Robust and Interpretable Vision Tasks at Larger Scales
par: Qi, Shuren, et autres
Publié: (2024)
par: Qi, Shuren, et autres
Publié: (2024)
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
par: Liang, Zichen, et autres
Publié: (2025)
par: Liang, Zichen, et autres
Publié: (2025)
Medical Image Spatial Grounding with Semantic Sampling
par: Yu, Andrew Seohwan, et autres
Publié: (2026)
par: Yu, Andrew Seohwan, et autres
Publié: (2026)
Documents similaires
-
Solution for Point Tracking Task of ICCV 1st Perception Test Challenge 2023
par: Pan, Hongpeng, et autres
Publié: (2024) -
The Solution for Temporal Sound Localisation Task of ICCV 1st Perception Test Challenge 2023
par: Huang, Yurui, et autres
Publié: (2024) -
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge
par: Chao, Dian, et autres
Publié: (2024) -
Solution for OOD-CV Workshop SSB Challenge 2024 (Open-Set Recognition Track)
par: Feng, Mingxu, et autres
Publié: (2024) -
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge
par: Peng, Yingzhe, et autres
Publié: (2024)