Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Heyward, Joseph, Carreira, João, Damen, Dima, Zisserman, Andrew, Pătrăucean, Viorica |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Perception Test 2025: Challenge Summary and a Unified VQA Extension
von: Heyward, Joseph, et al.
Veröffentlicht: (2026)
von: Heyward, Joseph, et al.
Veröffentlicht: (2026)
Learning from Streaming Video with Orthogonal Gradients
von: Han, Tengda, et al.
Veröffentlicht: (2025)
von: Han, Tengda, et al.
Veröffentlicht: (2025)
Learning from One Continuous Video Stream
von: Carreira, João, et al.
Veröffentlicht: (2023)
von: Carreira, João, et al.
Veröffentlicht: (2023)
A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
von: Papalampidi, Pinelopi, et al.
Veröffentlicht: (2023)
von: Papalampidi, Pinelopi, et al.
Veröffentlicht: (2023)
Unique Lives, Shared World: Learning from Single-Life Videos
von: Han, Tengda, et al.
Veröffentlicht: (2025)
von: Han, Tengda, et al.
Veröffentlicht: (2025)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
von: Perrett, Toby, et al.
Veröffentlicht: (2024)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
Seeing without Pixels: Perception from Camera Trajectories
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
TIM: A Time Interval Machine for Audio-Visual Action Recognition
von: Chalk, Jacob, et al.
Veröffentlicht: (2024)
von: Chalk, Jacob, et al.
Veröffentlicht: (2024)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
Reading Between the Lanes: Text VideoQA on the Road
von: Tom, George, et al.
Veröffentlicht: (2023)
von: Tom, George, et al.
Veröffentlicht: (2023)
VideoQA in the Era of LLMs: An Empirical Study
von: Xiao, Junbin, et al.
Veröffentlicht: (2024)
von: Xiao, Junbin, et al.
Veröffentlicht: (2024)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
Get a Grip: Reconstructing Hand-Object Stable Grasps in Egocentric Videos
von: Zhu, Zhifan, et al.
Veröffentlicht: (2023)
von: Zhu, Zhifan, et al.
Veröffentlicht: (2023)
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
von: Koppula, Skanda, et al.
Veröffentlicht: (2024)
von: Koppula, Skanda, et al.
Veröffentlicht: (2024)
TRecViT: A Recurrent Video Transformer
von: Pătrăucean, Viorica, et al.
Veröffentlicht: (2024)
von: Pătrăucean, Viorica, et al.
Veröffentlicht: (2024)
Understanding Complexity in VideoQA via Visual Program Generation
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
ENTER: Event Based Interpretable Reasoning for VideoQA
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
Dynamic Reflections: Probing Video Representations with Text Alignment
von: Zhu, Tyler, et al.
Veröffentlicht: (2025)
von: Zhu, Tyler, et al.
Veröffentlicht: (2025)
QTG-VQA: Question-Type-Guided Architectural for VideoQA Systems
von: He, Zhixian, et al.
Veröffentlicht: (2024)
von: He, Zhixian, et al.
Veröffentlicht: (2024)
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
von: Parikh, Chirag, et al.
Veröffentlicht: (2025)
von: Parikh, Chirag, et al.
Veröffentlicht: (2025)
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
von: Fragomeni, Adriano, et al.
Veröffentlicht: (2025)
von: Fragomeni, Adriano, et al.
Veröffentlicht: (2025)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
von: Fragomeni, Adriano, et al.
Veröffentlicht: (2025)
von: Fragomeni, Adriano, et al.
Veröffentlicht: (2025)
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
von: Flanagan, Kevin, et al.
Veröffentlicht: (2025)
von: Flanagan, Kevin, et al.
Veröffentlicht: (2025)
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
von: Song, Zijie, et al.
Veröffentlicht: (2025)
von: Song, Zijie, et al.
Veröffentlicht: (2025)
Moment Sampling in Video LLMs for Long-Form Video QA
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
von: Rawal, Ishaan Singh, et al.
Veröffentlicht: (2023)
von: Rawal, Ishaan Singh, et al.
Veröffentlicht: (2023)
Context-Aware Multimodal Pretraining
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
VideoQA-SC: Adaptive Semantic Communication for Video Question Answering
von: Guo, Jiangyuan, et al.
Veröffentlicht: (2024)
von: Guo, Jiangyuan, et al.
Veröffentlicht: (2024)
Recurrent Video Masked Autoencoders
von: Zoran, Daniel, et al.
Veröffentlicht: (2025)
von: Zoran, Daniel, et al.
Veröffentlicht: (2025)
Every Shot Counts: Using Exemplars for Repetition Counting in Videos
von: Sinha, Saptarshi, et al.
Veröffentlicht: (2024)
von: Sinha, Saptarshi, et al.
Veröffentlicht: (2024)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
von: Liang, Lili, et al.
Veröffentlicht: (2024)
von: Liang, Lili, et al.
Veröffentlicht: (2024)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
TempCore: Are Video QA Benchmarks Temporally Grounded? A Frame Selection Sensitivity Analysis and Benchmark
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Perception Test 2025: Challenge Summary and a Unified VQA Extension
von: Heyward, Joseph, et al.
Veröffentlicht: (2026) -
Learning from Streaming Video with Orthogonal Gradients
von: Han, Tengda, et al.
Veröffentlicht: (2025) -
Learning from One Continuous Video Stream
von: Carreira, João, et al.
Veröffentlicht: (2023) -
A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
von: Papalampidi, Pinelopi, et al.
Veröffentlicht: (2023) -
Unique Lives, Shared World: Learning from Single-Life Videos
von: Han, Tengda, et al.
Veröffentlicht: (2025)