TimeLogic: A Temporal Logic Benchmark for Video QA
Fuente:
arXiv
Saved in:
| Main Authors: | Swetha, Sirnam, Kuehne, Hilde, Shah, Mubarak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
by: Kim, Younggun, et al.
Published: (2025)
by: Kim, Younggun, et al.
Published: (2025)
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
by: Siddiqui, Nyle, et al.
Published: (2025)
by: Siddiqui, Nyle, et al.
Published: (2025)
TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval
by: Shatwell, David G., et al.
Published: (2026)
by: Shatwell, David G., et al.
Published: (2026)
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
by: Narnaware, Vishal, et al.
Published: (2025)
by: Narnaware, Vishal, et al.
Published: (2025)
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space
by: Shatwell, David G., et al.
Published: (2025)
by: Shatwell, David G., et al.
Published: (2025)
The Telephone Game: Evaluating Semantic Drift in Unified Models
by: Mollah, Sabbir, et al.
Published: (2025)
by: Mollah, Sabbir, et al.
Published: (2025)
Convolutional Differentiable Logic Gate Networks
by: Petersen, Felix, et al.
Published: (2024)
by: Petersen, Felix, et al.
Published: (2024)
NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning
by: Shah, Sahil, et al.
Published: (2025)
by: Shah, Sahil, et al.
Published: (2025)
Temporal Evidence Routing with Structured Visual Evidence for TimeLogicQA
by: Sun, Yuyang, et al.
Published: (2026)
by: Sun, Yuyang, et al.
Published: (2026)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
by: Swetha, Sirnam, et al.
Published: (2024)
by: Swetha, Sirnam, et al.
Published: (2024)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
by: Gupta, Animesh, et al.
Published: (2025)
by: Gupta, Animesh, et al.
Published: (2025)
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
by: Miao, Bo, et al.
Published: (2024)
by: Miao, Bo, et al.
Published: (2024)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
by: Shvetsova, Nina, et al.
Published: (2023)
by: Shvetsova, Nina, et al.
Published: (2023)
LogicQA: Logical Anomaly Detection with Vision Language Model Generated Questions
by: Kwon, Yejin, et al.
Published: (2025)
by: Kwon, Yejin, et al.
Published: (2025)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
by: Jahagirdar, Soumya, et al.
Published: (2026)
by: Jahagirdar, Soumya, et al.
Published: (2026)
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
by: Chen, Brian, et al.
Published: (2023)
by: Chen, Brian, et al.
Published: (2023)
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
by: Kulkarni, Parth Parag, et al.
Published: (2026)
by: Kulkarni, Parth Parag, et al.
Published: (2026)
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025)
by: Vogel, Felix, et al.
Published: (2025)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
by: Li, Kunyang, et al.
Published: (2026)
by: Li, Kunyang, et al.
Published: (2026)
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
by: Bousselham, Walid, et al.
Published: (2026)
by: Bousselham, Walid, et al.
Published: (2026)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
Dual Guidance Semi-Supervised Action Detection
by: Singh, Ankit, et al.
Published: (2025)
by: Singh, Ankit, et al.
Published: (2025)
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
by: Bousselham, Walid, et al.
Published: (2024)
by: Bousselham, Walid, et al.
Published: (2024)
LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity
by: Bousselham, Walid, et al.
Published: (2024)
by: Bousselham, Walid, et al.
Published: (2024)
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
by: Fioresi, Joseph, et al.
Published: (2025)
by: Fioresi, Joseph, et al.
Published: (2025)
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
by: Pillai, Manu S, et al.
Published: (2024)
by: Pillai, Manu S, et al.
Published: (2024)
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
by: Song, Zijie, et al.
Published: (2025)
by: Song, Zijie, et al.
Published: (2025)
WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition
by: Bock, Marius, et al.
Published: (2023)
by: Bock, Marius, et al.
Published: (2023)
CityGuessr: City-Level Video Geo-Localization on a Global Scale
by: Kulkarni, Parth Parag, et al.
Published: (2024)
by: Kulkarni, Parth Parag, et al.
Published: (2024)
TempCore: Are Video QA Benchmarks Temporally Grounded? A Frame Selection Sensitivity Analysis and Benchmark
by: Ok, Hyunjong, et al.
Published: (2025)
by: Ok, Hyunjong, et al.
Published: (2025)
Lost in Time: A New Temporal Benchmark for VideoLLMs
by: Cores, Daniel, et al.
Published: (2024)
by: Cores, Daniel, et al.
Published: (2024)
Investigating Memorization in Video Diffusion Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
by: Rasheed, Hanoona, et al.
Published: (2025)
by: Rasheed, Hanoona, et al.
Published: (2025)
Leveraging Pre-Trained Visual Models for AI-Generated Video Detection
by: Veeramachaneni, Keerthi, et al.
Published: (2025)
by: Veeramachaneni, Keerthi, et al.
Published: (2025)
Similar Items
-
Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
by: Kim, Younggun, et al.
Published: (2025) -
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
by: Siddiqui, Nyle, et al.
Published: (2025) -
TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval
by: Shatwell, David G., et al.
Published: (2026) -
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
by: Narnaware, Vishal, et al.
Published: (2025) -
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
by: Swetha, Sirnam, et al.
Published: (2025)