Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Meng, Hu, Pengfei, Wang, Yingyao, Gu, Jihao, Tang, Haoran, Zhao, Haoze, Wang, Chen, Dong, Jiahua, Yu, Wangbo, Zhang, Ge, Song, Jun, Li, Xiang, Zheng, Bo, Reid, Ian, Liang, Xiaodan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
von: He, Yancheng, et al.
Veröffentlicht: (2024)
von: He, Yancheng, et al.
Veröffentlicht: (2024)
SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
von: Haas, Lukas, et al.
Veröffentlicht: (2025)
von: Haas, Lukas, et al.
Veröffentlicht: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2026)
von: Cao, Meng, et al.
Veröffentlicht: (2026)
Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
von: Gu, Jihao, et al.
Veröffentlicht: (2024)
von: Gu, Jihao, et al.
Veröffentlicht: (2024)
CodeSimpleQA: Scaling Factuality in Code Large Language Models
von: Yang, Jian, et al.
Veröffentlicht: (2025)
von: Yang, Jian, et al.
Veröffentlicht: (2025)
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
von: Cao, Meng, et al.
Veröffentlicht: (2025)
von: Cao, Meng, et al.
Veröffentlicht: (2025)
Video Spatial Reasoning with Object-Centric 3D Rollout
von: Tang, Haoran, et al.
Veröffentlicht: (2025)
von: Tang, Haoran, et al.
Veröffentlicht: (2025)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
FunQA: Towards Surprising Video Comprehension
von: Xie, Binzhu, et al.
Veröffentlicht: (2023)
von: Xie, Binzhu, et al.
Veröffentlicht: (2023)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
Text-guided Fine-Grained Video Anomaly Understanding
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
Hierarchical Memory for Long Video QA
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
von: Ko, Donghyeon, et al.
Veröffentlicht: (2025)
von: Ko, Donghyeon, et al.
Veröffentlicht: (2025)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
AdsQA: Towards Advertisement Video Understanding
von: Long, Xinwei, et al.
Veröffentlicht: (2025)
von: Long, Xinwei, et al.
Veröffentlicht: (2025)
SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
von: Cao, Meng, et al.
Veröffentlicht: (2025)
von: Cao, Meng, et al.
Veröffentlicht: (2025)
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
von: Cao, Meng, et al.
Veröffentlicht: (2025)
von: Cao, Meng, et al.
Veröffentlicht: (2025)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning
von: Xu, Liangyu, et al.
Veröffentlicht: (2025)
von: Xu, Liangyu, et al.
Veröffentlicht: (2025)
VideoQA-SC: Adaptive Semantic Communication for Video Question Answering
von: Guo, Jiangyuan, et al.
Veröffentlicht: (2024)
von: Guo, Jiangyuan, et al.
Veröffentlicht: (2024)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
Team of One: Cracking Complex Video QA with Model Synergy
von: Xie, Jun, et al.
Veröffentlicht: (2025)
von: Xie, Jun, et al.
Veröffentlicht: (2025)
DAM: Dynamic Adapter Merging for Continual Video QA Learning
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
von: Cheng, Feng, et al.
Veröffentlicht: (2024)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
von: Wang, Cong, et al.
Veröffentlicht: (2023)
von: Wang, Cong, et al.
Veröffentlicht: (2023)
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
Artemis: Towards Referential Understanding in Complex Videos
von: Qiu, Jihao, et al.
Veröffentlicht: (2024)
von: Qiu, Jihao, et al.
Veröffentlicht: (2024)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
SimpleGVR: A Simple Baseline for Latent-Cascaded Video Super-Resolution
von: Xie, Liangbin, et al.
Veröffentlicht: (2025)
von: Xie, Liangbin, et al.
Veröffentlicht: (2025)
YTCommentQA: Video Question Answerability in Instructional Videos
von: Yang, Saelyne, et al.
Veröffentlicht: (2024)
von: Yang, Saelyne, et al.
Veröffentlicht: (2024)
MF2Summ: Multimodal Fusion for Video Summarization with Temporal Alignment
von: wang, Shuo, et al.
Veröffentlicht: (2025)
von: wang, Shuo, et al.
Veröffentlicht: (2025)
MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval
von: Tang, Haoran, et al.
Veröffentlicht: (2024)
von: Tang, Haoran, et al.
Veröffentlicht: (2024)
A Simple Low-bit Quantization Framework for Video Snapshot Compressive Imaging
von: Cao, Miao, et al.
Veröffentlicht: (2024)
von: Cao, Miao, et al.
Veröffentlicht: (2024)
ENTER: Event Based Interpretable Reasoning for VideoQA
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
von: Ai, Qihang, et al.
Veröffentlicht: (2025)
von: Ai, Qihang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
von: He, Yancheng, et al.
Veröffentlicht: (2024) -
SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
von: Haas, Lukas, et al.
Veröffentlicht: (2025) -
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
von: Hu, Pengfei, et al.
Veröffentlicht: (2025) -
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2024) -
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2026)