Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Lanxiang, Shankarampeta, Abhilash, Huang, Yixin, Dai, Zilin, Yu, Haoyang, Zhao, Yujie, Kang, Haoqiang, Zhao, Daniel, Rosing, Tajana, Zhang, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
by: Zhao, Daniel, et al.
Published: (2025)
by: Zhao, Daniel, et al.
Published: (2025)
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
by: Hu, Lanxiang, et al.
Published: (2024)
by: Hu, Lanxiang, et al.
Published: (2024)
lmgame-Bench: How Good are LLMs at Playing Games?
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
by: Zhao, Yujie, et al.
Published: (2026)
by: Zhao, Yujie, et al.
Published: (2026)
Evidence-Guided Schema Normalization for Temporal Tabular Reasoning
by: Thanga, Ashish, et al.
Published: (2025)
by: Thanga, Ashish, et al.
Published: (2025)
Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
CAN-QA: A Question-Answering Benchmark for Reasoning over In-Vehicle CAN Traffic
by: Chen, Jing, et al.
Published: (2026)
by: Chen, Jing, et al.
Published: (2026)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
by: Deng, Andong, et al.
Published: (2025)
by: Deng, Andong, et al.
Published: (2025)
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
by: Reichman, Benjamin, et al.
Published: (2025)
by: Reichman, Benjamin, et al.
Published: (2025)
LifeAgentBench: A Multi-dimensional Benchmark and Agent for Personal Health Assistants in Digital Health
by: Tian, Ye, et al.
Published: (2026)
by: Tian, Ye, et al.
Published: (2026)
TransientTables: Evaluating LLMs' Reasoning on Temporally Evolving Semi-structured Tables
by: Shankarampeta, Abhilash, et al.
Published: (2025)
by: Shankarampeta, Abhilash, et al.
Published: (2025)
MicroHD: An Accuracy-Driven Optimization of Hyperdimensional Computing Algorithms for TinyML systems
by: Ponzina, Flavio, et al.
Published: (2024)
by: Ponzina, Flavio, et al.
Published: (2024)
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
by: Xun, Shuhang, et al.
Published: (2025)
by: Xun, Shuhang, et al.
Published: (2025)
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
by: Kong, Fanqi, et al.
Published: (2025)
by: Kong, Fanqi, et al.
Published: (2025)
SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions
by: Yu, Xiaofan, et al.
Published: (2025)
by: Yu, Xiaofan, et al.
Published: (2025)
CyberMaskQA: A Privacy-Aware Benchmark for Evaluating Large Language Models in Cybersecurity Question Answering
by: Gaddi, Matilda, et al.
Published: (2026)
by: Gaddi, Matilda, et al.
Published: (2026)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
FedUHD: Unsupervised Federated Learning using Hyperdimensional Computing
by: Lee, You Hak, et al.
Published: (2025)
by: Lee, You Hak, et al.
Published: (2025)
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
by: Wang, Xiaodong, et al.
Published: (2026)
by: Wang, Xiaodong, et al.
Published: (2026)
InstructionBench: An Instructional Video Understanding Benchmark
by: Wei, Haiwan, et al.
Published: (2025)
by: Wei, Haiwan, et al.
Published: (2025)
ACORN-IDS: Adaptive Continual Novelty Detection for Intrusion Detection Systems
by: Fuhrman, Sean, et al.
Published: (2026)
by: Fuhrman, Sean, et al.
Published: (2026)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
by: Wang, Xuan, et al.
Published: (2024)
by: Wang, Xuan, et al.
Published: (2024)
FaTRQ: Tiered Residual Quantization for LLM Vector Search in Far-Memory-Aware ANNS Systems
by: Zhang, Tianqi, et al.
Published: (2026)
by: Zhang, Tianqi, et al.
Published: (2026)
SpANNS: Optimizing Approximate Nearest Neighbor Search for Sparse Vectors Using Near Memory Processing
by: Zhang, Tianqi, et al.
Published: (2026)
by: Zhang, Tianqi, et al.
Published: (2026)
CND-IDS: Continual Novelty Detection for Intrusion Detection Systems
by: Fuhrman, Sean, et al.
Published: (2025)
by: Fuhrman, Sean, et al.
Published: (2025)
V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
by: Luo, Yang, et al.
Published: (2025)
by: Luo, Yang, et al.
Published: (2025)
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
VideoMarkBench: Benchmarking Robustness of Video Watermarking
by: Jiang, Zhengyuan, et al.
Published: (2025)
by: Jiang, Zhengyuan, et al.
Published: (2025)
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
by: Lu, Tsung-Han, et al.
Published: (2026)
by: Lu, Tsung-Han, et al.
Published: (2026)
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
by: Feng, Hengyi, et al.
Published: (2026)
by: Feng, Hengyi, et al.
Published: (2026)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
by: Han, ZhaoYang, et al.
Published: (2025)
by: Han, ZhaoYang, et al.
Published: (2025)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
by: Wu, Haoning, et al.
Published: (2024)
by: Wu, Haoning, et al.
Published: (2024)
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
by: Zhang, Yuanhan, et al.
Published: (2025)
by: Zhang, Yuanhan, et al.
Published: (2025)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
by: Xu, Yicheng, et al.
Published: (2025)
by: Xu, Yicheng, et al.
Published: (2025)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
by: Zhou, Wenqi, et al.
Published: (2025)
by: Zhou, Wenqi, et al.
Published: (2025)
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
by: Zhao, Baining, et al.
Published: (2025)
by: Zhao, Baining, et al.
Published: (2025)
Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMs
by: Zhao, Yujie, et al.
Published: (2025)
by: Zhao, Yujie, et al.
Published: (2025)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
by: Pan, Yaning, et al.
Published: (2025)
by: Pan, Yaning, et al.
Published: (2025)
CITADEL: Continual Anomaly Detection for Enhanced Learning in IoT Intrusion Detection
by: Li, Elvin, et al.
Published: (2025)
by: Li, Elvin, et al.
Published: (2025)
Similar Items
-
Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
by: Zhao, Daniel, et al.
Published: (2025) -
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
by: Hu, Lanxiang, et al.
Published: (2024) -
lmgame-Bench: How Good are LLMs at Playing Games?
by: Hu, Lanxiang, et al.
Published: (2025) -
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
by: Zhao, Yujie, et al.
Published: (2026) -
Evidence-Guided Schema Normalization for Temporal Tabular Reasoning
by: Thanga, Ashish, et al.
Published: (2025)