LVBench: An Extreme Long Video Understanding Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Weihan, He, Zehai, Hong, Wenyi, Cheng, Yean, Zhang, Xiaohan, Qi, Ji, Gu, Xiaotao, Huang, Shiyu, Xu, Bin, Dong, Yuxiao, Ding, Ming, Tang, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
von: Hong, Wenyi, et al.
Veröffentlicht: (2025)
von: Hong, Wenyi, et al.
Veröffentlicht: (2025)
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2024)
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2024)
CogVLM2: Visual Language Models for Image and Video Understanding
von: Hong, Wenyi, et al.
Veröffentlicht: (2024)
von: Hong, Wenyi, et al.
Veröffentlicht: (2024)
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
von: He, Zehai, et al.
Veröffentlicht: (2026)
von: He, Zehai, et al.
Veröffentlicht: (2026)
DreamPolish: Domain Score Distillation With Progressive Geometry Generation
von: Cheng, Yean, et al.
Veröffentlicht: (2024)
von: Cheng, Yean, et al.
Veröffentlicht: (2024)
CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
von: Qi, Ji, et al.
Veröffentlicht: (2024)
von: Qi, Ji, et al.
Veröffentlicht: (2024)
CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
von: Zheng, Wendi, et al.
Veröffentlicht: (2024)
von: Zheng, Wendi, et al.
Veröffentlicht: (2024)
UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
LongSafety: Evaluating Long-Context Safety of Large Language Models
von: Lu, Yida, et al.
Veröffentlicht: (2025)
von: Lu, Yida, et al.
Veröffentlicht: (2025)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
von: Gui, Jiayi, et al.
Veröffentlicht: (2024)
CogAgent: A Visual Language Model for GUI Agents
von: Hong, Wenyi, et al.
Veröffentlicht: (2023)
von: Hong, Wenyi, et al.
Veröffentlicht: (2023)
CogVLM: Visual Expert for Pretrained Language Models
von: Wang, Weihan, et al.
Veröffentlicht: (2023)
von: Wang, Weihan, et al.
Veröffentlicht: (2023)
Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2024)
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2024)
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
von: Bai, Yushi, et al.
Veröffentlicht: (2023)
von: Bai, Yushi, et al.
Veröffentlicht: (2023)
The number of irreducibles in the plethysm $s_λ[s_m]$
von: Lim, Ming Yean
Veröffentlicht: (2025)
von: Lim, Ming Yean
Veröffentlicht: (2025)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
von: Zhou, Wenqi, et al.
Veröffentlicht: (2025)
von: Zhou, Wenqi, et al.
Veröffentlicht: (2025)
WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
von: Xu, Mingde, et al.
Veröffentlicht: (2025)
von: Xu, Mingde, et al.
Veröffentlicht: (2025)
VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
von: Xu, Jiazheng, et al.
Veröffentlicht: (2024)
von: Xu, Jiazheng, et al.
Veröffentlicht: (2024)
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
von: Zhong, Yangyang, et al.
Veröffentlicht: (2025)
von: Zhong, Yangyang, et al.
Veröffentlicht: (2025)
Understanding Emergent Abilities of Language Models from the Loss Perspective
von: Du, Zhengxiao, et al.
Veröffentlicht: (2024)
von: Du, Zhengxiao, et al.
Veröffentlicht: (2024)
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
Glyph: Scaling Context Windows via Visual-Text Compression
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
ELITE LST: FY-3E/MERSI Clear-Sky LST & E at Dawn and Dusk over CONUS (2023.05-2024.04)
von: Liu, Weihan, et al.
Veröffentlicht: (2025)
von: Liu, Weihan, et al.
Veröffentlicht: (2025)
InstructionBench: An Instructional Video Understanding Benchmark
von: Wei, Haiwan, et al.
Veröffentlicht: (2025)
von: Wei, Haiwan, et al.
Veröffentlicht: (2025)
LongAlign: A Recipe for Long Context Alignment of Large Language Models
von: Bai, Yushi, et al.
Veröffentlicht: (2024)
von: Bai, Yushi, et al.
Veröffentlicht: (2024)
MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
von: Liu, Shizhan, et al.
Veröffentlicht: (2025)
von: Liu, Shizhan, et al.
Veröffentlicht: (2025)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
von: Cheng, Jiale, et al.
Veröffentlicht: (2024)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
von: Wang, Youze, et al.
Veröffentlicht: (2025)
von: Wang, Youze, et al.
Veröffentlicht: (2025)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
von: Nagrani, Arsha, et al.
Veröffentlicht: (2024)
von: Nagrani, Arsha, et al.
Veröffentlicht: (2024)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
Temporal Preference Optimization for Long-Form Video Understanding
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Does Negative Sampling Matter? A Review with Insights into its Theory and Applications
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
von: Hong, Wenyi, et al.
Veröffentlicht: (2025) -
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2024) -
CogVLM2: Visual Language Models for Image and Video Understanding
von: Hong, Wenyi, et al.
Veröffentlicht: (2024) -
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
von: He, Zehai, et al.
Veröffentlicht: (2026) -
DreamPolish: Domain Score Distillation With Progressive Geometry Generation
von: Cheng, Yean, et al.
Veröffentlicht: (2024)