ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, David, Yuan, Huaqing, Wang, Xingjian, Zang, Qianbo, Liu, Tianci, He, Xinyang, Wei, Yanbin, Guo, Jiawei, Jiahui, Ni, Yang, Zhenzhu, Cao, Meng, Quan, Shanghaoran, Li, Yizhi, Zhou, Wangchunshu, Liu, Jiaheng, Huang, Wenhao, Zhang, Ge, Ni, Shiwen, Jin, Xiaojie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DMoERM: Recipes of Mixture-of-Experts for Effective Reward Modeling
by: Quan, Shanghaoran
Published: (2024)
by: Quan, Shanghaoran
Published: (2024)
Automatically Generating Numerous Context-Driven SFT Data for LLMs across Diverse Granularity
by: Quan, Shanghaoran
Published: (2024)
by: Quan, Shanghaoran
Published: (2024)
AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions
by: Li, Ziming, et al.
Published: (2024)
by: Li, Ziming, et al.
Published: (2024)
LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm
by: Wu, Siwei, et al.
Published: (2025)
by: Wu, Siwei, et al.
Published: (2025)
LongIns: A Challenging Long-context Instruction-based Exam for LLMs
by: Gavin, Shawn, et al.
Published: (2024)
by: Gavin, Shawn, et al.
Published: (2024)
CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
by: Li, Siyi, et al.
Published: (2026)
by: Li, Siyi, et al.
Published: (2026)
Educational-Psychological Dialogue Robot Based on Multi-Agent Collaboration
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
Language Models can Self-Lengthen to Generate Long Texts
by: Quan, Shanghaoran, et al.
Published: (2024)
by: Quan, Shanghaoran, et al.
Published: (2024)
LIME: Less Is More for MLLM Evaluation
by: Zhu, King, et al.
Published: (2024)
by: Zhu, King, et al.
Published: (2024)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
by: Gan, Chengguang, et al.
Published: (2025)
by: Gan, Chengguang, et al.
Published: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
by: Ma, David, et al.
Published: (2025)
by: Ma, David, et al.
Published: (2025)
I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
by: Liang, Yiming, et al.
Published: (2024)
by: Liang, Yiming, et al.
Published: (2024)
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
by: Xie, Yuan, et al.
Published: (2025)
by: Xie, Yuan, et al.
Published: (2025)
Pre-training, Fine-tuning and Re-ranking: A Three-Stage Framework for Legal Question Answering
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
by: Que, Haoran, et al.
Published: (2024)
by: Que, Haoran, et al.
Published: (2024)
KG-HTC: Integrating Knowledge Graphs into LLMs for Effective Zero-shot Hierarchical Text Classification
by: Zang, Qianbo, et al.
Published: (2025)
by: Zang, Qianbo, et al.
Published: (2025)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
by: Chen, Qiguang, et al.
Published: (2026)
by: Chen, Qiguang, et al.
Published: (2026)
MVAN: Multi-View Attention Networks for Fake News Detection on Social Media
by: Ni, Shiwen, et al.
Published: (2025)
by: Ni, Shiwen, et al.
Published: (2025)
Quantification of Large Language Model Distillation
by: Lee, Sunbowen, et al.
Published: (2025)
by: Lee, Sunbowen, et al.
Published: (2025)
LiPUP-MA: A Residential Experience-centric Multi-Agent Framework for Living-in-the-loop Participatory Urban Planning
by: Ni, Hang, et al.
Published: (2024)
by: Ni, Hang, et al.
Published: (2024)
SparseMap: Loop Mapping for Sparse CNNs on Streaming Coarse-grained Reconfigurable Array
by: Ni, Xiaobing, et al.
Published: (2024)
by: Ni, Xiaobing, et al.
Published: (2024)
How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality
by: Tu, Minzhu, et al.
Published: (2026)
by: Tu, Minzhu, et al.
Published: (2026)
PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment
by: Wang, Zekun Moore, et al.
Published: (2024)
by: Wang, Zekun Moore, et al.
Published: (2024)
Long-form RewardBench: Evaluating Reward Models for Long-form Generation
by: Huang, Hui, et al.
Published: (2026)
by: Huang, Hui, et al.
Published: (2026)
A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
by: Ge, Suyu, et al.
Published: (2024)
by: Ge, Suyu, et al.
Published: (2024)
Context as a Tool: Context Management for Long-Horizon SWE-Agents
by: Liu, Shukai, et al.
Published: (2025)
by: Liu, Shukai, et al.
Published: (2025)
XL$^2$Bench: A Benchmark for Extremely Long Context Understanding with Long-range Dependencies
by: Ni, Xuanfan, et al.
Published: (2024)
by: Ni, Xuanfan, et al.
Published: (2024)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
by: Lin, Kuanwei, et al.
Published: (2026)
by: Lin, Kuanwei, et al.
Published: (2026)
A Comprehensive Survey on Long Context Language Modeling
by: Liu, Jiaheng, et al.
Published: (2025)
by: Liu, Jiaheng, et al.
Published: (2025)
Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization
by: Chen, Qianben, et al.
Published: (2026)
by: Chen, Qianben, et al.
Published: (2026)
Long-Video Audio Synthesis with Multi-Agent Collaboration
by: Zhang, Yehang, et al.
Published: (2025)
by: Zhang, Yehang, et al.
Published: (2025)
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
by: Hu, Xavier, et al.
Published: (2026)
by: Hu, Xavier, et al.
Published: (2026)
O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents
by: Wang, Piaohong, et al.
Published: (2025)
by: Wang, Piaohong, et al.
Published: (2025)
Earley-Driven Dynamic Pruning for Efficient Structured Decoding
by: Sun, Xintong, et al.
Published: (2025)
by: Sun, Xintong, et al.
Published: (2025)
Hierarchical Memory for Long Video QA
by: Wang, Yiqin, et al.
Published: (2024)
by: Wang, Yiqin, et al.
Published: (2024)
RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding
by: Chen, Guanzheng, et al.
Published: (2025)
by: Chen, Guanzheng, et al.
Published: (2025)
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
by: He, Yancheng, et al.
Published: (2025)
by: He, Yancheng, et al.
Published: (2025)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
by: Zhang, Beichen, et al.
Published: (2024)
by: Zhang, Beichen, et al.
Published: (2024)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
by: Wang, Zexuan, et al.
Published: (2026)
by: Wang, Zexuan, et al.
Published: (2026)
Similar Items
-
DMoERM: Recipes of Mixture-of-Experts for Effective Reward Modeling
by: Quan, Shanghaoran
Published: (2024) -
Automatically Generating Numerous Context-Driven SFT Data for LLMs across Diverse Granularity
by: Quan, Shanghaoran
Published: (2024) -
AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions
by: Li, Ziming, et al.
Published: (2024) -
LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm
by: Wu, Siwei, et al.
Published: (2025) -
LongIns: A Challenging Long-context Instruction-based Exam for LLMs
by: Gavin, Shawn, et al.
Published: (2024)