VCBench: A Streaming Counting Benchmark for Spatial-Temporal State Maintenance in Long Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Pengyiang, Shi, Zhongyue, Hao, Hongye, Fu, Qi, Bi, Xueting, Zhang, Siwei, Hu, Xiaoyang, Wang, Zitian, Huang, Linjiang, Liu, Si |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VCBench: Benchmarking LLMs in Venture Capital
by: Chen, Rick, et al.
Published: (2025)
by: Chen, Rick, et al.
Published: (2025)
BiCoord: A Bimanual Manipulation Benchmark towards Long-Horizon Spatial-Temporal Coordination
by: Peng, Xingyu, et al.
Published: (2026)
by: Peng, Xingyu, et al.
Published: (2026)
StreamSTGS: Streaming Spatial and Temporal Gaussian Grids for Real-Time Free-Viewpoint Video
by: Ke, Zhihui, et al.
Published: (2025)
by: Ke, Zhihui, et al.
Published: (2025)
Audio-Sync Video Generation with Multi-Stream Temporal Control
by: Weng, Shuchen, et al.
Published: (2025)
by: Weng, Shuchen, et al.
Published: (2025)
Spacewalk-18: A Benchmark for Multimodal and Long-form Procedural Video Understanding in Novel Domains
by: Tang, Zitian, et al.
Published: (2023)
by: Tang, Zitian, et al.
Published: (2023)
Crowded Video Individual Counting Informed by Social Grouping and Spatial-Temporal Displacement Priors
by: Lu, Hao, et al.
Published: (2026)
by: Lu, Hao, et al.
Published: (2026)
LongStream: Long-Sequence Streaming Autoregressive Visual Geometry
by: Cheng, Chong, et al.
Published: (2026)
by: Cheng, Chong, et al.
Published: (2026)
Multilayer Collaborative Optimization for the System Configuration, Operation, and Maintenance of Smart Community Microgrids
by: Jiangshan Liu, et al.
Published: (2025)
by: Jiangshan Liu, et al.
Published: (2025)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
by: Cheng, Zixu, et al.
Published: (2025)
by: Cheng, Zixu, et al.
Published: (2025)
MAMBA4D: Efficient Long-Sequence Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models
by: Liu, Jiuming, et al.
Published: (2024)
by: Liu, Jiuming, et al.
Published: (2024)
Spatial-Conditioned Reasoning in Long-Egocentric Videos
by: Tribble, James, et al.
Published: (2026)
by: Tribble, James, et al.
Published: (2026)
EC-Bench: Enumeration and Counting Benchmark for Ultra-Long Videos
by: Tsuchiya, Fumihiko, et al.
Published: (2026)
by: Tsuchiya, Fumihiko, et al.
Published: (2026)
Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
by: Gao, Yuanyuan, et al.
Published: (2026)
by: Gao, Yuanyuan, et al.
Published: (2026)
Spatial Visibility and Temporal Dynamics: Revolutionizing Field of View Prediction in Adaptive Point Cloud Video Streaming
by: Li, Chen, et al.
Published: (2024)
by: Li, Chen, et al.
Published: (2024)
Deep Temporal Graph Clustering: A Comprehensive Benchmark and Datasets
by: Liu, Meng, et al.
Published: (2026)
by: Liu, Meng, et al.
Published: (2026)
Dual migration modes of unfaulted disconnections on curved twin boundaries
by: He, Hongrui, et al.
Published: (2026)
by: He, Hongrui, et al.
Published: (2026)
Efficient Approximate Temporal Triangle Counting in Streaming with Predictions
by: Venturin, Giorgio, et al.
Published: (2025)
by: Venturin, Giorgio, et al.
Published: (2025)
Kinematics-Aware Latent World Models for Data-Efficient Autonomous Driving
by: Li, Jiazhuo, et al.
Published: (2026)
by: Li, Jiazhuo, et al.
Published: (2026)
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
by: Gao, Jianxiong, et al.
Published: (2025)
by: Gao, Jianxiong, et al.
Published: (2025)
SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
by: Yang, Zhenyu, et al.
Published: (2025)
by: Yang, Zhenyu, et al.
Published: (2025)
L-STEC: Learned Video Compression with Long-term Spatio-Temporal Enhanced Context
by: Zhang, Tiange, et al.
Published: (2025)
by: Zhang, Tiange, et al.
Published: (2025)
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
by: Yu, Yifei, et al.
Published: (2025)
by: Yu, Yifei, et al.
Published: (2025)
EndoStreamDepth: Temporally Consistent Monocular Depth Estimation for Endoscopic Video Streams
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video
by: Qu, Jinyuan, et al.
Published: (2024)
by: Qu, Jinyuan, et al.
Published: (2024)
Sparse Reduced-rank Regression Methods for Spatially Misaligned Data with Application to Spatial Transcriptomics
by: Wu, Zitian, et al.
Published: (2026)
by: Wu, Zitian, et al.
Published: (2026)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
by: Fu, Honghao, et al.
Published: (2026)
by: Fu, Honghao, et al.
Published: (2026)
See, Remember, Explore: A Benchmark and Baselines for Streaming Spatial Reasoning
by: Wei, Yuxi, et al.
Published: (2026)
by: Wei, Yuxi, et al.
Published: (2026)
Described Spatial-Temporal Video Detection
by: Ji, Wei, et al.
Published: (2024)
by: Ji, Wei, et al.
Published: (2024)
Generative Neural Video Compression via Video Diffusion Prior
by: Mao, Qi, et al.
Published: (2025)
by: Mao, Qi, et al.
Published: (2025)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
FlexDrive: Toward Trajectory Flexibility in Driving Scene Reconstruction and Rendering
by: Zhou, Jingqiu, et al.
Published: (2025)
by: Zhou, Jingqiu, et al.
Published: (2025)
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
by: Cheng, Chong, et al.
Published: (2026)
by: Cheng, Chong, et al.
Published: (2026)
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
by: Huang, Haojian, et al.
Published: (2025)
by: Huang, Haojian, et al.
Published: (2025)
TFMQ-DM: Temporal Feature Maintenance Quantization for Diffusion Models
by: Huang, Yushi, et al.
Published: (2023)
by: Huang, Yushi, et al.
Published: (2023)
PDStream: Slashing Long-Tail Delay in Interactive Video Streaming via Pseudo-Dual Streaming
by: Xiao, Xuedou, et al.
Published: (2025)
by: Xiao, Xuedou, et al.
Published: (2025)
LaDy: Lagrangian-Dynamic Informed Network for Skeleton-based Action Segmentation via Spatial-Temporal Modulation
by: Ji, Haoyu, et al.
Published: (2026)
by: Ji, Haoyu, et al.
Published: (2026)
Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
by: Shen, Cuifeng, et al.
Published: (2025)
by: Shen, Cuifeng, et al.
Published: (2025)
Counting Fish with Temporal Representations of Sonar Video
by: Van Brunt, Kai, et al.
Published: (2025)
by: Van Brunt, Kai, et al.
Published: (2025)
MDSAM:Memory-Driven Sparse Attention Matrix for LVLMs Hallucination Mitigation
by: Lu, Shuaiye, et al.
Published: (2025)
by: Lu, Shuaiye, et al.
Published: (2025)
Promptus: Can Prompts Streaming Replace Video Streaming with Stable Diffusion
by: Wu, Jiangkai, et al.
Published: (2024)
by: Wu, Jiangkai, et al.
Published: (2024)
Similar Items
-
VCBench: Benchmarking LLMs in Venture Capital
by: Chen, Rick, et al.
Published: (2025) -
BiCoord: A Bimanual Manipulation Benchmark towards Long-Horizon Spatial-Temporal Coordination
by: Peng, Xingyu, et al.
Published: (2026) -
StreamSTGS: Streaming Spatial and Temporal Gaussian Grids for Real-Time Free-Viewpoint Video
by: Ke, Zhihui, et al.
Published: (2025) -
Audio-Sync Video Generation with Multi-Stream Temporal Control
by: Weng, Shuchen, et al.
Published: (2025) -
Spacewalk-18: A Benchmark for Multimodal and Long-form Procedural Video Understanding in Novel Domains
by: Tang, Zitian, et al.
Published: (2023)