PredBench: Benchmarking Spatio-Temporal Prediction across Diverse Disciplines
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, ZiDong, Lu, Zeyu, Huang, Di, He, Tong, Liu, Xihui, Ouyang, Wanli, Bai, Lei |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines
par: Tang, Chen, et autres
Publié: (2025)
par: Tang, Chen, et autres
Publié: (2025)
FiTv2: Scalable and Improved Flexible Vision Transformer for Diffusion Model
par: Wang, ZiDong, et autres
Publié: (2024)
par: Wang, ZiDong, et autres
Publié: (2024)
ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
par: Xue, Xiangyuan, et autres
Publié: (2024)
par: Xue, Xiangyuan, et autres
Publié: (2024)
FiT: Flexible Vision Transformer for Diffusion Model
par: Lu, Zeyu, et autres
Publié: (2024)
par: Lu, Zeyu, et autres
Publié: (2024)
Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction
par: Chen, Junyi, et autres
Publié: (2024)
par: Chen, Junyi, et autres
Publié: (2024)
WorldSimBench: Towards Video Generation Models as World Simulators
par: Qin, Yiran, et autres
Publié: (2024)
par: Qin, Yiran, et autres
Publié: (2024)
ReBA-Pred-Net: Weakly-Supervised Regional Brain Age Prediction on MRI
par: Shao, Shuai, et autres
Publié: (2026)
par: Shao, Shuai, et autres
Publié: (2026)
Improving Weakly Supervised Temporal Action Localization by Exploiting Multi-resolution Information in Temporal Domain
par: Su, Rui, et autres
Publié: (2025)
par: Su, Rui, et autres
Publié: (2025)
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
par: Yue, Xiaoyu, et autres
Publié: (2025)
par: Yue, Xiaoyu, et autres
Publié: (2025)
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
par: Su, Rui, et autres
Publié: (2019)
par: Su, Rui, et autres
Publié: (2019)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
par: Zhang, Mingfang, et autres
Publié: (2026)
par: Zhang, Mingfang, et autres
Publié: (2026)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
par: Lin, Jingli, et autres
Publié: (2025)
par: Lin, Jingli, et autres
Publié: (2025)
PrevPredMap: Exploring Temporal Modeling with Previous Predictions for Online Vectorized HD Map Construction
par: Peng, Nan, et autres
Publié: (2024)
par: Peng, Nan, et autres
Publié: (2024)
NeuRodin: A Two-stage Framework for High-Fidelity Neural Surface Reconstruction
par: Wang, Yifan, et autres
Publié: (2024)
par: Wang, Yifan, et autres
Publié: (2024)
RS-Mamba for Large Remote Sensing Image Dense Prediction
par: Zhao, Sijie, et autres
Publié: (2024)
par: Zhao, Sijie, et autres
Publié: (2024)
Diffusion Models Need Visual Priors for Image Generation
par: Yue, Xiaoyu, et autres
Publié: (2024)
par: Yue, Xiaoyu, et autres
Publié: (2024)
CompBench: Benchmarking Complex Instruction-guided Image Editing
par: Jia, Bohan, et autres
Publié: (2025)
par: Jia, Bohan, et autres
Publié: (2025)
T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
par: Huang, Kaiyi, et autres
Publié: (2023)
par: Huang, Kaiyi, et autres
Publié: (2023)
mmPred: Radar-based Human Motion Prediction in the Dark
par: Fan, Junqiao, et autres
Publié: (2025)
par: Fan, Junqiao, et autres
Publié: (2025)
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
par: Park, Jinho, et autres
Publié: (2026)
par: Park, Jinho, et autres
Publié: (2026)
Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning
par: Zhu, Haoyi, et autres
Publié: (2024)
par: Zhu, Haoyi, et autres
Publié: (2024)
T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
par: Sun, Kaiyue, et autres
Publié: (2025)
par: Sun, Kaiyue, et autres
Publié: (2025)
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
par: Zhang, Sha, et autres
Publié: (2024)
par: Zhang, Sha, et autres
Publié: (2024)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
par: Meng, Jiahao, et autres
Publié: (2026)
par: Meng, Jiahao, et autres
Publié: (2026)
T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
par: Sun, Kaiyue, et autres
Publié: (2024)
par: Sun, Kaiyue, et autres
Publié: (2024)
PredNext: Explicit Cross-View Temporal Prediction for Unsupervised Learning in Spiking Neural Networks
par: Dong, Yiting, et autres
Publié: (2025)
par: Dong, Yiting, et autres
Publié: (2025)
Point Transformer V3: Simpler, Faster, Stronger
par: Wu, Xiaoyang, et autres
Publié: (2023)
par: Wu, Xiaoyang, et autres
Publié: (2023)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
par: Xu, Qi'ao, et autres
Publié: (2025)
par: Xu, Qi'ao, et autres
Publié: (2025)
STAR: A Benchmark for Astronomical Star Fields Super-Resolution
par: Wu, Kuo-Cheng, et autres
Publié: (2025)
par: Wu, Kuo-Cheng, et autres
Publié: (2025)
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
par: Qiu, Lu, et autres
Publié: (2024)
par: Qiu, Lu, et autres
Publié: (2024)
GVGEN: Text-to-3D Generation with Volumetric Representation
par: He, Xianglong, et autres
Publié: (2024)
par: He, Xianglong, et autres
Publié: (2024)
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
par: Shen, Hao, et autres
Publié: (2024)
par: Shen, Hao, et autres
Publié: (2024)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
par: He, Yuping, et autres
Publié: (2025)
par: He, Yuping, et autres
Publié: (2025)
Native-Resolution Image Synthesis
par: Wang, Zidong, et autres
Publié: (2025)
par: Wang, Zidong, et autres
Publié: (2025)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
par: Fang, Rongyao, et autres
Publié: (2025)
par: Fang, Rongyao, et autres
Publié: (2025)
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
par: Aparcedo, Alejandro, et autres
Publié: (2026)
par: Aparcedo, Alejandro, et autres
Publié: (2026)
MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
par: Zhang, Yaqi, et autres
Publié: (2023)
par: Zhang, Yaqi, et autres
Publié: (2023)
TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation
par: Wu, Xiaopei, et autres
Publié: (2024)
par: Wu, Xiaopei, et autres
Publié: (2024)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
par: Chen, Guo, et autres
Publié: (2024)
par: Chen, Guo, et autres
Publié: (2024)
EMR-Merging: Tuning-Free High-Performance Model Merging
par: Huang, Chenyu, et autres
Publié: (2024)
par: Huang, Chenyu, et autres
Publié: (2024)
Documents similaires
-
UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines
par: Tang, Chen, et autres
Publié: (2025) -
FiTv2: Scalable and Improved Flexible Vision Transformer for Diffusion Model
par: Wang, ZiDong, et autres
Publié: (2024) -
ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
par: Xue, Xiangyuan, et autres
Publié: (2024) -
FiT: Flexible Vision Transformer for Diffusion Model
par: Lu, Zeyu, et autres
Publié: (2024) -
Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction
par: Chen, Junyi, et autres
Publié: (2024)