SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Matsuda, Ryosuke, Kudo, Keito, Yoshida, Haruto, Shimizu, Nobuyuki, Suzuki, Jun |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Video Text Preservation with Synthetic Text-Rich Videos
par: Liu, Ziyang, et autres
Publié: (2025)
par: Liu, Ziyang, et autres
Publié: (2025)
LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation
par: Zheng, Xiangqing, et autres
Publié: (2025)
par: Zheng, Xiangqing, et autres
Publié: (2025)
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
par: Li, Jie, et autres
Publié: (2025)
par: Li, Jie, et autres
Publié: (2025)
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
par: Oshima, Yuta, et autres
Publié: (2024)
par: Oshima, Yuta, et autres
Publié: (2024)
ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models
par: Rawte, Vipula, et autres
Publié: (2024)
par: Rawte, Vipula, et autres
Publié: (2024)
SeqBench: Benchmarking Sequential Narrative Generation in Text-to-Video Models
par: Tang, Zhengxu, et autres
Publié: (2025)
par: Tang, Zhengxu, et autres
Publié: (2025)
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
par: Tu, Rong-Cheng, et autres
Publié: (2024)
par: Tu, Rong-Cheng, et autres
Publié: (2024)
AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
par: Li, Jieyu, et autres
Publié: (2025)
par: Li, Jieyu, et autres
Publié: (2025)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
par: Feng, Weixi, et autres
Publié: (2024)
par: Feng, Weixi, et autres
Publié: (2024)
Image-Conditioned 3D Gaussian Splat Quantization
par: Liu, Xinshuang, et autres
Publié: (2025)
par: Liu, Xinshuang, et autres
Publié: (2025)
Nodes Are Early, Edges Are Late: Probing Diagram Representations in Large Vision-Language Models
par: Yoshida, Haruto, et autres
Publié: (2026)
par: Yoshida, Haruto, et autres
Publié: (2026)
LVBench: An Extreme Long Video Understanding Benchmark
par: Wang, Weihan, et autres
Publié: (2024)
par: Wang, Weihan, et autres
Publié: (2024)
OSCBench: Benchmarking Object State Change in Text-to-Video Generation
par: Han, Xianjing, et autres
Publié: (2026)
par: Han, Xianjing, et autres
Publié: (2026)
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
par: Zhou, Ziwei, et autres
Publié: (2026)
par: Zhou, Ziwei, et autres
Publié: (2026)
DAOVI: Distortion-Aware Omnidirectional Video Inpainting
par: Seshimo, Ryosuke, et autres
Publié: (2025)
par: Seshimo, Ryosuke, et autres
Publié: (2025)
VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation
par: Jiang, Longteng, et autres
Publié: (2026)
par: Jiang, Longteng, et autres
Publié: (2026)
Video-Infinity: Distributed Long Video Generation
par: Tan, Zhenxiong, et autres
Publié: (2024)
par: Tan, Zhenxiong, et autres
Publié: (2024)
Audio-centric Video Understanding Benchmark without Text Shortcut
par: Yang, Yudong, et autres
Publié: (2025)
par: Yang, Yudong, et autres
Publié: (2025)
Video-Bench: Human-Aligned Video Generation Benchmark
par: Han, Hui, et autres
Publié: (2025)
par: Han, Hui, et autres
Publié: (2025)
DiffuSyn Bench: Evaluating Vision-Language Models on Real-World Complexities with Diffusion-Generated Synthetic Benchmarks
par: Zhou, Haokun, et autres
Publié: (2024)
par: Zhou, Haokun, et autres
Publié: (2024)
VideoLLM Benchmarks and Evaluation: A Survey
par: Kumar, Yogesh
Publié: (2025)
par: Kumar, Yogesh
Publié: (2025)
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding
par: Zou, Heqing, et autres
Publié: (2025)
par: Zou, Heqing, et autres
Publié: (2025)
Video-Based Performance Evaluation for ECR Drills in Synthetic Training Environments
par: Rayala, Surya, et autres
Publié: (2025)
par: Rayala, Surya, et autres
Publié: (2025)
Synthetic Human Action Video Data Generation with Pose Transfer
par: Knapp, Vaclav, et autres
Publié: (2025)
par: Knapp, Vaclav, et autres
Publié: (2025)
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
par: Lu, Hao, et autres
Publié: (2025)
par: Lu, Hao, et autres
Publié: (2025)
Guidance Matters: Rethinking the Evaluation Pitfall for Text-to-Image Generation
par: Xie, Dian, et autres
Publié: (2026)
par: Xie, Dian, et autres
Publié: (2026)
Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation
par: Tang, Zhijiang, et autres
Publié: (2026)
par: Tang, Zhijiang, et autres
Publié: (2026)
STREAM: Spatio-TempoRal Evaluation and Analysis Metric for Video Generative Models
par: Kim, Pum Jun, et autres
Publié: (2024)
par: Kim, Pum Jun, et autres
Publié: (2024)
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
par: Wang, Ruotong, et autres
Publié: (2025)
par: Wang, Ruotong, et autres
Publié: (2025)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
par: Matsuda, Kazuki, et autres
Publié: (2025)
par: Matsuda, Kazuki, et autres
Publié: (2025)
Efficient Zero-Shot AI-Generated Image Detection
par: Sonoda, Ryosuke, et autres
Publié: (2026)
par: Sonoda, Ryosuke, et autres
Publié: (2026)
Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval
par: Cho, CH, et autres
Publié: (2025)
par: Cho, CH, et autres
Publié: (2025)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
par: Yuan, Huaying, et autres
Publié: (2025)
par: Yuan, Huaying, et autres
Publié: (2025)
MLVU: Benchmarking Multi-task Long Video Understanding
par: Zhou, Junjie, et autres
Publié: (2024)
par: Zhou, Junjie, et autres
Publié: (2024)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
par: Nagrani, Arsha, et autres
Publié: (2024)
par: Nagrani, Arsha, et autres
Publié: (2024)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
par: Feng, Weixi, et autres
Publié: (2025)
par: Feng, Weixi, et autres
Publié: (2025)
HARIVO: Harnessing Text-to-Image Models for Video Generation
par: Kwon, Mingi, et autres
Publié: (2024)
par: Kwon, Mingi, et autres
Publié: (2024)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
par: Kamath, Amita, et autres
Publié: (2025)
par: Kamath, Amita, et autres
Publié: (2025)
Script-to-Slide Grounding: Grounding Script Sentences to Slide Objects for Automatic Instructional Video Generation
par: Suzuki, Rena, et autres
Publié: (2026)
par: Suzuki, Rena, et autres
Publié: (2026)
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
par: Hu, Lanxiang, et autres
Publié: (2025)
par: Hu, Lanxiang, et autres
Publié: (2025)
Documents similaires
-
Video Text Preservation with Synthetic Text-Rich Videos
par: Liu, Ziyang, et autres
Publié: (2025) -
LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation
par: Zheng, Xiangqing, et autres
Publié: (2025) -
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
par: Li, Jie, et autres
Publié: (2025) -
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
par: Oshima, Yuta, et autres
Publié: (2024) -
ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models
par: Rawte, Vipula, et autres
Publié: (2024)