OSCBench: Benchmarking Object State Change in Text-to-Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Xianjing, Zhu, Bin, Hu, Shiqi, Li, Franklin Mingzhe, Carrington, Patrick, Zimmermann, Roger, Chen, Jingjing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OSCAR: Object Status and Contextual Awareness for Recipes to Support Non-Visual Cooking
von: Li, Franklin Mingzhe, et al.
Veröffentlicht: (2025)
von: Li, Franklin Mingzhe, et al.
Veröffentlicht: (2025)
Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking
von: Li, Franklin Mingzhe, et al.
Veröffentlicht: (2025)
von: Li, Franklin Mingzhe, et al.
Veröffentlicht: (2025)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
von: Li, Huiqiong, et al.
Veröffentlicht: (2026)
von: Li, Huiqiong, et al.
Veröffentlicht: (2026)
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models
von: Cheng, Ming, et al.
Veröffentlicht: (2025)
von: Cheng, Ming, et al.
Veröffentlicht: (2025)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
von: Huang, Haoyang, et al.
Veröffentlicht: (2025)
von: Huang, Haoyang, et al.
Veröffentlicht: (2025)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
American Sign Language Video to Text Translation
von: Roy, Parsheeta, et al.
Veröffentlicht: (2024)
von: Roy, Parsheeta, et al.
Veröffentlicht: (2024)
More than One Step at a Time: Designing Procedural Feedback for Non-visual Makeup Routines
von: Li, Franklin Mingzhe, et al.
Veröffentlicht: (2025)
von: Li, Franklin Mingzhe, et al.
Veröffentlicht: (2025)
AutoArabic: A Three-Stage Framework for Localizing Video-Text Retrieval Benchmarks
von: Eltahir, Mohamed, et al.
Veröffentlicht: (2025)
von: Eltahir, Mohamed, et al.
Veröffentlicht: (2025)
Bilingual Text-to-Motion Generation: A New Benchmark and Baselines
von: Weng, Wanjiang, et al.
Veröffentlicht: (2026)
von: Weng, Wanjiang, et al.
Veröffentlicht: (2026)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
von: Gao, Xin, et al.
Veröffentlicht: (2026)
von: Gao, Xin, et al.
Veröffentlicht: (2026)
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
von: Zhou, Ziwei, et al.
Veröffentlicht: (2026)
von: Zhou, Ziwei, et al.
Veröffentlicht: (2026)
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
Integrating Video and Text: A Balanced Approach to Multimodal Summary Generation and Evaluation
von: Pennec, Galann, et al.
Veröffentlicht: (2025)
von: Pennec, Galann, et al.
Veröffentlicht: (2025)
T$^3$Bench: Benchmarking Current Progress in Text-to-3D Generation
von: He, Yuze, et al.
Veröffentlicht: (2023)
von: He, Yuze, et al.
Veröffentlicht: (2023)
MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models
von: Ruan, Chenxi, et al.
Veröffentlicht: (2026)
von: Ruan, Chenxi, et al.
Veröffentlicht: (2026)
LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text
von: yunhan, Li, et al.
Veröffentlicht: (2025)
von: yunhan, Li, et al.
Veröffentlicht: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
von: Wu, Yin, et al.
Veröffentlicht: (2025)
von: Wu, Yin, et al.
Veröffentlicht: (2025)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
von: Du, Yang, et al.
Veröffentlicht: (2025)
von: Du, Yang, et al.
Veröffentlicht: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
OSCaR: Object State Captioning and State Change Representation
von: Nguyen, Nguyen, et al.
Veröffentlicht: (2024)
von: Nguyen, Nguyen, et al.
Veröffentlicht: (2024)
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
von: Li, Xuchen, et al.
Veröffentlicht: (2024)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
von: Li, Jialu, et al.
Veröffentlicht: (2025)
von: Li, Jialu, et al.
Veröffentlicht: (2025)
Temporal Reasoning Transfer from Text to Video
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models
von: Luo, Sha, et al.
Veröffentlicht: (2026)
von: Luo, Sha, et al.
Veröffentlicht: (2026)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2025)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
von: Zhu, Kejian, et al.
Veröffentlicht: (2025)
von: Zhu, Kejian, et al.
Veröffentlicht: (2025)
T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
Actions and Objects Pathways for Domain Adaptation in Video Question Answering
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2024)
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2024)
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning
von: Zhao, Yusong, et al.
Veröffentlicht: (2026)
von: Zhao, Yusong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OSCAR: Object Status and Contextual Awareness for Recipes to Support Non-Visual Cooking
von: Li, Franklin Mingzhe, et al.
Veröffentlicht: (2025) -
Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking
von: Li, Franklin Mingzhe, et al.
Veröffentlicht: (2025) -
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
von: Li, Huiqiong, et al.
Veröffentlicht: (2026) -
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024) -
VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models
von: Cheng, Ming, et al.
Veröffentlicht: (2025)