Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chatterjee, Agneet, Entezari, Rahim, Zhuravinskyi, Maksym, Lapin, Maksim, Adithyan, Reshinth, Raj, Amit, Baral, Chitta, Yang, Yezhou, Jampani, Varun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
by: Bandyopadhyay, Hmrishav, et al.
Published: (2025)
by: Bandyopadhyay, Hmrishav, et al.
Published: (2025)
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
by: Chatterjee, Agneet, et al.
Published: (2024)
by: Chatterjee, Agneet, et al.
Published: (2024)
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
by: Malaviya, Vatsal, et al.
Published: (2025)
by: Malaviya, Vatsal, et al.
Published: (2025)
Chimera: Compositional Image Generation using Part-based Concepting
by: Singh, Shivam, et al.
Published: (2025)
by: Singh, Shivam, et al.
Published: (2025)
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
by: Chatterjee, Agneet, et al.
Published: (2024)
by: Chatterjee, Agneet, et al.
Published: (2024)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
by: Fallah, Forouzan, et al.
Published: (2025)
by: Fallah, Forouzan, et al.
Published: (2025)
Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
Dual Caption Preference Optimization for Diffusion Models
by: Saeidi, Amir, et al.
Published: (2025)
by: Saeidi, Amir, et al.
Published: (2025)
Stable Video-Driven Portraits
by: R., Mallikarjun B., et al.
Published: (2025)
by: R., Mallikarjun B., et al.
Published: (2025)
Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic
by: Alyafeai, Zaid, et al.
Published: (2024)
by: Alyafeai, Zaid, et al.
Published: (2024)
Stable Code Technical Report
by: Pinnaparaju, Nikhil, et al.
Published: (2024)
by: Pinnaparaju, Nikhil, et al.
Published: (2024)
Block Cascading: Training Free Acceleration of Block-Causal Video Models
by: Bandyopadhyay, Hmrishav, et al.
Published: (2025)
by: Bandyopadhyay, Hmrishav, et al.
Published: (2025)
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation
by: Varshney, Neeraj, et al.
Published: (2024)
by: Varshney, Neeraj, et al.
Published: (2024)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
by: Patel, Maitreya, et al.
Published: (2023)
by: Patel, Maitreya, et al.
Published: (2023)
ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training
by: Pieler, Michael, et al.
Published: (2024)
by: Pieler, Michael, et al.
Published: (2024)
HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
by: Vani, Sameep, et al.
Published: (2025)
by: Vani, Sameep, et al.
Published: (2025)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
WordRobe: Text-Guided Generation of Textured 3D Garments
by: Srivastava, Astitva, et al.
Published: (2024)
by: Srivastava, Astitva, et al.
Published: (2024)
Getting it Right: Improving Spatial Consistency in Text-to-Image Models
by: Chatterjee, Agneet, et al.
Published: (2024)
by: Chatterjee, Agneet, et al.
Published: (2024)
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
by: Pathiraja, Bimsara, et al.
Published: (2025)
by: Pathiraja, Bimsara, et al.
Published: (2025)
Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images
by: Luo, Yiran, et al.
Published: (2024)
by: Luo, Yiran, et al.
Published: (2024)
Stable LM 2 1.6B Technical Report
by: Bellagente, Marco, et al.
Published: (2024)
by: Bellagente, Marco, et al.
Published: (2024)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
by: Yilmaz, Nilay, et al.
Published: (2025)
by: Yilmaz, Nilay, et al.
Published: (2025)
Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
by: Kusumba, Abhiram, et al.
Published: (2025)
by: Kusumba, Abhiram, et al.
Published: (2025)
Rethinking Training for De-biasing Text-to-Image Generation: Unlocking the Potential of Stable Diffusion
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
by: Saxon, Michael, et al.
Published: (2024)
by: Saxon, Michael, et al.
Published: (2024)
DiffusionLight-Turbo: Accelerated Light Probes for Free via Single-Pass Chrome Ball Inpainting
by: Chinchuthakun, Worameth, et al.
Published: (2025)
by: Chinchuthakun, Worameth, et al.
Published: (2025)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
LightHeadEd: Relightable & Editable Head Avatars from a Smartphone
by: Manu, Pranav, et al.
Published: (2025)
by: Manu, Pranav, et al.
Published: (2025)
SViM3D: Stable Video Material Diffusion for Single Image 3D Generation
by: Engelhardt, Andreas, et al.
Published: (2025)
by: Engelhardt, Andreas, et al.
Published: (2025)
SF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglement
by: Boss, Mark, et al.
Published: (2024)
by: Boss, Mark, et al.
Published: (2024)
VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
Dress-Me-Up: A Dataset & Method for Self-Supervised 3D Garment Retargeting
by: Naik, Shanthika, et al.
Published: (2024)
by: Naik, Shanthika, et al.
Published: (2024)
ICE-G: Image Conditional Editing of 3D Gaussian Splats
by: Jaganathan, Vishnu, et al.
Published: (2024)
by: Jaganathan, Vishnu, et al.
Published: (2024)
DiffusionLight: Light Probes for Free by Painting a Chrome Ball
by: Phongthawee, Pakkapon, et al.
Published: (2023)
by: Phongthawee, Pakkapon, et al.
Published: (2023)
The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
by: Varshney, Neeraj, et al.
Published: (2023)
by: Varshney, Neeraj, et al.
Published: (2023)
Similar Items
-
SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
by: Bandyopadhyay, Hmrishav, et al.
Published: (2025) -
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
by: Chatterjee, Agneet, et al.
Published: (2024) -
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
by: Malaviya, Vatsal, et al.
Published: (2025) -
Chimera: Compositional Image Generation using Part-based Concepting
by: Singh, Shivam, et al.
Published: (2025) -
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
by: Chatterjee, Agneet, et al.
Published: (2024)