What You See Is What Matters: A Novel Visual and Physics-Based Metric for Evaluating Video Generation Quality
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Zihan, Li, Songlin, Hao, Lingyan, Hu, Xinyu, Song, Bowen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
por: Ni, Ziqi, et al.
Publicado: (2025)
por: Ni, Ziqi, et al.
Publicado: (2025)
What You See is What You Ask: Evaluating Audio Descriptions
por: Kala, Divy, et al.
Publicado: (2025)
por: Kala, Divy, et al.
Publicado: (2025)
FrameOracle: Learning What to See and How Much to See in Videos
por: Li, Chaoyu, et al.
Publicado: (2025)
por: Li, Chaoyu, et al.
Publicado: (2025)
Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
por: Pei, Gensheng, et al.
Publicado: (2025)
por: Pei, Gensheng, et al.
Publicado: (2025)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
por: Choi, Yura, et al.
Publicado: (2026)
por: Choi, Yura, et al.
Publicado: (2026)
What You See is What You Classify: Black Box Attributions
por: Stalder, Steven, et al.
Publicado: (2022)
por: Stalder, Steven, et al.
Publicado: (2022)
What You See is (Usually) What You Get: Multimodal Prototype Networks that Abstain from Expensive Modalities
por: Bahng, Muchang, et al.
Publicado: (2025)
por: Bahng, Muchang, et al.
Publicado: (2025)
Seeing What You Say: Expressive Image Generation from Speech
por: Lee, Jiyoung, et al.
Publicado: (2025)
por: Lee, Jiyoung, et al.
Publicado: (2025)
Tell What You Hear From What You See -- Video to Audio Generation Through Text
por: Liu, Xiulong, et al.
Publicado: (2024)
por: Liu, Xiulong, et al.
Publicado: (2024)
Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation
por: Corvi, Riccardo, et al.
Publicado: (2025)
por: Corvi, Riccardo, et al.
Publicado: (2025)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
por: Kang, Seil, et al.
Publicado: (2025)
por: Kang, Seil, et al.
Publicado: (2025)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
por: Song, Junha, et al.
Publicado: (2026)
por: Song, Junha, et al.
Publicado: (2026)
Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection
por: Yi, Ariana, et al.
Publicado: (2025)
por: Yi, Ariana, et al.
Publicado: (2025)
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
por: Bora, Maheswar, et al.
Publicado: (2025)
por: Bora, Maheswar, et al.
Publicado: (2025)
What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze Estimation
por: Cheng, Yihua, et al.
Publicado: (2024)
por: Cheng, Yihua, et al.
Publicado: (2024)
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
por: De Nadai, Marco, et al.
Publicado: (2025)
por: De Nadai, Marco, et al.
Publicado: (2025)
Focus On What Matters: Separated Models For Visual-Based RL Generalization
por: Zhang, Di, et al.
Publicado: (2024)
por: Zhang, Di, et al.
Publicado: (2024)
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
por: Lin, Weifeng, et al.
Publicado: (2024)
por: Lin, Weifeng, et al.
Publicado: (2024)
Generating 360° Video is What You Need For a 3D Scene
por: Zhang, Zhaoyang, et al.
Publicado: (2025)
por: Zhang, Zhaoyang, et al.
Publicado: (2025)
Visually Dehallucinative Instruction Generation: Know What You Don't Know
por: Cha, Sungguk, et al.
Publicado: (2024)
por: Cha, Sungguk, et al.
Publicado: (2024)
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark
por: Chen, Seng Nam, et al.
Publicado: (2026)
por: Chen, Seng Nam, et al.
Publicado: (2026)
What Matters to You? Towards Visual Representation Alignment for Robot Learning
por: Tian, Ran, et al.
Publicado: (2023)
por: Tian, Ran, et al.
Publicado: (2023)
Smart Feature is What You Need
por: Hu, Zhaoxin, et al.
Publicado: (2024)
por: Hu, Zhaoxin, et al.
Publicado: (2024)
Point What You Mean: Visually Grounded Instruction Policy
por: Yu, Hang, et al.
Publicado: (2025)
por: Yu, Hang, et al.
Publicado: (2025)
What Are You Doing? A Closer Look at Controllable Human Video Generation
por: Bugliarello, Emanuele, et al.
Publicado: (2025)
por: Bugliarello, Emanuele, et al.
Publicado: (2025)
What You Have is What You Track: Adaptive and Robust Multimodal Tracking
por: Tan, Yuedong, et al.
Publicado: (2025)
por: Tan, Yuedong, et al.
Publicado: (2025)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
por: Li, Senmao, et al.
Publicado: (2024)
por: Li, Senmao, et al.
Publicado: (2024)
From Sora What We Can See: A Survey of Text-to-Video Generation
por: Sun, Rui, et al.
Publicado: (2024)
por: Sun, Rui, et al.
Publicado: (2024)
See What You Seek: Semantic Contextual Integration for Cloth-Changing Person Re-Identification
por: Han, Xiyu, et al.
Publicado: (2024)
por: Han, Xiyu, et al.
Publicado: (2024)
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
por: Cho, Beomsik, et al.
Publicado: (2025)
por: Cho, Beomsik, et al.
Publicado: (2025)
What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
por: Lin, Jianghang, et al.
Publicado: (2025)
por: Lin, Jianghang, et al.
Publicado: (2025)
Can Current AI Models Count What We Mean, Not What They See? A Benchmark and Systematic Evaluation
por: Nguyen, Gia Khanh, et al.
Publicado: (2025)
por: Nguyen, Gia Khanh, et al.
Publicado: (2025)
What Do You See? Enhancing Zero-Shot Image Classification with Multimodal Large Language Models
por: Abdelhamed, Abdelrahman, et al.
Publicado: (2024)
por: Abdelhamed, Abdelrahman, et al.
Publicado: (2024)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
por: Sun, Boyuan, et al.
Publicado: (2026)
por: Sun, Boyuan, et al.
Publicado: (2026)
NowYouSee Me: Context-Aware Automatic Audio Description
por: Lee, Seon-Ho, et al.
Publicado: (2024)
por: Lee, Seon-Ho, et al.
Publicado: (2024)
What You See is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANs
por: Trevithick, Alex, et al.
Publicado: (2024)
por: Trevithick, Alex, et al.
Publicado: (2024)
What Do You See in Common? Learning Hierarchical Prototypes over Tree-of-Life to Discover Evolutionary Traits
por: Manogaran, Harish Babu, et al.
Publicado: (2024)
por: Manogaran, Harish Babu, et al.
Publicado: (2024)
Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models
por: Magid, Salma Abdel, et al.
Publicado: (2024)
por: Magid, Salma Abdel, et al.
Publicado: (2024)
Beyond FVD: Enhanced Evaluation Metrics for Video Generation Quality
por: Luo, Ge Ya, et al.
Publicado: (2024)
por: Luo, Ge Ya, et al.
Publicado: (2024)
Masked Generative Transformer Is What You Need for Image Editing
por: Chow, Wei, et al.
Publicado: (2026)
por: Chow, Wei, et al.
Publicado: (2026)
Ejemplares similares
-
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
por: Ni, Ziqi, et al.
Publicado: (2025) -
What You See is What You Ask: Evaluating Audio Descriptions
por: Kala, Divy, et al.
Publicado: (2025) -
FrameOracle: Learning What to See and How Much to See in Videos
por: Li, Chaoyu, et al.
Publicado: (2025) -
Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
por: Pei, Gensheng, et al.
Publicado: (2025) -
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
por: Choi, Yura, et al.
Publicado: (2026)