FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Jing, Liqiang, Lai, Viet, Yoon, Seunghyun, Bui, Trung, Du, Xinya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AniClipart: Clipart Animation with Text-to-Video Priors
by: Wu, Ronghuan, et al.
Published: (2024)
by: Wu, Ronghuan, et al.
Published: (2024)
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
by: S, Sridhar, et al.
Published: (2025)
by: S, Sridhar, et al.
Published: (2025)
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
by: Gal, Rinon, et al.
Published: (2024)
by: Gal, Rinon, et al.
Published: (2024)
TextToon: Real-Time Text Toonify Head Avatar from Single Video
by: Song, Luchuan, et al.
Published: (2024)
by: Song, Luchuan, et al.
Published: (2024)
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
by: Wang, Yiping, et al.
Published: (2024)
by: Wang, Yiping, et al.
Published: (2024)
Bridging Text and Video Generation: A Survey
by: Kumar, Nilay, et al.
Published: (2025)
by: Kumar, Nilay, et al.
Published: (2025)
CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
by: Jing, Liqiang, et al.
Published: (2024)
by: Jing, Liqiang, et al.
Published: (2024)
Text-to-Vector Generation with Neural Path Representation
by: Zhang, Peiying, et al.
Published: (2024)
by: Zhang, Peiying, et al.
Published: (2024)
MVMR: A New Framework for Evaluating Faithfulness of Video Moment Retrieval against Multiple Distractors
by: Yang, Nakyeong, et al.
Published: (2023)
by: Yang, Nakyeong, et al.
Published: (2023)
Text2NeRF: Text-Driven 3D Scene Generation with Neural Radiance Fields
by: Zhang, Jingbo, et al.
Published: (2023)
by: Zhang, Jingbo, et al.
Published: (2023)
Enhancing Sketch Animation: Text-to-Video Diffusion Models with Temporal Consistency and Rigidity Constraints
by: Rai, Gaurav, et al.
Published: (2024)
by: Rai, Gaurav, et al.
Published: (2024)
Style Customization of Text-to-Vector Generation with Image Diffusion Priors
by: Zhang, Peiying, et al.
Published: (2025)
by: Zhang, Peiying, et al.
Published: (2025)
Adaptive Hybrid Caching for Efficient Text-to-Video Diffusion Model Acceleration
by: Wei, Yuanxin, et al.
Published: (2025)
by: Wei, Yuanxin, et al.
Published: (2025)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
CAP: Evaluation of Persuasive and Creative Image Generation
by: Aghazadeh, Aysan, et al.
Published: (2024)
by: Aghazadeh, Aysan, et al.
Published: (2024)
SketchVideo: Sketch-based Video Generation and Editing
by: Liu, Feng-Lin, et al.
Published: (2025)
by: Liu, Feng-Lin, et al.
Published: (2025)
PALP: Prompt Aligned Personalization of Text-to-Image Models
by: Arar, Moab, et al.
Published: (2024)
by: Arar, Moab, et al.
Published: (2024)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention Guidance
by: Sala, Nathan, et al.
Published: (2026)
by: Sala, Nathan, et al.
Published: (2026)
Generative Video Bi-flow
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance
by: Zhou, Xilong, et al.
Published: (2026)
by: Zhou, Xilong, et al.
Published: (2026)
A Text-to-3D Framework for Joint Generation of CG-Ready Humans and Compatible Garments
by: Sun, Zhiyao, et al.
Published: (2025)
by: Sun, Zhiyao, et al.
Published: (2025)
VideoNeuMat: Neural Material Extraction from Generative Video Models
by: Xue, Bowen, et al.
Published: (2026)
by: Xue, Bowen, et al.
Published: (2026)
Text2CAD: Generating Sequential CAD Models from Beginner-to-Expert Level Text Prompts
by: Khan, Mohammad Sadil, et al.
Published: (2024)
by: Khan, Mohammad Sadil, et al.
Published: (2024)
Video4DGen: Enhancing Video and 4D Generation through Mutual Optimization
by: Wang, Yikai, et al.
Published: (2025)
by: Wang, Yikai, et al.
Published: (2025)
Controllable Video Generation: A Survey
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023)
by: Girdhar, Rohit, et al.
Published: (2023)
Generating Human Interaction Motions in Scenes with Text Control
by: Yi, Hongwei, et al.
Published: (2024)
by: Yi, Hongwei, et al.
Published: (2024)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025)
by: Zang, Yuan, et al.
Published: (2025)
A Survey On Text-to-3D Contents Generation In The Wild
by: Jiang, Chenhan
Published: (2024)
by: Jiang, Chenhan
Published: (2024)
Real-time 3D-aware Portrait Video Relighting
by: Cai, Ziqi, et al.
Published: (2024)
by: Cai, Ziqi, et al.
Published: (2024)
From Mannequin to Human: A Pose-Aware and Identity-Preserving Video Generation Framework for Lifelike Clothing Display
by: Mu, Xiangyu, et al.
Published: (2025)
by: Mu, Xiangyu, et al.
Published: (2025)
Portrait Video Editing Empowered by Multimodal Generative Priors
by: Gao, Xuan, et al.
Published: (2024)
by: Gao, Xuan, et al.
Published: (2024)
FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
by: Jing, Liqiang, et al.
Published: (2023)
by: Jing, Liqiang, et al.
Published: (2023)
DragVideo: Interactive Drag-style Video Editing
by: Deng, Yufan, et al.
Published: (2023)
by: Deng, Yufan, et al.
Published: (2023)
Expressive Text-to-Image Generation with Rich Text
by: Ge, Songwei, et al.
Published: (2023)
by: Ge, Songwei, et al.
Published: (2023)
NIVeL: Neural Implicit Vector Layers for Text-to-Vector Generation
by: Thamizharasan, Vikas, et al.
Published: (2024)
by: Thamizharasan, Vikas, et al.
Published: (2024)
DressCode: Autoregressively Sewing and Generating Garments from Text Guidance
by: He, Kai, et al.
Published: (2024)
by: He, Kai, et al.
Published: (2024)
WordRobe: Text-Guided Generation of Textured 3D Garments
by: Srivastava, Astitva, et al.
Published: (2024)
by: Srivastava, Astitva, et al.
Published: (2024)
Similar Items
-
AniClipart: Clipart Animation with Text-to-Video Priors
by: Wu, Ronghuan, et al.
Published: (2024) -
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
by: S, Sridhar, et al.
Published: (2025) -
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
by: Gal, Rinon, et al.
Published: (2024) -
TextToon: Real-Time Text Toonify Head Avatar from Single Video
by: Song, Luchuan, et al.
Published: (2024) -
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
by: Wang, Yiping, et al.
Published: (2024)