PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Qiyao, Yin, Xiangyu, Yang, Boyuan, Gao, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis Users
von: Yin, Xiangyu, et al.
Veröffentlicht: (2025)
von: Yin, Xiangyu, et al.
Veröffentlicht: (2025)
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
von: Wu, Shang, et al.
Veröffentlicht: (2026)
von: Wu, Shang, et al.
Veröffentlicht: (2026)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
von: Wang, Haoming, et al.
Veröffentlicht: (2026)
von: Wang, Haoming, et al.
Veröffentlicht: (2026)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2025)
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2025)
PhiP-G: Physics-Guided Text-to-3D Compositional Scene Generation
von: Li, Qixuan, et al.
Veröffentlicht: (2025)
von: Li, Qixuan, et al.
Veröffentlicht: (2025)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
von: Huang, Yidong, et al.
Veröffentlicht: (2026)
von: Huang, Yidong, et al.
Veröffentlicht: (2026)
VideoPhy: Evaluating Physical Commonsense for Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
PhyDrawGen: Physically Grounded Diagram Generation from Natural Language
von: Haque, Nafiul, et al.
Veröffentlicht: (2026)
von: Haque, Nafiul, et al.
Veröffentlicht: (2026)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
von: Gu, Jing, et al.
Veröffentlicht: (2025)
von: Gu, Jing, et al.
Veröffentlicht: (2025)
PhyGround: Benchmarking Physical Reasoning in Generative World Models
von: Lin, Juyi, et al.
Veröffentlicht: (2026)
von: Lin, Juyi, et al.
Veröffentlicht: (2026)
TimeRefine: Temporal Grounding with Time Refining Video LLM
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
PhyGile: Physics-Prefix Guided Motion Generation for Agile General Humanoid Motion Tracking
von: Bao, Jiacheng, et al.
Veröffentlicht: (2026)
von: Bao, Jiacheng, et al.
Veröffentlicht: (2026)
Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Yang, Xiaomeng, et al.
Veröffentlicht: (2025)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
von: Jeong, Suchae, et al.
Veröffentlicht: (2025)
von: Jeong, Suchae, et al.
Veröffentlicht: (2025)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
von: Lee, Daeun, et al.
Veröffentlicht: (2024)
Exploring Iterative Refinement with Diffusion Models for Video Grounding
von: Liang, Xiao, et al.
Veröffentlicht: (2023)
von: Liang, Xiao, et al.
Veröffentlicht: (2023)
Physics-Guided Motion Loss for Video Generation Model
von: Xue, Bowen, et al.
Veröffentlicht: (2025)
von: Xue, Bowen, et al.
Veröffentlicht: (2025)
Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement
von: Guo, Junrong, et al.
Veröffentlicht: (2026)
von: Guo, Junrong, et al.
Veröffentlicht: (2026)
FAIRT2V: Training-Free Debiasing for Text-to-Video Diffusion Models
von: Zhong, Haonan, et al.
Veröffentlicht: (2026)
von: Zhong, Haonan, et al.
Veröffentlicht: (2026)
PhyWorld: Physics-Faithful World Model for Video Generation
von: Zhao, Pu, et al.
Veröffentlicht: (2026)
von: Zhao, Pu, et al.
Veröffentlicht: (2026)
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
von: Xu, Chengzhi, et al.
Veröffentlicht: (2025)
von: Xu, Chengzhi, et al.
Veröffentlicht: (2025)
PhyTracker: An Online Tracker for Phytoplankton
von: Yu, Yang, et al.
Veröffentlicht: (2024)
von: Yu, Yang, et al.
Veröffentlicht: (2024)
VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
von: Lei, Bin, et al.
Veröffentlicht: (2025)
von: Lei, Bin, et al.
Veröffentlicht: (2025)
Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
FAGER: Factually Grounded Evaluation and Refinement of Text-to-Image Models
von: Lim, Youngsun, et al.
Veröffentlicht: (2026)
von: Lim, Youngsun, et al.
Veröffentlicht: (2026)
PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation
von: Wu, Fan, et al.
Veröffentlicht: (2025)
von: Wu, Fan, et al.
Veröffentlicht: (2025)
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
von: Mao, Jiawei, et al.
Veröffentlicht: (2026)
von: Mao, Jiawei, et al.
Veröffentlicht: (2026)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
von: Cai, Yuanhao, et al.
Veröffentlicht: (2025)
von: Cai, Yuanhao, et al.
Veröffentlicht: (2025)
Motion-o: Trajectory-Grounded Video Reasoning
von: Galoaa, Bishoy, et al.
Veröffentlicht: (2026)
von: Galoaa, Bishoy, et al.
Veröffentlicht: (2026)
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
von: M, Megha Mariam K., et al.
Veröffentlicht: (2026)
Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness
von: Wan, Hanwen, et al.
Veröffentlicht: (2025)
von: Wan, Hanwen, et al.
Veröffentlicht: (2025)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative Refinement
von: Zou, Yude, et al.
Veröffentlicht: (2026)
von: Zou, Yude, et al.
Veröffentlicht: (2026)
AIR: Zero-shot Generative Model Adaptation with Iterative Refinement
von: Liu, Guimeng, et al.
Veröffentlicht: (2025)
von: Liu, Guimeng, et al.
Veröffentlicht: (2025)
Online Iterative Self-Alignment for Radiology Report Generation
von: Xiao, Ting, et al.
Veröffentlicht: (2025)
von: Xiao, Ting, et al.
Veröffentlicht: (2025)
VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
von: Cao, Xinye, et al.
Veröffentlicht: (2025)
von: Cao, Xinye, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis Users
von: Yin, Xiangyu, et al.
Veröffentlicht: (2025) -
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
von: Wu, Shang, et al.
Veröffentlicht: (2026) -
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
von: Liu, Yang, et al.
Veröffentlicht: (2025) -
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
von: Wang, Haoming, et al.
Veröffentlicht: (2026) -
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2025)