Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yang, Zhao, Xilin, Wen, Peisong, Dai, Siran, Huang, Qingming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Exploring Structural Degradation in Dense Representations for Self-supervised Learning
by: Dai, Siran, et al.
Published: (2025)
by: Dai, Siran, et al.
Published: (2025)
Self-supervised Representation Learning with Local Aggregation for Image-based Profiling
by: Dai, Siran, et al.
Published: (2025)
by: Dai, Siran, et al.
Published: (2025)
Semantic Concentration for Self-Supervised Dense Representations Learning
by: Wen, Peisong, et al.
Published: (2025)
by: Wen, Peisong, et al.
Published: (2025)
HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models
by: Lu, Zhiguang, et al.
Published: (2025)
by: Lu, Zhiguang, et al.
Published: (2025)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
by: Xue, Qiyao, et al.
Published: (2024)
by: Xue, Qiyao, et al.
Published: (2024)
Exploring Iterative Refinement with Diffusion Models for Video Grounding
by: Liang, Xiao, et al.
Published: (2023)
by: Liang, Xiao, et al.
Published: (2023)
Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness
by: Wan, Hanwen, et al.
Published: (2025)
by: Wan, Hanwen, et al.
Published: (2025)
ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
by: Pan, Panwang, et al.
Published: (2025)
by: Pan, Panwang, et al.
Published: (2025)
VLM-Guided Iterative Refinement for Surgical Image Segmentation with Foundation Models
by: Lou, Ange, et al.
Published: (2026)
by: Lou, Ange, et al.
Published: (2026)
Regularized Contrastive Partial Multi-view Outlier Detection
by: Wang, Yijia, et al.
Published: (2024)
by: Wang, Yijia, et al.
Published: (2024)
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
by: Wang, Zun, et al.
Published: (2024)
by: Wang, Zun, et al.
Published: (2024)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
by: Zheng, Zelin, et al.
Published: (2026)
by: Zheng, Zelin, et al.
Published: (2026)
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance
by: Wang, Zhao, et al.
Published: (2025)
by: Wang, Zhao, et al.
Published: (2025)
PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Top-K Pairwise Ranking: Bridging the Gap Among Ranking-Based Measures for Multi-Label Classification
by: Wang, Zitai, et al.
Published: (2024)
by: Wang, Zitai, et al.
Published: (2024)
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual Representation
by: Han, Boyu, et al.
Published: (2026)
by: Han, Boyu, et al.
Published: (2026)
Enhancing Image Matting in Real-World Scenes with Mask-Guided Iterative Refinement
by: Liu, Rui
Published: (2025)
by: Liu, Rui
Published: (2025)
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
NEWTON: Agentic Planning for Physically Grounded Video Generation
by: Feng, Yuxiang, et al.
Published: (2026)
by: Feng, Yuxiang, et al.
Published: (2026)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
by: Ma, Yue, et al.
Published: (2023)
by: Ma, Yue, et al.
Published: (2023)
Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in Video
by: Qi, Zhaobo, et al.
Published: (2024)
by: Qi, Zhaobo, et al.
Published: (2024)
Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
by: Yang, Zhengyuan, et al.
Published: (2023)
by: Yang, Zhengyuan, et al.
Published: (2023)
Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation
by: Hassan, Mariam, et al.
Published: (2026)
by: Hassan, Mariam, et al.
Published: (2026)
Concepts Worth Having: Refining VLM-Guided Concept Bottleneck Models with Minimal Annotations
by: Debole, Nicola, et al.
Published: (2026)
by: Debole, Nicola, et al.
Published: (2026)
Context-Guided Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
AUCSeg: AUC-oriented Pixel-level Long-tail Semantic Segmentation
by: Han, Boyu, et al.
Published: (2024)
by: Han, Boyu, et al.
Published: (2024)
Learning Procedural-aware Video Representations through State-Grounded Hierarchy Unfolding
by: Zhao, Jinghan, et al.
Published: (2025)
by: Zhao, Jinghan, et al.
Published: (2025)
Event-based Video Frame Interpolation with Edge Guided Motion Refinement
by: Liu, Yuhan, et al.
Published: (2024)
by: Liu, Yuhan, et al.
Published: (2024)
Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
by: Ma, Yunchuan, et al.
Published: (2026)
by: Ma, Yunchuan, et al.
Published: (2026)
Physics-Guided VLM Priors for All-Cloud Removal
by: Xu, Liying, et al.
Published: (2026)
by: Xu, Liying, et al.
Published: (2026)
Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment
by: Song, Yizhi, et al.
Published: (2024)
by: Song, Yizhi, et al.
Published: (2024)
HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
by: Han, Tingting, et al.
Published: (2026)
by: Han, Tingting, et al.
Published: (2026)
PrecisionCUA: Iterative Visual Refinement for Pixel-Precise Cursor Grounding in Code Editors
by: Mittal, Himangi, et al.
Published: (2026)
by: Mittal, Himangi, et al.
Published: (2026)
SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents
by: Huang-Menders, Alexander, et al.
Published: (2025)
by: Huang-Menders, Alexander, et al.
Published: (2025)
DiffQRCoder: Diffusion-based Aesthetic QR Code Generation with Scanning Robustness Guided Iterative Refinement
by: Liao, Jia-Wei, et al.
Published: (2024)
by: Liao, Jia-Wei, et al.
Published: (2024)
Learn Faster and Remember More: Balancing Exploration and Exploitation for Continual Test-time Adaptation
by: Yang, Pinci, et al.
Published: (2025)
by: Yang, Pinci, et al.
Published: (2025)
Similar Items
-
From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning
by: Liu, Yang, et al.
Published: (2026) -
When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning
by: Liu, Yang, et al.
Published: (2025) -
Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval
by: Liu, Yang, et al.
Published: (2024) -
Exploring Structural Degradation in Dense Representations for Self-supervised Learning
by: Dai, Siran, et al.
Published: (2025) -
Self-supervised Representation Learning with Local Aggregation for Image-based Profiling
by: Dai, Siran, et al.
Published: (2025)