Exploring Iterative Refinement with Diffusion Models for Video Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Xiao, Shi, Tao, Liang, Yaoyuan, Tao, Te, Huang, Shao-Lun |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
by: Han, Tingting, et al.
Published: (2026)
by: Han, Tingting, et al.
Published: (2026)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
by: Liu, Qihao, et al.
Published: (2025)
by: Liu, Qihao, et al.
Published: (2025)
Fine-grained Spatiotemporal Grounding on Egocentric Videos
by: Liang, Shuo, et al.
Published: (2025)
by: Liang, Shuo, et al.
Published: (2025)
PSDiff: Diffusion Model for Person Search with Iterative and Collaborative Refinement
by: Jia, Chengyou, et al.
Published: (2023)
by: Jia, Chengyou, et al.
Published: (2023)
On Exploring PDE Modeling for Point Cloud Video Representation Learning
by: Huang, Zhuoxu, et al.
Published: (2024)
by: Huang, Zhuoxu, et al.
Published: (2024)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
by: Wu, Jianzong, et al.
Published: (2025)
by: Wu, Jianzong, et al.
Published: (2025)
LATTE: Improving Latex Recognition for Tables and Formulae with Iterative Refinement
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation
by: Zhang, Haojie, et al.
Published: (2024)
by: Zhang, Haojie, et al.
Published: (2024)
NaRCan: Natural Refined Canonical Image with Integration of Diffusion Prior for Video Editing
by: Chen, Ting-Hsuan, et al.
Published: (2024)
by: Chen, Ting-Hsuan, et al.
Published: (2024)
CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning
by: Guo, Tengda, et al.
Published: (2026)
by: Guo, Tengda, et al.
Published: (2026)
Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion
by: Chen, Jingyuan, et al.
Published: (2025)
by: Chen, Jingyuan, et al.
Published: (2025)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
by: Liang, Cheng, et al.
Published: (2026)
by: Liang, Cheng, et al.
Published: (2026)
HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting
by: Lu, Wenquan, et al.
Published: (2023)
by: Lu, Wenquan, et al.
Published: (2023)
Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
by: Cheng, Junhao, et al.
Published: (2025)
by: Cheng, Junhao, et al.
Published: (2025)
SDM-Car: A Dataset for Small and Dim Moving Vehicles Detection in Satellite Videos
by: Zhang, Zhen, et al.
Published: (2024)
by: Zhang, Zhen, et al.
Published: (2024)
Improving Visual Reasoning with Iterative Evidence Refinement
by: Shi, Zeru, et al.
Published: (2026)
by: Shi, Zeru, et al.
Published: (2026)
DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label Recognition
by: Liu, Haijing, et al.
Published: (2025)
by: Liu, Haijing, et al.
Published: (2025)
PrecisionCUA: Iterative Visual Refinement for Pixel-Precise Cursor Grounding in Code Editors
by: Mittal, Himangi, et al.
Published: (2026)
by: Mittal, Himangi, et al.
Published: (2026)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
by: Xue, Qiyao, et al.
Published: (2024)
by: Xue, Qiyao, et al.
Published: (2024)
Pseudo-labeling with Keyword Refining for Few-Supervised Video Captioning
by: Li, Ping, et al.
Published: (2024)
by: Li, Ping, et al.
Published: (2024)
Exploring Data-Free LoRA Transferability for Video Diffusion Models
by: Wang, Yuchen, et al.
Published: (2026)
by: Wang, Yuchen, et al.
Published: (2026)
Category-Adaptive Cross-Modal Semantic Refinement and Transfer for Open-Vocabulary Multi-Label Recognition
by: Liu, Haijing, et al.
Published: (2024)
by: Liu, Haijing, et al.
Published: (2024)
ReflexFlow: Rethinking Learning Objective for Exposure Bias Alleviation in Flow Matching
by: Huang, Guanbo, et al.
Published: (2025)
by: Huang, Guanbo, et al.
Published: (2025)
GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
by: Dou, Huanzhang, et al.
Published: (2024)
by: Dou, Huanzhang, et al.
Published: (2024)
Iterative Diffusion-Refined Neural Attenuation Fields for Multi-Source Stationary CT Reconstruction: NAF Meets Diffusion Model
by: Fang, Jiancheng, et al.
Published: (2025)
by: Fang, Jiancheng, et al.
Published: (2025)
AICL: Action In-Context Learning for Video Diffusion Model
by: Liu, Jianzhi, et al.
Published: (2024)
by: Liu, Jianzhi, et al.
Published: (2024)
IRNet: Iterative Refinement Network for Noisy Partial Label Learning
by: Lian, Zheng, et al.
Published: (2022)
by: Lian, Zheng, et al.
Published: (2022)
Representation Space Constrained Learning with Modality Decoupling for Multimodal Object Detection
by: Shao, YiKang, et al.
Published: (2025)
by: Shao, YiKang, et al.
Published: (2025)
OrbitNVS: Harnessing Video Diffusion Priors for Novel View Synthesis
by: Liang, Jinglin, et al.
Published: (2026)
by: Liang, Jinglin, et al.
Published: (2026)
HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
by: Li, Peiyan, et al.
Published: (2026)
by: Li, Peiyan, et al.
Published: (2026)
AutoRefiner: Improving Autoregressive Video Diffusion Models via Reflective Refinement Over the Stochastic Sampling Path
by: Yu, Zhengyang, et al.
Published: (2025)
by: Yu, Zhengyang, et al.
Published: (2025)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
TimeRefine: Temporal Grounding with Time Refining Video LLM
by: Wang, Xizi, et al.
Published: (2024)
by: Wang, Xizi, et al.
Published: (2024)
End-to-End Dense Video Grounding via Parallel Regression
by: Shi, Fengyuan, et al.
Published: (2021)
by: Shi, Fengyuan, et al.
Published: (2021)
DreamJourney: Perpetual View Generation with Video Diffusion Models
by: Pan, Bo, et al.
Published: (2025)
by: Pan, Bo, et al.
Published: (2025)
SDMatte: Grafting Diffusion Models for Interactive Matting
by: Huang, Longfei, et al.
Published: (2025)
by: Huang, Longfei, et al.
Published: (2025)
Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models
by: Liang, Hanwen, et al.
Published: (2024)
by: Liang, Hanwen, et al.
Published: (2024)
E2VIDiff: Perceptual Events-to-Video Reconstruction using Diffusion Priors
by: Liang, Jinxiu, et al.
Published: (2024)
by: Liang, Jinxiu, et al.
Published: (2024)
Similar Items
-
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
by: Liu, Yang, et al.
Published: (2025) -
HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
by: Han, Tingting, et al.
Published: (2026) -
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
by: Liu, Qihao, et al.
Published: (2025) -
Fine-grained Spatiotemporal Grounding on Egocentric Videos
by: Liang, Shuo, et al.
Published: (2025) -
PSDiff: Diffusion Model for Person Search with Iterative and Collaborative Refinement
by: Jia, Chengyou, et al.
Published: (2023)