One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
Fuente:
arXiv
Saved in:
| Main Authors: | Esmaeilzadeh, Mehdi, Jolicoeur-Martineau, Alexia, Vashist, Chirag, Li, Ke |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rejection Sampling IMLE: Designing Priors for Better Few-Shot Image Synthesis
by: Vashist, Chirag, et al.
Published: (2024)
by: Vashist, Chirag, et al.
Published: (2024)
Ctrl-V: Higher Fidelity Video Generation with Bounding-Box Controlled Object Motion
by: Luo, Ge Ya, et al.
Published: (2024)
by: Luo, Ge Ya, et al.
Published: (2024)
PopulAtion Parameter Averaging (PAPA)
by: Jolicoeur-Martineau, Alexia, et al.
Published: (2023)
by: Jolicoeur-Martineau, Alexia, et al.
Published: (2023)
Beyond FVD: Enhanced Evaluation Metrics for Video Generation Quality
by: Luo, Ge Ya, et al.
Published: (2024)
by: Luo, Ge Ya, et al.
Published: (2024)
Less is More: Recursive Reasoning with Tiny Networks
by: Jolicoeur-Martineau, Alexia
Published: (2025)
by: Jolicoeur-Martineau, Alexia
Published: (2025)
Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes
by: Gosselin, Anthony, et al.
Published: (2025)
by: Gosselin, Anthony, et al.
Published: (2025)
OMG-Seg: Is One Model Good Enough For All Segmentation?
by: Li, Xiangtai, et al.
Published: (2024)
by: Li, Xiangtai, et al.
Published: (2024)
One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
One Look is Enough: Seamless Patchwise Refinement for Zero-Shot Monocular Depth Estimation on High-Resolution Images
by: Kwon, Byeongjun, et al.
Published: (2025)
by: Kwon, Byeongjun, et al.
Published: (2025)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
by: Guo, Ziyu, et al.
Published: (2026)
by: Guo, Ziyu, et al.
Published: (2026)
One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
by: Rahary, Adrien Ramanana, et al.
Published: (2026)
by: Rahary, Adrien Ramanana, et al.
Published: (2026)
OSV: One Step is Enough for High-Quality Image to Video Generation
by: Mao, Xiaofeng, et al.
Published: (2024)
by: Mao, Xiaofeng, et al.
Published: (2024)
Vision Tiny Recursion Model (ViTRM): Parameter-Efficient Image Classification via Recursive State Refinement
by: Akazan, Ange-Clément, et al.
Published: (2026)
by: Akazan, Ange-Clément, et al.
Published: (2026)
Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization
by: Bharadwaj, Siddhant, et al.
Published: (2026)
by: Bharadwaj, Siddhant, et al.
Published: (2026)
One Shot is Enough for Sequential Infrared Small Target Segmentation
by: Dan, Bingbing, et al.
Published: (2024)
by: Dan, Bingbing, et al.
Published: (2024)
Pool-Select-Refine: Allocation-Aware Generative Dataset Distillation with Soft-Label-Guided Latent Refinement
by: Li, Wenmin, et al.
Published: (2026)
by: Li, Wenmin, et al.
Published: (2026)
Generative Video Compression with One-Dimensional Latent Representation
by: Zheng, Zihan, et al.
Published: (2026)
by: Zheng, Zihan, et al.
Published: (2026)
Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning
by: Zafar, Anas, et al.
Published: (2026)
by: Zafar, Anas, et al.
Published: (2026)
Is One GPU Enough? Pushing Image Generation at Higher-Resolutions with Foundation Models
by: Tragakis, Athanasios, et al.
Published: (2024)
by: Tragakis, Athanasios, et al.
Published: (2024)
One Language-Free Foundation Model Is Enough for Universal Vision Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2026)
by: Gao, Bin-Bin, et al.
Published: (2026)
Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion
by: Kim, Jiwon, et al.
Published: (2025)
by: Kim, Jiwon, et al.
Published: (2025)
Omni-3DEdit: Generalized Versatile 3D Editing in One-Pass
by: Liyi, Chen, et al.
Published: (2026)
by: Liyi, Chen, et al.
Published: (2026)
Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation
by: Hassan, Mariam, et al.
Published: (2026)
by: Hassan, Mariam, et al.
Published: (2026)
One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
Multi-Agent Game Generation and Evaluation via Audio-Visual Recordings
by: Jolicoeur-Martineau, Alexia
Published: (2025)
by: Jolicoeur-Martineau, Alexia
Published: (2025)
COLLAR: Cascaded Object-Level Latent Refinement for High-Fidelity Conditional Generation
by: Zhang, Xinlong, et al.
Published: (2026)
by: Zhang, Xinlong, et al.
Published: (2026)
One-step Latent-free Image Generation with Pixel Mean Flows
by: Lu, Yiyang, et al.
Published: (2026)
by: Lu, Yiyang, et al.
Published: (2026)
Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation
by: Li, Muquan, et al.
Published: (2026)
by: Li, Muquan, et al.
Published: (2026)
From One to More: Contextual Part Latents for 3D Generation
by: Dong, Shaocong, et al.
Published: (2025)
by: Dong, Shaocong, et al.
Published: (2025)
One Small Step in Latent, One Giant Leap for Pixels: Fast Latent Upscale Adapter for Your Diffusion Models
by: Razin, Aleksandr, et al.
Published: (2025)
by: Razin, Aleksandr, et al.
Published: (2025)
Probabilistic Tiny Recursive Model
by: Sghaier, Amin, et al.
Published: (2026)
by: Sghaier, Amin, et al.
Published: (2026)
Text-Driven Diverse Facial Texture Generation via Progressive Latent-Space Refinement
by: Wang, Chi, et al.
Published: (2024)
by: Wang, Chi, et al.
Published: (2024)
MGTraj: Multi-Granularity Goal-Guided Human Trajectory Prediction with Recursive Refinement Network
by: Sun, Ge, et al.
Published: (2025)
by: Sun, Ge, et al.
Published: (2025)
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
by: Sun, Yuwei, et al.
Published: (2026)
by: Sun, Yuwei, et al.
Published: (2026)
Latent Harmony: Synergistic Unified UHD Image Restoration via Latent Space Regularization and Controllable Refinement
by: Liu, Yidi, et al.
Published: (2025)
by: Liu, Yidi, et al.
Published: (2025)
LongDiff: Training-Free Long Video Generation in One Go
by: Li, Zhuoling, et al.
Published: (2025)
by: Li, Zhuoling, et al.
Published: (2025)
InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image Generation
by: Liu, Xingchao, et al.
Published: (2023)
by: Liu, Xingchao, et al.
Published: (2023)
HighSync: High-Quality Lip Synchronization via Latent Diffusion Models
by: Daghigh, Saeed Firouzi, et al.
Published: (2026)
by: Daghigh, Saeed Firouzi, et al.
Published: (2026)
One Model, Many Budgets: Elastic Latent Interfaces for Diffusion Transformers
by: Haji-Ali, Moayed, et al.
Published: (2026)
by: Haji-Ali, Moayed, et al.
Published: (2026)
Diffusion As Self-Distillation: End-to-End Latent Diffusion In One Model
by: Wang, Xiyuan, et al.
Published: (2025)
by: Wang, Xiyuan, et al.
Published: (2025)
Similar Items
-
Rejection Sampling IMLE: Designing Priors for Better Few-Shot Image Synthesis
by: Vashist, Chirag, et al.
Published: (2024) -
Ctrl-V: Higher Fidelity Video Generation with Bounding-Box Controlled Object Motion
by: Luo, Ge Ya, et al.
Published: (2024) -
PopulAtion Parameter Averaging (PAPA)
by: Jolicoeur-Martineau, Alexia, et al.
Published: (2023) -
Beyond FVD: Enhanced Evaluation Metrics for Video Generation Quality
by: Luo, Ge Ya, et al.
Published: (2024) -
Less is More: Recursive Reasoning with Tiny Networks
by: Jolicoeur-Martineau, Alexia
Published: (2025)