One Diffusion to Generate Them All
Fuente:
arXiv
Saved in:
| Main Authors: | Le, Duong H., Pham, Tuan, Lee, Sangho, Clark, Christopher, Kembhavi, Aniruddha, Mandt, Stephan, Krishna, Ranjay, Lu, Jiasen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preserving Identity with Variational Score for General-purpose 3D Editing
by: Le, Duong H., et al.
Published: (2024)
by: Le, Duong H., et al.
Published: (2024)
Iterated Learning Improves Compositionality in Large Vision-Language Models
by: Zheng, Chenhao, et al.
Published: (2024)
by: Zheng, Chenhao, et al.
Published: (2024)
Neural NeRF Compression
by: Pham, Tuan, et al.
Published: (2024)
by: Pham, Tuan, et al.
Published: (2024)
GeoDiff: Geometry-Guided Diffusion for Metric Depth Estimation
by: Pham, Tuan, et al.
Published: (2025)
by: Pham, Tuan, et al.
Published: (2025)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
UMAMI: Unifying Masked Autoregressive Models and Deterministic Rendering for View Synthesis
by: Le, Thanh-Tung, et al.
Published: (2025)
by: Le, Thanh-Tung, et al.
Published: (2025)
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
by: Kim, Chris Dongjoo, et al.
Published: (2025)
by: Kim, Chris Dongjoo, et al.
Published: (2025)
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
by: Duggal, Shivam, et al.
Published: (2025)
by: Duggal, Shivam, et al.
Published: (2025)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
by: Eftekhar, Ainaz, et al.
Published: (2023)
by: Eftekhar, Ainaz, et al.
Published: (2023)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
by: Maharana, Adyasha, et al.
Published: (2023)
by: Maharana, Adyasha, et al.
Published: (2023)
One Noise to Rule Them All: Multi-View Adversarial Attacks with Universal Perturbation
by: Ergezer, Mehmet, et al.
Published: (2024)
by: Ergezer, Mehmet, et al.
Published: (2024)
ATATA: One Algorithm to Align Them All
by: Pang, Boyi, et al.
Published: (2026)
by: Pang, Boyi, et al.
Published: (2026)
The One RING: a Robotic Indoor Navigation Generalist
by: Eftekhar, Ainaz, et al.
Published: (2024)
by: Eftekhar, Ainaz, et al.
Published: (2024)
OmniCaptioner: One Captioner to Rule Them All
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
Lossy Image Compression with Conditional Diffusion Models
by: Yang, Ruihan, et al.
Published: (2022)
by: Yang, Ruihan, et al.
Published: (2022)
Progressive Compression with Universally Quantized Diffusion Models
by: Yang, Yibo, et al.
Published: (2024)
by: Yang, Yibo, et al.
Published: (2024)
MegaLoc: One Retrieval to Place Them All
by: Berton, Gabriele, et al.
Published: (2025)
by: Berton, Gabriele, et al.
Published: (2025)
SupeRANSAC: One RANSAC to Rule Them All
by: Barath, Daniel
Published: (2025)
by: Barath, Daniel
Published: (2025)
Task Me Anything
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
A4O: All Trigger for One sample
by: Vu, Duc Anh, et al.
Published: (2025)
by: Vu, Duc Anh, et al.
Published: (2025)
Diffusion-Guided Gaussian Splatting for Large-Scale Unconstrained 3D Reconstruction and Novel View Synthesis
by: Mithun, Niluthpol Chowdhury, et al.
Published: (2025)
by: Mithun, Niluthpol Chowdhury, et al.
Published: (2025)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
by: Yang, Yue, et al.
Published: (2023)
by: Yang, Yue, et al.
Published: (2023)
MIMIC: Masked Image Modeling with Image Correspondences
by: Marathe, Kalyani, et al.
Published: (2023)
by: Marathe, Kalyani, et al.
Published: (2023)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
by: Fan, Xiang, et al.
Published: (2024)
by: Fan, Xiang, et al.
Published: (2024)
One Graph to Track Them All: Dynamic GNNs for Single- and Multi-View Tracking
by: Engilberge, Martin, et al.
Published: (2025)
by: Engilberge, Martin, et al.
Published: (2025)
One Channel to Rule Them All: Rethinking Input Representation for Visual Place Recognition
by: Ismagilov, Timur, et al.
Published: (2026)
by: Ismagilov, Timur, et al.
Published: (2026)
One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
by: Bianchi, Lorenzo, et al.
Published: (2025)
by: Bianchi, Lorenzo, et al.
Published: (2025)
Seeing the Unseen: Visual Common Sense for Semantic Placement
by: Ramrakhya, Ram, et al.
Published: (2024)
by: Ramrakhya, Ram, et al.
Published: (2024)
One-Dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing Applications
by: Lyu, Mengyao, et al.
Published: (2023)
by: Lyu, Mengyao, et al.
Published: (2023)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection
by: Guo, Jia, et al.
Published: (2025)
by: Guo, Jia, et al.
Published: (2025)
SwiftPie: Lightning-fast Subject-driven Image Personalization via One step Diffusion
by: Duong, Huy, et al.
Published: (2026)
by: Duong, Huy, et al.
Published: (2026)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
ALTER: All-in-One Layer Pruning and Temporal Expert Routing for Efficient Diffusion Generation
by: Yang, Xiaomeng, et al.
Published: (2025)
by: Yang, Xiaomeng, et al.
Published: (2025)
Fast Samplers for Inverse Problems in Iterative Refinement Models
by: Pandey, Kushagra, et al.
Published: (2024)
by: Pandey, Kushagra, et al.
Published: (2024)
A Hyperdimensional One Place Signature to Represent Them All: Stackable Descriptors For Visual Place Recognition
by: Malone, Connor, et al.
Published: (2024)
by: Malone, Connor, et al.
Published: (2024)
One RL to See Them All: Visual Triple Unified Reinforcement Learning
by: Ma, Yan, et al.
Published: (2025)
by: Ma, Yan, et al.
Published: (2025)
Similar Items
-
Preserving Identity with Variational Score for General-purpose 3D Editing
by: Le, Duong H., et al.
Published: (2024) -
Iterated Learning Improves Compositionality in Large Vision-Language Models
by: Zheng, Chenhao, et al.
Published: (2024) -
Neural NeRF Compression
by: Pham, Tuan, et al.
Published: (2024) -
GeoDiff: Geometry-Guided Diffusion for Metric Depth Estimation
by: Pham, Tuan, et al.
Published: (2025) -
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)