ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge
Fuente:
arXiv
Guardado en:
| Autores principales: | Abdelrahman, Eslam, Zhao, Liangbing, Hu, Vincent Tao, Cord, Matthieu, Perez, Patrick, Elhoseiny, Mohamed |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
por: Abdelrahman, Eslam, et al.
Publicado: (2023)
From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
por: Zhao, Liangbing, et al.
Publicado: (2026)
por: Zhao, Liangbing, et al.
Publicado: (2026)
iMotion-LLM: Instruction-Conditioned Trajectory Generation
por: Felemban, Abdulwahab, et al.
Publicado: (2024)
por: Felemban, Abdulwahab, et al.
Publicado: (2024)
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
por: Zhuo, Le, et al.
Publicado: (2025)
por: Zhuo, Le, et al.
Publicado: (2025)
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
por: Ataallah, Kirolos, et al.
Publicado: (2024)
por: Ataallah, Kirolos, et al.
Publicado: (2024)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
por: Ataallah, Kirolos, et al.
Publicado: (2024)
por: Ataallah, Kirolos, et al.
Publicado: (2024)
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
por: Vallaeys, Théophane, et al.
Publicado: (2025)
por: Vallaeys, Théophane, et al.
Publicado: (2025)
Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation
por: Yang, Zhongyu, et al.
Publicado: (2025)
por: Yang, Zhongyu, et al.
Publicado: (2025)
Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
por: Ahmed, Mahmoud, et al.
Publicado: (2024)
por: Ahmed, Mahmoud, et al.
Publicado: (2024)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
por: Ataallah, Kirolos, et al.
Publicado: (2024)
por: Ataallah, Kirolos, et al.
Publicado: (2024)
Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation
por: Eldesokey, Abdelrahman, et al.
Publicado: (2024)
por: Eldesokey, Abdelrahman, et al.
Publicado: (2024)
A Dynamic Programming Framework for Discovering Count and Values of Multilevel Image Thresholding
por: Hegazy, Eslam, et al.
Publicado: (2026)
por: Hegazy, Eslam, et al.
Publicado: (2026)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
Halton Scheduler For Masked Generative Image Transformer
por: Besnier, Victor, et al.
Publicado: (2025)
por: Besnier, Victor, et al.
Publicado: (2025)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
por: Shukor, Mustafa, et al.
Publicado: (2024)
por: Shukor, Mustafa, et al.
Publicado: (2024)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
por: Shen, Xiaoqian, et al.
Publicado: (2023)
por: Shen, Xiaoqian, et al.
Publicado: (2023)
INSTA-YOLO: Real-Time Instance Segmentation
por: Mohamed, Eslam, et al.
Publicado: (2021)
por: Mohamed, Eslam, et al.
Publicado: (2021)
FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
Skipping Computations in Multimodal LLMs
por: Shukor, Mustafa, et al.
Publicado: (2024)
por: Shukor, Mustafa, et al.
Publicado: (2024)
AsyncDSB: Schedule-Asynchronous Diffusion Schrödinger Bridge for Image Inpainting
por: Han, Zihao, et al.
Publicado: (2024)
por: Han, Zihao, et al.
Publicado: (2024)
Overcoming Generic Knowledge Loss with Selective Parameter Update
por: Zhang, Wenxuan, et al.
Publicado: (2023)
por: Zhang, Wenxuan, et al.
Publicado: (2023)
Efficient Self-supervised Vision Pretraining with Local Masked Reconstruction
por: Chen, Jun, et al.
Publicado: (2022)
por: Chen, Jun, et al.
Publicado: (2022)
PointBeV: A Sparse Approach to BeV Predictions
por: Chambon, Loick, et al.
Publicado: (2023)
por: Chambon, Loick, et al.
Publicado: (2023)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
por: Shen, Xiaoqian, et al.
Publicado: (2025)
por: Shen, Xiaoqian, et al.
Publicado: (2025)
Simplified Diffusion Schrödinger Bridge
por: Tang, Zhicong, et al.
Publicado: (2024)
por: Tang, Zhicong, et al.
Publicado: (2024)
Ambiguous Medical Image Segmentation Using Diffusion Schrödinger Bridge
por: Baru, Lalith Bharadwaj, et al.
Publicado: (2025)
por: Baru, Lalith Bharadwaj, et al.
Publicado: (2025)
LatentMan: Generating Consistent Animated Characters using Image Diffusion Models
por: Eldesokey, Abdelrahman, et al.
Publicado: (2023)
por: Eldesokey, Abdelrahman, et al.
Publicado: (2023)
Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting
por: Messaoud, Kaouther, et al.
Publicado: (2025)
por: Messaoud, Kaouther, et al.
Publicado: (2025)
IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
por: Wang, Hanting, et al.
Publicado: (2025)
por: Wang, Hanting, et al.
Publicado: (2025)
Reliability in Semantic Segmentation: Can We Use Synthetic Data?
por: Loiseau, Thibaut, et al.
Publicado: (2023)
por: Loiseau, Thibaut, et al.
Publicado: (2023)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
por: Mohamed, Abdelrahman, et al.
Publicado: (2025)
por: Mohamed, Abdelrahman, et al.
Publicado: (2025)
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
por: Caselles-Dupré, Hugo, et al.
Publicado: (2026)
por: Caselles-Dupré, Hugo, et al.
Publicado: (2026)
DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut
por: Couairon, Paul, et al.
Publicado: (2024)
por: Couairon, Paul, et al.
Publicado: (2024)
Neural Catalog: Scaling Species Recognition with Catalog of Life-Augmented Generation
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
BridgeShape: Latent Diffusion Schrödinger Bridge for 3D Shape Completion
por: Kong, Dequan, et al.
Publicado: (2025)
por: Kong, Dequan, et al.
Publicado: (2025)
Efficient Generative Modeling beyond Memoryless Diffusion via Adjoint Schrödinger Bridge Matching
por: Shin, Jeongwoo, et al.
Publicado: (2026)
por: Shin, Jeongwoo, et al.
Publicado: (2026)
ToddlerAct: A Toddler Action Recognition Dataset for Gross Motor Development Assessment
por: Huang, Hsiang-Wei, et al.
Publicado: (2024)
por: Huang, Hsiang-Wei, et al.
Publicado: (2024)
FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion Models
por: Corradini, Barbara Toniella, et al.
Publicado: (2024)
por: Corradini, Barbara Toniella, et al.
Publicado: (2024)
Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?
por: Xu, Yihong, et al.
Publicado: (2023)
por: Xu, Yihong, et al.
Publicado: (2023)
Ejemplares similares
-
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
por: Abdelrahman, Eslam, et al.
Publicado: (2023) -
From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
por: Zhao, Liangbing, et al.
Publicado: (2026) -
iMotion-LLM: Instruction-Conditioned Trajectory Generation
por: Felemban, Abdulwahab, et al.
Publicado: (2024) -
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
por: Zhuo, Le, et al.
Publicado: (2025) -
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
por: Ataallah, Kirolos, et al.
Publicado: (2024)