JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Byung-Ki, Kwon, Dai, Qi, Hyoseok, Lee, Luo, Chong, Oh, Tae-Hyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zero-shot Depth Completion via Test-time Alignment with Affine-invariant Depth Prior
von: Hyoseok, Lee, et al.
Veröffentlicht: (2025)
von: Hyoseok, Lee, et al.
Veröffentlicht: (2025)
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
von: Lim, Sohwi, et al.
Veröffentlicht: (2026)
von: Lim, Sohwi, et al.
Veröffentlicht: (2026)
Early Failure Detection and Intervention in Video Diffusion Models
von: Byung-Ki, Kwon, et al.
Veröffentlicht: (2026)
von: Byung-Ki, Kwon, et al.
Veröffentlicht: (2026)
Measurement-Consistent Langevin Corrector for Stabilizing Latent Diffusion Inverse Problem Solvers
von: Hyoseok, Lee, et al.
Veröffentlicht: (2026)
von: Hyoseok, Lee, et al.
Veröffentlicht: (2026)
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
Co-synthesis of Histopathology Nuclei Image-Label Pairs using a Context-Conditioned Joint Diffusion Model
von: Min, Seonghui, et al.
Veröffentlicht: (2024)
von: Min, Seonghui, et al.
Veröffentlicht: (2024)
HDR-NSFF: High Dynamic Range Neural Scene Flow Fields
von: Dong-Yeon, Shin, et al.
Veröffentlicht: (2026)
von: Dong-Yeon, Shin, et al.
Veröffentlicht: (2026)
Learning-based Axial Video Motion Magnification
von: Byung-Ki, Kwon, et al.
Veröffentlicht: (2023)
von: Byung-Ki, Kwon, et al.
Veröffentlicht: (2023)
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and TEst-time Generative Adaptation
von: Youwang, Kim, et al.
Veröffentlicht: (2026)
von: Youwang, Kim, et al.
Veröffentlicht: (2026)
The Devil is in the Details: Simple Remedies for Image-to-LiDAR Representation Learning
von: Jo, Wonjun, et al.
Veröffentlicht: (2025)
von: Jo, Wonjun, et al.
Veröffentlicht: (2025)
Semantic Guidance Tuning for Text-To-Image Diffusion Models
von: Kang, Hyun, et al.
Veröffentlicht: (2023)
von: Kang, Hyun, et al.
Veröffentlicht: (2023)
Localized Concept Erasure for Text-to-Image Diffusion Models Using Training-Free Gated Low-Rank Adaptation
von: Lee, Byung Hyun, et al.
Veröffentlicht: (2025)
von: Lee, Byung Hyun, et al.
Veröffentlicht: (2025)
JointSplat: Probabilistic Joint Flow-Depth Optimization for Sparse-View Gaussian Splatting
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
Joint PET-MRI Reconstruction with Diffusion Stochastic Differential Model
von: Xie, Taofeng, et al.
Veröffentlicht: (2024)
von: Xie, Taofeng, et al.
Veröffentlicht: (2024)
ESC: Erasing Space Concept for Knowledge Deletion
von: Lee, Tae-Young, et al.
Veröffentlicht: (2025)
von: Lee, Tae-Young, et al.
Veröffentlicht: (2025)
VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
von: Hyeon-Woo, Nam, et al.
Veröffentlicht: (2024)
von: Hyeon-Woo, Nam, et al.
Veröffentlicht: (2024)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
von: Wang, Kai, et al.
Veröffentlicht: (2024)
von: Wang, Kai, et al.
Veröffentlicht: (2024)
Co-SemDepth: Fast Joint Semantic Segmentation and Depth Estimation on Aerial Images
von: AlaaEldin, Yara, et al.
Veröffentlicht: (2025)
von: AlaaEldin, Yara, et al.
Veröffentlicht: (2025)
Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models
von: Seo, Hoigi, et al.
Veröffentlicht: (2026)
von: Seo, Hoigi, et al.
Veröffentlicht: (2026)
RDFC-GAN: RGB-Depth Fusion CycleGAN for Indoor Depth Completion
von: Wang, Haowen, et al.
Veröffentlicht: (2023)
von: Wang, Haowen, et al.
Veröffentlicht: (2023)
Controllable and Efficient Multi-Class Pathology Nuclei Data Augmentation using Text-Conditioned Diffusion Models
von: Oh, Hyun-Jic, et al.
Veröffentlicht: (2024)
von: Oh, Hyun-Jic, et al.
Veröffentlicht: (2024)
MetaVoxel: Joint Diffusion Modeling of Imaging and Clinical Metadata
von: Liu, Yihao, et al.
Veröffentlicht: (2025)
von: Liu, Yihao, et al.
Veröffentlicht: (2025)
MemBench: Memorized Image Trigger Prompt Dataset for Diffusion Models
von: Hong, Chunsan, et al.
Veröffentlicht: (2024)
von: Hong, Chunsan, et al.
Veröffentlicht: (2024)
Diffusion Models for Joint Audio-Video Generation
von: La Torre, Alejandro Paredes
Veröffentlicht: (2026)
von: La Torre, Alejandro Paredes
Veröffentlicht: (2026)
SiNGER: A Clearer Voice Distills Vision Transformers Further
von: Yu, Geunhyeok, et al.
Veröffentlicht: (2025)
von: Yu, Geunhyeok, et al.
Veröffentlicht: (2025)
Joint Learning of Pose Regression and Denoising Diffusion with Score Scaling Sampling for Category-level 6D Pose Estimation
von: Lee, Seunghyun, et al.
Veröffentlicht: (2025)
von: Lee, Seunghyun, et al.
Veröffentlicht: (2025)
UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry
von: Um, Tae-Wook, et al.
Veröffentlicht: (2025)
von: Um, Tae-Wook, et al.
Veröffentlicht: (2025)
Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework
von: Zeng, Weixuan, et al.
Veröffentlicht: (2026)
von: Zeng, Weixuan, et al.
Veröffentlicht: (2026)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
DiagNet: Detecting Objects using Diagonal Constraints on Adjacency Matrix of Graph Neural Network
von: Lee, Chong Hyun, et al.
Veröffentlicht: (2025)
von: Lee, Chong Hyun, et al.
Veröffentlicht: (2025)
GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers
von: Song, Zhiye, et al.
Veröffentlicht: (2025)
von: Song, Zhiye, et al.
Veröffentlicht: (2025)
Diffusion-Based sRGB Real Noise Generation via Prompt-Driven Noise Representation Learning
von: Ko, Jaekyun, et al.
Veröffentlicht: (2026)
von: Ko, Jaekyun, et al.
Veröffentlicht: (2026)
DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling
von: Go, Hyojun, et al.
Veröffentlicht: (2025)
von: Go, Hyojun, et al.
Veröffentlicht: (2025)
PackDiT: Joint Human Motion and Text Generation via Mutual Prompting
von: Jiang, Zhongyu, et al.
Veröffentlicht: (2025)
von: Jiang, Zhongyu, et al.
Veröffentlicht: (2025)
LuxDiT: Lighting Estimation with Video Diffusion Transformer
von: Liang, Ruofan, et al.
Veröffentlicht: (2025)
von: Liang, Ruofan, et al.
Veröffentlicht: (2025)
FMA-Net++: Motion- and Exposure-Aware Real-World Joint Video Super-Resolution and Deblurring
von: Youk, Geunhyuk, et al.
Veröffentlicht: (2025)
von: Youk, Geunhyuk, et al.
Veröffentlicht: (2025)
HQ-DiT: Efficient Diffusion Transformer with FP4 Hybrid Quantization
von: Liu, Wenxuan, et al.
Veröffentlicht: (2024)
von: Liu, Wenxuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Zero-shot Depth Completion via Test-time Alignment with Affine-invariant Depth Prior
von: Hyoseok, Lee, et al.
Veröffentlicht: (2025) -
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
von: Lim, Sohwi, et al.
Veröffentlicht: (2026) -
Early Failure Detection and Intervention in Video Diffusion Models
von: Byung-Ki, Kwon, et al.
Veröffentlicht: (2026) -
Measurement-Consistent Langevin Corrector for Stabilizing Latent Diffusion Inverse Problem Solvers
von: Hyoseok, Lee, et al.
Veröffentlicht: (2026) -
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)