Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Bader, Jessica, Pach, Mateusz, Bravo, Maria A., Belongie, Serge, Akata, Zeynep |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
by: Pach, Mateusz, et al.
Published: (2026)
by: Pach, Mateusz, et al.
Published: (2026)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
by: Pach, Mateusz, et al.
Published: (2025)
by: Pach, Mateusz, et al.
Published: (2025)
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
Stitched Value Model for Diffusion Alignment
by: Go, Hyojun, et al.
Published: (2026)
by: Go, Hyojun, et al.
Published: (2026)
TORE: Token Recycling in Vision Transformers for Efficient Active Visual Exploration
by: Olszewski, Jan, et al.
Published: (2023)
by: Olszewski, Jan, et al.
Published: (2023)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
by: Wu, Boyong, et al.
Published: (2026)
by: Wu, Boyong, et al.
Published: (2026)
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers
by: Roschmann, Simon, et al.
Published: (2025)
by: Roschmann, Simon, et al.
Published: (2025)
Unlearning-based Neural Interpretations
by: Choi, Ching Lam, et al.
Published: (2024)
by: Choi, Ching Lam, et al.
Published: (2024)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
by: Singhi, Nishad, et al.
Published: (2024)
by: Singhi, Nishad, et al.
Published: (2024)
LucidPPN: Unambiguous Prototypical Parts Network for User-centric Interpretable Computer Vision
by: Pach, Mateusz, et al.
Published: (2024)
by: Pach, Mateusz, et al.
Published: (2024)
(Almost) Free Modality Stitching of Foundation Models
by: Singh, Jaisidh, et al.
Published: (2025)
by: Singh, Jaisidh, et al.
Published: (2025)
MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning
by: Nedungadi, Vishal, et al.
Published: (2024)
by: Nedungadi, Vishal, et al.
Published: (2024)
Concept-Guided Interpretability via Neural Chunking
by: Wu, Shuchen, et al.
Published: (2025)
by: Wu, Shuchen, et al.
Published: (2025)
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
by: Li, Bingyu, et al.
Published: (2024)
by: Li, Bingyu, et al.
Published: (2024)
DataDream: Few-shot Guided Dataset Generation
by: Kim, Jae Myung, et al.
Published: (2024)
by: Kim, Jae Myung, et al.
Published: (2024)
Unbalancedness in Neural Monge Maps Improves Unpaired Domain Translation
by: Eyring, Luca, et al.
Published: (2023)
by: Eyring, Luca, et al.
Published: (2023)
RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers
by: Gokmen, Ahmet Berke, et al.
Published: (2025)
by: Gokmen, Ahmet Berke, et al.
Published: (2025)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
by: Kim, Sanghwan, et al.
Published: (2024)
by: Kim, Sanghwan, et al.
Published: (2024)
OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer Models
by: Chu, Huanpeng, et al.
Published: (2025)
by: Chu, Huanpeng, et al.
Published: (2025)
Fast Training of Diffusion Models with Masked Transformers
by: Zheng, Hongkai, et al.
Published: (2023)
by: Zheng, Hongkai, et al.
Published: (2023)
Multiple Object Stitching for Unsupervised Representation Learning
by: Shen, Chengchao, et al.
Published: (2025)
by: Shen, Chengchao, et al.
Published: (2025)
Revisiting Model Stitching In the Foundation Model Era
by: Mai, Zheda, et al.
Published: (2026)
by: Mai, Zheda, et al.
Published: (2026)
OminiControl: Minimal and Universal Control for Diffusion Transformer
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
Sparse Autoencoders are Topic Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Towards Training-Free Underwater 3D Object Detection from Sonar Point Clouds: A Comparison of Traditional and Deep Learning Approaches
by: Shaukat, M. Salman, et al.
Published: (2025)
by: Shaukat, M. Salman, et al.
Published: (2025)
Geometric 4D Stitching for Grounded 4D Generation
by: Park, Sunwoo, et al.
Published: (2026)
by: Park, Sunwoo, et al.
Published: (2026)
Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy Guidance
by: Sani, Matina Mahdizadeh, et al.
Published: (2026)
by: Sani, Matina Mahdizadeh, et al.
Published: (2026)
T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory Stitching
by: Pan, Zizheng, et al.
Published: (2024)
by: Pan, Zizheng, et al.
Published: (2024)
When Training-Free NAS Meets Vision Transformer: A Neural Tangent Kernel Perspective
by: Zhou, Qiqi, et al.
Published: (2024)
by: Zhou, Qiqi, et al.
Published: (2024)
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
TextDestroyer: A Training- and Annotation-Free Diffusion Method for Destroying Anomal Text from Images
by: Li, Mengcheng, et al.
Published: (2024)
by: Li, Mengcheng, et al.
Published: (2024)
Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces
by: Rojas, Kevin, et al.
Published: (2025)
by: Rojas, Kevin, et al.
Published: (2025)
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
Training-Free Restoration of Pruned Neural Networks
by: Lee, Keonho, et al.
Published: (2025)
by: Lee, Keonho, et al.
Published: (2025)
Training Diffusion Models with Reinforcement Learning
by: Black, Kevin, et al.
Published: (2023)
by: Black, Kevin, et al.
Published: (2023)
Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
by: Eyring, Luca, et al.
Published: (2025)
by: Eyring, Luca, et al.
Published: (2025)
Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM
by: Kim, Jaemin, et al.
Published: (2024)
by: Kim, Jaemin, et al.
Published: (2024)
Video Motion Transfer with Diffusion Transformers
by: Pondaven, Alexander, et al.
Published: (2024)
by: Pondaven, Alexander, et al.
Published: (2024)
Taming Outlier Tokens in Diffusion Transformers
by: Wu, Xiaoyu, et al.
Published: (2026)
by: Wu, Xiaoyu, et al.
Published: (2026)
A Multimodal Architecture for Endpoint Position Prediction in Team-based Multiplayer Games
by: Peche, Jonas, et al.
Published: (2025)
by: Peche, Jonas, et al.
Published: (2025)
Similar Items
-
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
by: Pach, Mateusz, et al.
Published: (2026) -
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
by: Pach, Mateusz, et al.
Published: (2025) -
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
by: Bader, Jessica, et al.
Published: (2025) -
Stitched Value Model for Diffusion Alignment
by: Go, Hyojun, et al.
Published: (2026) -
TORE: Token Recycling in Vision Transformers for Efficient Active Visual Exploration
by: Olszewski, Jan, et al.
Published: (2023)