X-Fusion: Introducing New Modality to Frozen Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mo, Sicheng, Nguyen, Thao, Huang, Xun, Iyer, Siddharth Srinivasan, Li, Yijun, Liu, Yuchen, Tandon, Abhishek, Shechtman, Eli, Singh, Krishna Kumar, Lee, Yong Jae, Zhou, Bolei, Li, Yuheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
von: Mo, Sicheng, et al.
Veröffentlicht: (2025)
von: Mo, Sicheng, et al.
Veröffentlicht: (2025)
Relational Visual Similarity
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
von: Li, Yuheng, et al.
Veröffentlicht: (2024)
von: Li, Yuheng, et al.
Veröffentlicht: (2024)
YoChameleon: Personalized Vision and Language Generation
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
Edit One for All: Interactive Batch Image Editing
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
von: Huang, Xun, et al.
Veröffentlicht: (2025)
von: Huang, Xun, et al.
Veröffentlicht: (2025)
Can abstract concepts from LLM improve SLM performance?
von: Tandon, Siddharth
Veröffentlicht: (2025)
von: Tandon, Siddharth
Veröffentlicht: (2025)
Yo'LLaVA: Your Personalized Language and Vision Assistant
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
Learning an Image Editing Model without Image Editing Pairs
von: Kumari, Nupur, et al.
Veröffentlicht: (2025)
von: Kumari, Nupur, et al.
Veröffentlicht: (2025)
Causality in Video Diffusers is Separable from Denoising
von: Bai, Xingjian, et al.
Veröffentlicht: (2026)
von: Bai, Xingjian, et al.
Veröffentlicht: (2026)
Personal Visual Memory from Explicit and Implicit Evidence
von: Nguyen, Viet, et al.
Veröffentlicht: (2026)
von: Nguyen, Viet, et al.
Veröffentlicht: (2026)
Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2026)
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2026)
Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
von: Lin, Kuan Heng, et al.
Veröffentlicht: (2024)
von: Lin, Kuan Heng, et al.
Veröffentlicht: (2024)
Lifting for Arbitrary Gadgets
von: Iyer, Siddharth
Veröffentlicht: (2025)
von: Iyer, Siddharth
Veröffentlicht: (2025)
Gaps between quadratic forms
von: Iyer, Siddharth
Veröffentlicht: (2025)
von: Iyer, Siddharth
Veröffentlicht: (2025)
Cubic Polynomials and Sums of Two Squares
von: Iyer, Siddharth
Veröffentlicht: (2025)
von: Iyer, Siddharth
Veröffentlicht: (2025)
Distribution of sums of square roots modulo $1$
von: Iyer, Siddharth
Veröffentlicht: (2024)
von: Iyer, Siddharth
Veröffentlicht: (2024)
Rational approximation with digit-restricted denominators
von: Iyer, Siddharth
Veröffentlicht: (2023)
von: Iyer, Siddharth
Veröffentlicht: (2023)
Distribution of $θ-$powers and their sums
von: Iyer, Siddharth
Veröffentlicht: (2025)
von: Iyer, Siddharth
Veröffentlicht: (2025)
On the Digits of Partition Functions
von: Iyer, Siddharth
Veröffentlicht: (2026)
von: Iyer, Siddharth
Veröffentlicht: (2026)
Dreamland: Controllable World Creation with Simulator and Generative Models
von: Mo, Sicheng, et al.
Veröffentlicht: (2025)
von: Mo, Sicheng, et al.
Veröffentlicht: (2025)
MotionStream: Real-Time Video Generation with Interactive Motion Controls
von: Shin, Joonghyuk, et al.
Veröffentlicht: (2025)
von: Shin, Joonghyuk, et al.
Veröffentlicht: (2025)
Towards Universal Fake Image Detectors that Generalize Across Generative Models
von: Ojha, Utkarsh, et al.
Veröffentlicht: (2023)
von: Ojha, Utkarsh, et al.
Veröffentlicht: (2023)
A Generalized Multi-Modal Fusion Detection Framework
von: Cui, Leichao, et al.
Veröffentlicht: (2023)
von: Cui, Leichao, et al.
Veröffentlicht: (2023)
MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
von: Lin, Xi Victoria, et al.
Veröffentlicht: (2024)
von: Lin, Xi Victoria, et al.
Veröffentlicht: (2024)
FLAME: Empowering Frozen LLMs for Knowledge Graph Completion
von: Xue, Bo, et al.
Veröffentlicht: (2024)
von: Xue, Bo, et al.
Veröffentlicht: (2024)
From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
von: Yin, Tianwei, et al.
Veröffentlicht: (2024)
von: Yin, Tianwei, et al.
Veröffentlicht: (2024)
MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
DAE-Fuse: An Adaptive Discriminative Autoencoder for Multi-Modality Image Fusion
von: Guo, Yuchen, et al.
Veröffentlicht: (2024)
von: Guo, Yuchen, et al.
Veröffentlicht: (2024)
Improved Baselines with Visual Instruction Tuning
von: Liu, Haotian, et al.
Veröffentlicht: (2023)
von: Liu, Haotian, et al.
Veröffentlicht: (2023)
Detection of lumpy skin disease virus reads in the human upper respiratory tract microbiome requires further investigation
von: Siddharth Singh Tomar, et al.
Veröffentlicht: (2024)
von: Siddharth Singh Tomar, et al.
Veröffentlicht: (2024)
Improved Baselines with Representation Autoencoders
von: Singh, Jaskirat, et al.
Veröffentlicht: (2026)
von: Singh, Jaskirat, et al.
Veröffentlicht: (2026)
SimGen: Simulator-conditioned Driving Scene Generation
von: Zhou, Yunsong, et al.
Veröffentlicht: (2024)
von: Zhou, Yunsong, et al.
Veröffentlicht: (2024)
SPLD polynomial optimization and bounded degree SOS hierarchies
von: Jiao, Liguo, et al.
Veröffentlicht: (2025)
von: Jiao, Liguo, et al.
Veröffentlicht: (2025)
Alternating projections between two inconsistent affine subspaces with varying relaxation
von: Thao, Nguyen T.
Veröffentlicht: (2025)
von: Thao, Nguyen T.
Veröffentlicht: (2025)
An XOR Lemma for Deterministic Communication Complexity
von: Iyer, Siddharth, et al.
Veröffentlicht: (2024)
von: Iyer, Siddharth, et al.
Veröffentlicht: (2024)
XOR Lemmas for Communication via Marginal Information
von: Iyer, Siddharth, et al.
Veröffentlicht: (2023)
von: Iyer, Siddharth, et al.
Veröffentlicht: (2023)
SnAG: Scalable and Accurate Video Grounding
von: Mu, Fangzhou, et al.
Veröffentlicht: (2024)
von: Mu, Fangzhou, et al.
Veröffentlicht: (2024)
Exact Federated Continual Unlearning for Ridge Heads on Frozen Foundation Models
von: Quan, Yijun, et al.
Veröffentlicht: (2026)
von: Quan, Yijun, et al.
Veröffentlicht: (2026)
FW-Shapley: Real-time Estimation of Weighted Shapley Values
von: Panda, Pranoy, et al.
Veröffentlicht: (2025)
von: Panda, Pranoy, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
von: Mo, Sicheng, et al.
Veröffentlicht: (2025) -
Relational Visual Similarity
von: Nguyen, Thao, et al.
Veröffentlicht: (2025) -
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
von: Li, Yuheng, et al.
Veröffentlicht: (2024) -
YoChameleon: Personalized Vision and Language Generation
von: Nguyen, Thao, et al.
Veröffentlicht: (2025) -
Edit One for All: Interactive Batch Image Editing
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)