FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zehan, Zhang, Ziang, Cheng, Xize, Huang, Rongjie, Liu, Luping, Ye, Zhenhui, Huang, Haifeng, Zhao, Yang, Jin, Tao, Gao, Peng, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
Unleashing the Power of Natural Audio Featuring Multiple Sound Sources
by: Cheng, Xize, et al.
Published: (2025)
by: Cheng, Xize, et al.
Published: (2025)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024)
by: Cheng, Xize, et al.
Published: (2024)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
by: Huang, Haifeng, et al.
Published: (2023)
by: Huang, Haifeng, et al.
Published: (2023)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
by: Ye, Zhenhui, et al.
Published: (2024)
by: Ye, Zhenhui, et al.
Published: (2024)
Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding
by: Jiang, Yibo, et al.
Published: (2026)
by: Jiang, Yibo, et al.
Published: (2026)
FreeInv: Free Lunch for Improving DDIM Inversion
by: Bao, Yuxiang, et al.
Published: (2025)
by: Bao, Yuxiang, et al.
Published: (2025)
Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
by: Hu, Yijie, et al.
Published: (2025)
by: Hu, Yijie, et al.
Published: (2025)
No Free Lunch with Guardrails
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
Leveraging Semantic Attribute Binding for Free-Lunch Color Control in Diffusion Models
by: Laria, Héctor, et al.
Published: (2025)
by: Laria, Héctor, et al.
Published: (2025)
Data Augmentation as Free Lunch: Exploring the Test-Time Augmentation for Sequential Recommendation
by: Dang, Yizhou, et al.
Published: (2025)
by: Dang, Yizhou, et al.
Published: (2025)
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
by: Guo, Zirun, et al.
Published: (2024)
by: Guo, Zirun, et al.
Published: (2024)
Switch EMA: A Free Lunch for Better Flatness and Sharpness
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
No Free Lunch for Approximate MCMC
by: Johndrow, James E., et al.
Published: (2020)
by: Johndrow, James E., et al.
Published: (2020)
FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers
by: Zhang, Yanbing, et al.
Published: (2025)
by: Zhang, Yanbing, et al.
Published: (2025)
A No Free Lunch Theorem for Human-AI Collaboration
by: Peng, Kenny, et al.
Published: (2024)
by: Peng, Kenny, et al.
Published: (2024)
Free Lunch in Pathology Foundation Model: Task-specific Model Adaptation with Concept-Guided Feature Enhancement
by: Huang, Yanyan, et al.
Published: (2024)
by: Huang, Yanyan, et al.
Published: (2024)
MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
by: Zhang, Zihan, et al.
Published: (2025)
by: Zhang, Zihan, et al.
Published: (2025)
Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment
by: Hong, Zhiqing, et al.
Published: (2024)
by: Hong, Zhiqing, et al.
Published: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
by: Liu, Huadai, et al.
Published: (2024)
by: Liu, Huadai, et al.
Published: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
by: Cheng, Xize, et al.
Published: (2025)
by: Cheng, Xize, et al.
Published: (2025)
Energy-Calibrated VAE with Test Time Free Lunch
by: Luo, Yihong, et al.
Published: (2023)
by: Luo, Yihong, et al.
Published: (2023)
GenSpace: Benchmarking Spatially-Aware Image Generation
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
by: Huang, Haifeng, et al.
Published: (2026)
by: Huang, Haifeng, et al.
Published: (2026)
MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
by: Li, Ruiqi, et al.
Published: (2024)
by: Li, Ruiqi, et al.
Published: (2024)
No Free Lunch Theorem for Privacy-Preserving LLM Inference
by: Zhang, Xiaojin, et al.
Published: (2024)
by: Zhang, Xiaojin, et al.
Published: (2024)
CineScale: Free Lunch in High-Resolution Cinematic Visual Generation
by: Qiu, Haonan, et al.
Published: (2025)
by: Qiu, Haonan, et al.
Published: (2025)
No-Free-Lunch Theories for Tensor-Network Machine Learning Models
by: Wu, Jing-Chuan, et al.
Published: (2024)
by: Wu, Jing-Chuan, et al.
Published: (2024)
RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)
by: Zhao, Min, et al.
Published: (2025)
No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference
by: Mani, Pranav, et al.
Published: (2025)
by: Mani, Pranav, et al.
Published: (2025)
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
by: Zhang, Yanzhi, et al.
Published: (2025)
by: Zhang, Yanzhi, et al.
Published: (2025)
Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
by: Wang, Yongqi, et al.
Published: (2024)
by: Wang, Yongqi, et al.
Published: (2024)
FreeCond: Free Lunch in the Input Conditions of Text-Guided Inpainting
by: Hsiao, Teng-Fang, et al.
Published: (2024)
by: Hsiao, Teng-Fang, et al.
Published: (2024)
Free Lunch for Stabilizing Rectified Flow Inversion
by: Wang, Chenru, et al.
Published: (2026)
by: Wang, Chenru, et al.
Published: (2026)
Private Means and the Curious Incident of the Free Lunch
by: Fitzsimons, Jack, et al.
Published: (2024)
by: Fitzsimons, Jack, et al.
Published: (2024)
Free Lunch for Generating Effective Outlier Supervision
by: Pei, Sen, et al.
Published: (2023)
by: Pei, Sen, et al.
Published: (2023)
No Free Lunch: Research Software Testing in Teaching
by: Dorner, Michael, et al.
Published: (2024)
by: Dorner, Michael, et al.
Published: (2024)
No Free Lunch for Stochastic Gradient Langevin Dynamics
by: Pillai, Natesh S., et al.
Published: (2024)
by: Pillai, Natesh S., et al.
Published: (2024)
Orient Anything V2: Unifying Orientation and Rotation Understanding
by: Wang, Zehan, et al.
Published: (2026)
by: Wang, Zehan, et al.
Published: (2026)
Similar Items
-
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
by: Wang, Zehan, et al.
Published: (2024) -
Unleashing the Power of Natural Audio Featuring Multiple Sound Sources
by: Cheng, Xize, et al.
Published: (2025) -
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024) -
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
by: Huang, Haifeng, et al.
Published: (2023) -
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
by: Wang, Zehan, et al.
Published: (2025)