FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zehan, Zhang, Ziang, Cheng, Xize, Huang, Rongjie, Liu, Luping, Ye, Zhenhui, Huang, Haifeng, Zhao, Yang, Jin, Tao, Gao, Peng, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
Unleashing the Power of Natural Audio Featuring Multiple Sound Sources
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
von: Huang, Haifeng, et al.
Veröffentlicht: (2023)
von: Huang, Haifeng, et al.
Veröffentlicht: (2023)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
von: Ye, Zhenhui, et al.
Veröffentlicht: (2024)
von: Ye, Zhenhui, et al.
Veröffentlicht: (2024)
Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding
von: Jiang, Yibo, et al.
Veröffentlicht: (2026)
von: Jiang, Yibo, et al.
Veröffentlicht: (2026)
FreeInv: Free Lunch for Improving DDIM Inversion
von: Bao, Yuxiang, et al.
Veröffentlicht: (2025)
von: Bao, Yuxiang, et al.
Veröffentlicht: (2025)
Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
von: Hu, Yijie, et al.
Veröffentlicht: (2025)
von: Hu, Yijie, et al.
Veröffentlicht: (2025)
No Free Lunch with Guardrails
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2025)
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2025)
Leveraging Semantic Attribute Binding for Free-Lunch Color Control in Diffusion Models
von: Laria, Héctor, et al.
Veröffentlicht: (2025)
von: Laria, Héctor, et al.
Veröffentlicht: (2025)
Data Augmentation as Free Lunch: Exploring the Test-Time Augmentation for Sequential Recommendation
von: Dang, Yizhou, et al.
Veröffentlicht: (2025)
von: Dang, Yizhou, et al.
Veröffentlicht: (2025)
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
Switch EMA: A Free Lunch for Better Flatness and Sharpness
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
No Free Lunch for Approximate MCMC
von: Johndrow, James E., et al.
Veröffentlicht: (2020)
von: Johndrow, James E., et al.
Veröffentlicht: (2020)
FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers
von: Zhang, Yanbing, et al.
Veröffentlicht: (2025)
von: Zhang, Yanbing, et al.
Veröffentlicht: (2025)
A No Free Lunch Theorem for Human-AI Collaboration
von: Peng, Kenny, et al.
Veröffentlicht: (2024)
von: Peng, Kenny, et al.
Veröffentlicht: (2024)
Free Lunch in Pathology Foundation Model: Task-specific Model Adaptation with Concept-Guided Feature Enhancement
von: Huang, Yanyan, et al.
Veröffentlicht: (2024)
von: Huang, Yanyan, et al.
Veröffentlicht: (2024)
MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
von: Zhang, Zihan, et al.
Veröffentlicht: (2025)
von: Zhang, Zihan, et al.
Veröffentlicht: (2025)
Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment
von: Hong, Zhiqing, et al.
Veröffentlicht: (2024)
von: Hong, Zhiqing, et al.
Veröffentlicht: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
Energy-Calibrated VAE with Test Time Free Lunch
von: Luo, Yihong, et al.
Veröffentlicht: (2023)
von: Luo, Yihong, et al.
Veröffentlicht: (2023)
GenSpace: Benchmarking Spatially-Aware Image Generation
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
von: Huang, Haifeng, et al.
Veröffentlicht: (2026)
von: Huang, Haifeng, et al.
Veröffentlicht: (2026)
MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
No Free Lunch Theorem for Privacy-Preserving LLM Inference
von: Zhang, Xiaojin, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaojin, et al.
Veröffentlicht: (2024)
CineScale: Free Lunch in High-Resolution Cinematic Visual Generation
von: Qiu, Haonan, et al.
Veröffentlicht: (2025)
von: Qiu, Haonan, et al.
Veröffentlicht: (2025)
No-Free-Lunch Theories for Tensor-Network Machine Learning Models
von: Wu, Jing-Chuan, et al.
Veröffentlicht: (2024)
von: Wu, Jing-Chuan, et al.
Veröffentlicht: (2024)
RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers
von: Zhao, Min, et al.
Veröffentlicht: (2025)
von: Zhao, Min, et al.
Veröffentlicht: (2025)
No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference
von: Mani, Pranav, et al.
Veröffentlicht: (2025)
von: Mani, Pranav, et al.
Veröffentlicht: (2025)
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
von: Zhang, Yanzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Yanzhi, et al.
Veröffentlicht: (2025)
Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
FreeCond: Free Lunch in the Input Conditions of Text-Guided Inpainting
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2024)
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2024)
Free Lunch for Stabilizing Rectified Flow Inversion
von: Wang, Chenru, et al.
Veröffentlicht: (2026)
von: Wang, Chenru, et al.
Veröffentlicht: (2026)
Private Means and the Curious Incident of the Free Lunch
von: Fitzsimons, Jack, et al.
Veröffentlicht: (2024)
von: Fitzsimons, Jack, et al.
Veröffentlicht: (2024)
Free Lunch for Generating Effective Outlier Supervision
von: Pei, Sen, et al.
Veröffentlicht: (2023)
von: Pei, Sen, et al.
Veröffentlicht: (2023)
No Free Lunch: Research Software Testing in Teaching
von: Dorner, Michael, et al.
Veröffentlicht: (2024)
von: Dorner, Michael, et al.
Veröffentlicht: (2024)
No Free Lunch for Stochastic Gradient Langevin Dynamics
von: Pillai, Natesh S., et al.
Veröffentlicht: (2024)
von: Pillai, Natesh S., et al.
Veröffentlicht: (2024)
Orient Anything V2: Unifying Orientation and Rotation Understanding
von: Wang, Zehan, et al.
Veröffentlicht: (2026)
von: Wang, Zehan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
von: Wang, Zehan, et al.
Veröffentlicht: (2024) -
Unleashing the Power of Natural Audio Featuring Multiple Sound Sources
von: Cheng, Xize, et al.
Veröffentlicht: (2025) -
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024) -
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
von: Huang, Haifeng, et al.
Veröffentlicht: (2023) -
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
von: Wang, Zehan, et al.
Veröffentlicht: (2025)