Wings: Learning Multimodal LLMs without Text-only Forgetting
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yi-Kai, Lu, Shiyin, Li, Yang, Ma, Yanqing, Chen, Qing-Guo, Xu, Zhao, Luo, Weihua, Zhang, Kaifu, Zhan, De-Chuan, Ye, Han-Jia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
by: Lu, Shiyin, et al.
Published: (2024)
by: Lu, Shiyin, et al.
Published: (2024)
Multimodal Tabular Reasoning with Privileged Structured Information
by: Jiang, Jun-Peng, et al.
Published: (2025)
by: Jiang, Jun-Peng, et al.
Published: (2025)
Parrot: Multilingual Visual Instruction Tuning
by: Sun, Hai-Long, et al.
Published: (2024)
by: Sun, Hai-Long, et al.
Published: (2024)
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
by: Zhang, Yi-Kai, et al.
Published: (2024)
by: Zhang, Yi-Kai, et al.
Published: (2024)
MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
by: Sun, Hui, et al.
Published: (2025)
by: Sun, Hui, et al.
Published: (2025)
Learning without Forgetting for Vision-Language Models
by: Zhou, Da-Wei, et al.
Published: (2023)
by: Zhou, Da-Wei, et al.
Published: (2023)
Capability Instruction Tuning: A New Paradigm for Dynamic LLM Routing
by: Zhang, Yi-Kai, et al.
Published: (2025)
by: Zhang, Yi-Kai, et al.
Published: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
by: Yang, Wenhao, et al.
Published: (2025)
by: Yang, Wenhao, et al.
Published: (2025)
CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization
by: Yang, Wenhao, et al.
Published: (2026)
by: Yang, Wenhao, et al.
Published: (2026)
SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
by: Chen, Zhi-Kai, et al.
Published: (2025)
by: Chen, Zhi-Kai, et al.
Published: (2025)
Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification
by: Yi, Chao, et al.
Published: (2024)
by: Yi, Chao, et al.
Published: (2024)
Model Assembly Learning with Heterogeneous Layer Weight Merging
by: Zhang, Yi-Kai, et al.
Published: (2025)
by: Zhang, Yi-Kai, et al.
Published: (2025)
Improving LLMs for Recommendation with Out-Of-Vocabulary Tokens
by: Huang, Ting-Ji, et al.
Published: (2024)
by: Huang, Ting-Ji, et al.
Published: (2024)
Building Decision Making Models Through Language Model Regime
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
by: Gao, Sensen, et al.
Published: (2025)
by: Gao, Sensen, et al.
Published: (2025)
Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees
by: Chen, Sijia, et al.
Published: (2024)
by: Chen, Sijia, et al.
Published: (2024)
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
SeeingEye: Agentic Information Flow Unlocks Multimodal Reasoning In Text-only LLMs
by: Zhang, Weijia, et al.
Published: (2025)
by: Zhang, Weijia, et al.
Published: (2025)
Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
(Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
SOFTS: Efficient Multivariate Time Series Forecasting with Series-Core Fusion
by: Han, Lu, et al.
Published: (2024)
by: Han, Lu, et al.
Published: (2024)
MIETT: Multi-Instance Encrypted Traffic Transformer for Encrypted Traffic Classification
by: Chen, Xu-Yang, et al.
Published: (2024)
by: Chen, Xu-Yang, et al.
Published: (2024)
Can LLMs Learn New Concepts Incrementally without Forgetting?
by: Zheng, Junhao, et al.
Published: (2024)
by: Zheng, Junhao, et al.
Published: (2024)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025)
by: Zhao, Shanshan, et al.
Published: (2025)
Bridge the Modality and Capability Gaps in Vision-Language Model Selection
by: Yi, Chao, et al.
Published: (2024)
by: Yi, Chao, et al.
Published: (2024)
Challenging Multilingual LLMs: A New Taxonomy and Benchmark for Unraveling Hallucination in Translation
by: Wu, Xinwei, et al.
Published: (2025)
by: Wu, Xinwei, et al.
Published: (2025)
LLM-I: LLMs are Naturally Interleaved Multimodal Creators
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
Sub‐100 nm manipulation of blue light over a large field of view using Si nanolens array
by: Zhiyuan Shi, et al.
Published: (2025)
by: Zhiyuan Shi, et al.
Published: (2025)
Towards Lightweight, Adaptive and Attribute-Aware Multi-Aspect Controllable Text Generation with Large Language Models
by: Zhu, Chenyu, et al.
Published: (2025)
by: Zhu, Chenyu, et al.
Published: (2025)
$V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
by: Zhang, Yi-Kai, et al.
Published: (2026)
by: Zhang, Yi-Kai, et al.
Published: (2026)
Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs
by: Wu, Xinwei, et al.
Published: (2026)
by: Wu, Xinwei, et al.
Published: (2026)
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
Not Just Object, But State: Compositional Incremental Learning without Forgetting
by: Zhang, Yanyi, et al.
Published: (2024)
by: Zhang, Yanyi, et al.
Published: (2024)
One-Embedding-Fits-All: Efficient Zero-Shot Time Series Forecasting by a Model Zoo
by: Shi, Hao-Nan, et al.
Published: (2025)
by: Shi, Hao-Nan, et al.
Published: (2025)
Similar Items
-
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
by: Lu, Shiyin, et al.
Published: (2024) -
Multimodal Tabular Reasoning with Privileged Structured Information
by: Jiang, Jun-Peng, et al.
Published: (2025) -
Parrot: Multilingual Visual Instruction Tuning
by: Sun, Hai-Long, et al.
Published: (2024) -
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
by: Zhang, Yi-Kai, et al.
Published: (2024) -
MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
by: Sun, Hui, et al.
Published: (2025)