LLMs Can Evolve Continually on Modality for X-Modal Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Jiazuo, Xiong, Haomiao, Zhang, Lu, Diao, Haiwen, Zhuge, Yunzhi, Hong, Lanqing, Wang, Dong, Lu, Huchuan, He, You, Chen, Long |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
von: Zhang, Lu, et al.
Veröffentlicht: (2025)
von: Zhang, Lu, et al.
Veröffentlicht: (2025)
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
Regularizing Subspace Redundancy of Low-Rank Adaptation
von: Zhu, Yue, et al.
Veröffentlicht: (2025)
von: Zhu, Yue, et al.
Veröffentlicht: (2025)
Towards Cross-Platform Generalization: Domain Adaptive 3D Detection with Augmentation and Pseudo-Labeling
von: Feng, Xiyan, et al.
Veröffentlicht: (2026)
von: Feng, Xiyan, et al.
Veröffentlicht: (2026)
Reinforcing Video Reasoning Segmentation to Think Before It Segments
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2023)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2023)
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
von: Liu, Qing'an, et al.
Veröffentlicht: (2026)
von: Liu, Qing'an, et al.
Veröffentlicht: (2026)
Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
von: Zhuge, Yunzhi, et al.
Veröffentlicht: (2025)
von: Zhuge, Yunzhi, et al.
Veröffentlicht: (2025)
Complementary and Contrastive Learning for Audio-Visual Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
MM-LLMs: Recent Advances in MultiModal Large Language Models
von: Zhang, Duzhen, et al.
Veröffentlicht: (2024)
von: Zhang, Duzhen, et al.
Veröffentlicht: (2024)
Parameter Aware Mamba Model for Multi-task Dense Prediction
von: Yu, Xinzhuo, et al.
Veröffentlicht: (2025)
von: Yu, Xinzhuo, et al.
Veröffentlicht: (2025)
AEQ-Bench: Measuring Empathy of Omni-Modal Large Models
von: Luo, Xuan, et al.
Veröffentlicht: (2026)
von: Luo, Xuan, et al.
Veröffentlicht: (2026)
KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification
von: Zhu, Yue, et al.
Veröffentlicht: (2025)
von: Zhu, Yue, et al.
Veröffentlicht: (2025)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
Can LLMs Reason in the Wild with Programs?
von: Yang, Yuan, et al.
Veröffentlicht: (2024)
von: Yang, Yuan, et al.
Veröffentlicht: (2024)
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
von: Liu, Qianchu, et al.
Veröffentlicht: (2025)
von: Liu, Qianchu, et al.
Veröffentlicht: (2025)
Learning Universal Features for Generalizable Image Forgery Localization
von: Zhao, Hengrun, et al.
Veröffentlicht: (2025)
von: Zhao, Hengrun, et al.
Veröffentlicht: (2025)
Closing the Modality Reasoning Gap for Speech Large Language Models
von: Wang, Chaoren, et al.
Veröffentlicht: (2026)
von: Wang, Chaoren, et al.
Veröffentlicht: (2026)
Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding
von: Zhang, Wenbo, et al.
Veröffentlicht: (2024)
von: Zhang, Wenbo, et al.
Veröffentlicht: (2024)
G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
von: Gao, Jiahui, et al.
Veröffentlicht: (2023)
von: Gao, Jiahui, et al.
Veröffentlicht: (2023)
Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization
von: Xiao, Cihan, et al.
Veröffentlicht: (2026)
von: Xiao, Cihan, et al.
Veröffentlicht: (2026)
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
von: Guan, Xinyu, et al.
Veröffentlicht: (2025)
von: Guan, Xinyu, et al.
Veröffentlicht: (2025)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
von: Park, Simon, et al.
Veröffentlicht: (2025)
von: Park, Simon, et al.
Veröffentlicht: (2025)
X-VILA: Cross-Modality Alignment for Large Language Model
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
When Modalities Remember: Continual Learning for Multimodal Knowledge Graphs
von: Li, Linyu, et al.
Veröffentlicht: (2026)
von: Li, Linyu, et al.
Veröffentlicht: (2026)
Comparing Discrete and Continuous Space LLMs for Speech Recognition
von: Xu, Yaoxun, et al.
Veröffentlicht: (2024)
von: Xu, Yaoxun, et al.
Veröffentlicht: (2024)
GSSF: Generalized Structural Sparse Function for Deep Cross-modal Metric Learning
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory
von: Diao, Haiwen, et al.
Veröffentlicht: (2023)
von: Diao, Haiwen, et al.
Veröffentlicht: (2023)
MMGR: Multi-Modal Generative Reasoning
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
Deep Boosting Learning: A Brand-new Cooperative Approach for Image-Text Matching
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
Can LLMs Reason Over Non-Text Modalities in a Training-Free Manner? A Case Study with In-Context Representation Learning
von: Zhang, Tianle, et al.
Veröffentlicht: (2025)
von: Zhang, Tianle, et al.
Veröffentlicht: (2025)
Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling
von: Li, Junlin, et al.
Veröffentlicht: (2025)
von: Li, Junlin, et al.
Veröffentlicht: (2025)
Is Extending Modality The Right Path Towards Omni-Modality?
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory
von: Zhang, Ce, et al.
Veröffentlicht: (2026)
von: Zhang, Ce, et al.
Veröffentlicht: (2026)
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
von: Zhang, Lu, et al.
Veröffentlicht: (2025) -
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024) -
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025) -
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025) -
SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)