Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Zirun, Jin, Tao, Chen, Jingyuan, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
LLM-I: LLMs are Naturally Interleaved Multimodal Creators
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
von: Hong, Minjie, et al.
Veröffentlicht: (2025)
von: Hong, Minjie, et al.
Veröffentlicht: (2025)
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion
von: Chen, Peiyuan, et al.
Veröffentlicht: (2024)
von: Chen, Peiyuan, et al.
Veröffentlicht: (2024)
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
von: Sun, Linzhuang, et al.
Veröffentlicht: (2025)
von: Sun, Linzhuang, et al.
Veröffentlicht: (2025)
Efficient Prompting for Continual Adaptation to Missing Modalities
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning
von: Tang, Jun-Tao, et al.
Veröffentlicht: (2026)
von: Tang, Jun-Tao, et al.
Veröffentlicht: (2026)
Thinking with Programming Vision: Towards a Unified View for Thinking with Images
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
LLM as a Complementary Optimizer to Gradient Descent: A Case Study in Prompt Tuning
von: Guo, Zixian, et al.
Veröffentlicht: (2024)
von: Guo, Zixian, et al.
Veröffentlicht: (2024)
LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning
von: Kowsher, Md, et al.
Veröffentlicht: (2026)
von: Kowsher, Md, et al.
Veröffentlicht: (2026)
Controlling Multimodal LLMs via Reward-guided Decoding
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
Text Role Classification in Scientific Charts Using Multimodal Transformers
von: Kim, Hye Jin, et al.
Veröffentlicht: (2024)
von: Kim, Hye Jin, et al.
Veröffentlicht: (2024)
Hyperbolic Multimodal Representation Learning for Biological Taxonomies
von: Gong, ZeMing, et al.
Veröffentlicht: (2025)
von: Gong, ZeMing, et al.
Veröffentlicht: (2025)
The Double Dilemma in Multi-Task Radiology Report Generation: A Gradient Dynamics Analysis and Solution
von: Zhang, Erjian, et al.
Veröffentlicht: (2026)
von: Zhang, Erjian, et al.
Veröffentlicht: (2026)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
von: Joshi, Siddharth, et al.
Veröffentlicht: (2025)
von: Joshi, Siddharth, et al.
Veröffentlicht: (2025)
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
von: Zou, Yicheng, et al.
Veröffentlicht: (2026)
von: Zou, Yicheng, et al.
Veröffentlicht: (2026)
DreamLLM: Synergistic Multimodal Comprehension and Creation
von: Dong, Runpei, et al.
Veröffentlicht: (2023)
von: Dong, Runpei, et al.
Veröffentlicht: (2023)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
von: Chen, Zhaorun, et al.
Veröffentlicht: (2024)
von: Chen, Zhaorun, et al.
Veröffentlicht: (2024)
Intern-S1: A Scientific Multimodal Foundation Model
von: Bai, Lei, et al.
Veröffentlicht: (2025)
von: Bai, Lei, et al.
Veröffentlicht: (2025)
Improving Multimodal Large Language Models Using Continual Learning
von: Srivastava, Shikhar, et al.
Veröffentlicht: (2024)
von: Srivastava, Shikhar, et al.
Veröffentlicht: (2024)
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
von: Wang, Jinyin, et al.
Veröffentlicht: (2024)
von: Wang, Jinyin, et al.
Veröffentlicht: (2024)
Efficient Domain Adaptation of Multimodal Embeddings using Constrastive Learning
von: Margaritis, Georgios, et al.
Veröffentlicht: (2025)
von: Margaritis, Georgios, et al.
Veröffentlicht: (2025)
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections
von: Imam, Mohamed Fazli, et al.
Veröffentlicht: (2024)
von: Imam, Mohamed Fazli, et al.
Veröffentlicht: (2024)
BEEM: Boosting Performance of Early Exit DNNs using Multi-Exit Classifiers as Experts
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
von: Arazi, Alan, et al.
Veröffentlicht: (2026)
von: Arazi, Alan, et al.
Veröffentlicht: (2026)
C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
von: Chen, Xiuwei, et al.
Veröffentlicht: (2025)
von: Chen, Xiuwei, et al.
Veröffentlicht: (2025)
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation
von: Wang, Yongxin, et al.
Veröffentlicht: (2024)
von: Wang, Yongxin, et al.
Veröffentlicht: (2024)
Lumos : Empowering Multimodal LLMs with Scene Text Recognition
von: Shenoy, Ashish, et al.
Veröffentlicht: (2024)
von: Shenoy, Ashish, et al.
Veröffentlicht: (2024)
Gradient descent with generalized Newton's method
von: Bu, Zhiqi, et al.
Veröffentlicht: (2024)
von: Bu, Zhiqi, et al.
Veröffentlicht: (2024)
Balancing Multimodal Domain Generalization via Gradient Modulation and Projection
von: Li, Hongzhao, et al.
Veröffentlicht: (2026)
von: Li, Hongzhao, et al.
Veröffentlicht: (2026)
MuJo: Multimodal Joint Feature Space Learning for Human Activity Recognition
von: Fritsch, Stefan Gerd, et al.
Veröffentlicht: (2024)
von: Fritsch, Stefan Gerd, et al.
Veröffentlicht: (2024)
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
von: Chen, Shuang, et al.
Veröffentlicht: (2025)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
Modality-Aware SAM: Sharpness-Aware-Minimization Driven Gradient Modulation for Harmonized Multimodal Learning
von: Nowdeh, Hossein R., et al.
Veröffentlicht: (2025)
von: Nowdeh, Hossein R., et al.
Veröffentlicht: (2025)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024) -
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
von: Guo, Zirun, et al.
Veröffentlicht: (2025) -
A Wander Through the Multimodal Landscape: Efficient Transfer Learning via Low-rank Sequence Multimodal Adapter
von: Guo, Zirun, et al.
Veröffentlicht: (2024) -
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
von: Guo, Zirun, et al.
Veröffentlicht: (2025) -
LLM-I: LLMs are Naturally Interleaved Multimodal Creators
von: Guo, Zirun, et al.
Veröffentlicht: (2025)