Re-Imagining Multimodal Instruction Tuning: A Representation View
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Yiyang, Liang, James Chenhao, Tang, Ruixiang, Lee, Yugyung, Rabbani, Majid, Dianat, Sohail, Rao, Raghuveer, Huang, Lifu, Liu, Dongfang, Wang, Qifan, Han, Cheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AMD: Automatic Multi-step Distillation of Large-scale Vision Models
por: Han, Cheng, et al.
Publicado: (2024)
por: Han, Cheng, et al.
Publicado: (2024)
Image Translation as Diffusion Visual Programmers
por: Han, Cheng, et al.
Publicado: (2024)
por: Han, Cheng, et al.
Publicado: (2024)
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
por: Zeng, Runjia, et al.
Publicado: (2025)
por: Zeng, Runjia, et al.
Publicado: (2025)
Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
por: Wang, Jiamian, et al.
Publicado: (2024)
por: Wang, Jiamian, et al.
Publicado: (2024)
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
por: Pulakurthi, Prasanna Reddy, et al.
Publicado: (2025)
por: Pulakurthi, Prasanna Reddy, et al.
Publicado: (2025)
M$^2$PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning
por: Wang, Taowen, et al.
Publicado: (2024)
por: Wang, Taowen, et al.
Publicado: (2024)
Effective Dual-Region Augmentation for Reduced Reliance on Large Amounts of Labeled Data
por: Pulakurthi, Prasanna Reddy, et al.
Publicado: (2025)
por: Pulakurthi, Prasanna Reddy, et al.
Publicado: (2025)
Shuffle PatchMix Augmentation with Confidence-Margin Weighted Pseudo-Labels for Enhanced Source-Free Domain Adaptation
por: Pulakurthi, Prasanna Reddy, et al.
Publicado: (2025)
por: Pulakurthi, Prasanna Reddy, et al.
Publicado: (2025)
Latent Chain-of-Thought for Visual Reasoning
por: Sun, Guohao, et al.
Publicado: (2025)
por: Sun, Guohao, et al.
Publicado: (2025)
Visual Self-Refinement for Autoregressive Models
por: Wang, Jiamian, et al.
Publicado: (2025)
por: Wang, Jiamian, et al.
Publicado: (2025)
TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token Ditching
por: Zeng, Runjia, et al.
Publicado: (2026)
por: Zeng, Runjia, et al.
Publicado: (2026)
All You Need is One: Capsule Prompt Tuning with a Single Vector
por: Liu, Yiyang, et al.
Publicado: (2025)
por: Liu, Yiyang, et al.
Publicado: (2025)
Prototypical Transformer as Unified Motion Learners
por: Han, Cheng, et al.
Publicado: (2024)
por: Han, Cheng, et al.
Publicado: (2024)
Multimodal Instruction Tuning with Conditional Mixture of LoRA
por: Shen, Ying, et al.
Publicado: (2024)
por: Shen, Ying, et al.
Publicado: (2024)
Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning?
por: Han, Cheng, et al.
Publicado: (2024)
por: Han, Cheng, et al.
Publicado: (2024)
Visual Fourier Prompt Tuning
por: Zeng, Runjia, et al.
Publicado: (2024)
por: Zeng, Runjia, et al.
Publicado: (2024)
Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics
por: Wang, Taowen, et al.
Publicado: (2024)
por: Wang, Taowen, et al.
Publicado: (2024)
Error-driven Data-efficient Large Multimodal Model Tuning
por: Yao, Barry Menglong, et al.
Publicado: (2024)
por: Yao, Barry Menglong, et al.
Publicado: (2024)
A-SelecT: Automatic Timestep Selection for Diffusion Transformer Representation Learning
por: Liu, Changyu, et al.
Publicado: (2026)
por: Liu, Changyu, et al.
Publicado: (2026)
On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning
por: Liu, Changyu, et al.
Publicado: (2026)
por: Liu, Changyu, et al.
Publicado: (2026)
Probabilistic Token Alignment for Large Language Model Fusion
por: Zeng, Runjia, et al.
Publicado: (2025)
por: Zeng, Runjia, et al.
Publicado: (2025)
Self-supervised Adversarial Training of Monocular Depth Estimation against Physical-World Attacks
por: Cheng, Zhiyuan, et al.
Publicado: (2024)
por: Cheng, Zhiyuan, et al.
Publicado: (2024)
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
por: Xu, Zhiyang, et al.
Publicado: (2024)
por: Xu, Zhiyang, et al.
Publicado: (2024)
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
por: Miller, Luke James, et al.
Publicado: (2026)
por: Miller, Luke James, et al.
Publicado: (2026)
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
por: Shan, Xiaojun, et al.
Publicado: (2025)
por: Shan, Xiaojun, et al.
Publicado: (2025)
Q-Bridge: Code Translation for Quantum Machine Learning via LLMs
por: Zeng, Runjia, et al.
Publicado: (2026)
por: Zeng, Runjia, et al.
Publicado: (2026)
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
por: Zhang, Boxuan, et al.
Publicado: (2026)
por: Zhang, Boxuan, et al.
Publicado: (2026)
Micro-Defects Expose Macro-Fakes: Detecting AI-Generated Images via Local Distributional Shifts
por: Zhang, Boxuan, et al.
Publicado: (2026)
por: Zhang, Boxuan, et al.
Publicado: (2026)
Data Augmentation for High-Fidelity Generation of CAR-T/NK Immunological Synapse Images
por: Zhang, Xiang, et al.
Publicado: (2026)
por: Zhang, Xiang, et al.
Publicado: (2026)
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
por: Xie, Zhen-Hao, et al.
Publicado: (2026)
por: Xie, Zhen-Hao, et al.
Publicado: (2026)
Comparing theApproaches of International Political Economy MacroStrategy of the UnitedStates of America and China Towards the Islamic Republic of Iran
por: Dianat, Hossein
Publicado: (2024)
por: Dianat, Hossein
Publicado: (2024)
Generative Representational Instruction Tuning
por: Muennighoff, Niklas, et al.
Publicado: (2024)
por: Muennighoff, Niklas, et al.
Publicado: (2024)
ReMatch: Boosting Representation through Matching for Multimodal Retrieval
por: Liu, Qianying, et al.
Publicado: (2025)
por: Liu, Qianying, et al.
Publicado: (2025)
A Unified Causal View of Instruction Tuning
por: Chen, Lu, et al.
Publicado: (2024)
por: Chen, Lu, et al.
Publicado: (2024)
Predicting When to Trust Vision-Language Models for Spatial Reasoning
por: Imran, Muhammad, et al.
Publicado: (2026)
por: Imran, Muhammad, et al.
Publicado: (2026)
DGTEN: A Robust Deep Gaussian based Graph Neural Network for Dynamic Trust Evaluation with Uncertainty-Quantification Support
por: Usman, Muhammad, et al.
Publicado: (2025)
por: Usman, Muhammad, et al.
Publicado: (2025)
QuIC: A Training-Free Quantum Graph Embedding from Ideal Analysis to Practical Hardware Evaluation
por: Miller, Luke, et al.
Publicado: (2026)
por: Miller, Luke, et al.
Publicado: (2026)
Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models
por: Imran, Muhammad, et al.
Publicado: (2025)
por: Imran, Muhammad, et al.
Publicado: (2025)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
por: Wu, Rui, et al.
Publicado: (2026)
por: Wu, Rui, et al.
Publicado: (2026)
Multimodal Instruction Tuning with Hybrid State Space Models
por: Zhou, Jianing, et al.
Publicado: (2024)
por: Zhou, Jianing, et al.
Publicado: (2024)
Ejemplares similares
-
AMD: Automatic Multi-step Distillation of Large-scale Vision Models
por: Han, Cheng, et al.
Publicado: (2024) -
Image Translation as Diffusion Visual Programmers
por: Han, Cheng, et al.
Publicado: (2024) -
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
por: Zeng, Runjia, et al.
Publicado: (2025) -
Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
por: Wang, Jiamian, et al.
Publicado: (2024) -
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
por: Pulakurthi, Prasanna Reddy, et al.
Publicado: (2025)