Towards Understanding Multimodal Fine-Tuning: Spatial Features
Fuente:
arXiv
Saved in:
| Main Authors: | Naghashyar, Lachin, Batra, Hunar, Khakzar, Ashkan, Torr, Philip, Clark, Ronald, de Witt, Christian Schroeder, Venhoff, Constantin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Visual Representations Map to Language Feature Space in Multimodal LLMs
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Improving Myocardial Infarction Detection via Synthetic ECG Pretraining
by: Naghashyar, Lachin
Published: (2025)
by: Naghashyar, Lachin
Published: (2025)
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
by: Deb, Oishi, et al.
Published: (2025)
by: Deb, Oishi, et al.
Published: (2025)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
by: Batra, Hunar, et al.
Published: (2025)
by: Batra, Hunar, et al.
Published: (2025)
Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
by: Rezaei, Razieh, et al.
Published: (2024)
by: Rezaei, Razieh, et al.
Published: (2024)
DreamPolisher: Towards High-Quality Text-to-3D Generation via Geometric Diffusion
by: Lin, Yuanze, et al.
Published: (2024)
by: Lin, Yuanze, et al.
Published: (2024)
Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
by: Hemmat, Arshia, et al.
Published: (2024)
by: Hemmat, Arshia, et al.
Published: (2024)
Latent Guard: a Safety Framework for Text-to-image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Learnable Sparsity for Vision Generative Models
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Mixture of Experts Made Intrinsically Interpretable
by: Yang, Xingyi, et al.
Published: (2025)
by: Yang, Xingyi, et al.
Published: (2025)
Minimalist Concept Erasure in Generative Models
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning
by: Qin, Chuan, et al.
Published: (2026)
by: Qin, Chuan, et al.
Published: (2026)
SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders
by: Venhoff, Constantin, et al.
Published: (2024)
by: Venhoff, Constantin, et al.
Published: (2024)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
A Survey on Transferability of Adversarial Examples across Deep Neural Networks
by: Gu, Jindong, et al.
Published: (2023)
by: Gu, Jindong, et al.
Published: (2023)
Parameter Efficient Fine-Tuning of Segment Anything Model for Biomedical Imaging
by: Teuber, Carolin, et al.
Published: (2025)
by: Teuber, Carolin, et al.
Published: (2025)
EVCL: Elastic Variational Continual Learning with Weight Consolidation
by: Batra, Hunar, et al.
Published: (2024)
by: Batra, Hunar, et al.
Published: (2024)
Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2026)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2026)
Temporal Gains, Spatial Costs: Revisiting Video Fine-Tuning in Multimodal Large Language Models
by: Zhang, Linghao, et al.
Published: (2026)
by: Zhang, Linghao, et al.
Published: (2026)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
by: Hufe, Lorenz, et al.
Published: (2025)
by: Hufe, Lorenz, et al.
Published: (2025)
Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge
by: Lin, Yuanze, et al.
Published: (2024)
by: Lin, Yuanze, et al.
Published: (2024)
MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
by: Cao, Yue, et al.
Published: (2024)
by: Cao, Yue, et al.
Published: (2024)
Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
by: Yuan, Jianhao, et al.
Published: (2022)
by: Yuan, Jianhao, et al.
Published: (2022)
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
by: Wu, Yixuan, et al.
Published: (2025)
by: Wu, Yixuan, et al.
Published: (2025)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
Scene-Conditional 3D Object Stylization and Composition
by: Zhou, Jinghao, et al.
Published: (2023)
by: Zhou, Jinghao, et al.
Published: (2023)
MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
by: Ashraf, Tajamul, et al.
Published: (2025)
by: Ashraf, Tajamul, et al.
Published: (2025)
SAFT: Towards Out-of-Distribution Generalization in Fine-Tuning
by: Nguyen, Bac, et al.
Published: (2024)
by: Nguyen, Bac, et al.
Published: (2024)
Olympus: A Universal Task Router for Computer Vision Tasks
by: Lin, Yuanze, et al.
Published: (2024)
by: Lin, Yuanze, et al.
Published: (2024)
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
by: Brown, Ellis, et al.
Published: (2025)
by: Brown, Ellis, et al.
Published: (2025)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
by: Tong, Shengbang, et al.
Published: (2024)
by: Tong, Shengbang, et al.
Published: (2024)
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
by: Ye, Qilang, et al.
Published: (2024)
by: Ye, Qilang, et al.
Published: (2024)
NeRF-VPT: Learning Novel View Representations with Neural Radiance Fields via View Prompt Tuning
by: Chen, Linsheng, et al.
Published: (2024)
by: Chen, Linsheng, et al.
Published: (2024)
Fine-tuning can cripple your foundation model; preserving features may be the solution
by: Mukhoti, Jishnu, et al.
Published: (2023)
by: Mukhoti, Jishnu, et al.
Published: (2023)
Directly Fine-Tuning Diffusion Models on Differentiable Rewards
by: Clark, Kevin, et al.
Published: (2023)
by: Clark, Kevin, et al.
Published: (2023)
Improving 2D Feature Representations by 3D-Aware Fine-Tuning
by: Yue, Yuanwen, et al.
Published: (2024)
by: Yue, Yuanwen, et al.
Published: (2024)
LDP: Parameter-Efficient Fine-Tuning of Multimodal LLM for Medical Report Generation
by: Zhou, Tianyu, et al.
Published: (2025)
by: Zhou, Tianyu, et al.
Published: (2025)
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
PanAdapter: Two-Stage Fine-Tuning with Spatial-Spectral Priors Injecting for Pansharpening
by: Wu, RuoCheng, et al.
Published: (2024)
by: Wu, RuoCheng, et al.
Published: (2024)
Similar Items
-
How Visual Representations Map to Language Feature Space in Multimodal LLMs
by: Venhoff, Constantin, et al.
Published: (2025) -
Improving Myocardial Infarction Detection via Synthetic ECG Pretraining
by: Naghashyar, Lachin
Published: (2025) -
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
by: Deb, Oishi, et al.
Published: (2025) -
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
by: Batra, Hunar, et al.
Published: (2025) -
Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval
by: Venhoff, Constantin, et al.
Published: (2025)