Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gan, Woody Haosheng, Fu, Deqing, Asilis, Julian, Liu, Ollie, Yogatama, Dani, Sharan, Vatsal, Jia, Robin, Neiswanger, Willie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeLLMa: Decision Making Under Uncertainty with Large Language Models
by: Liu, Ollie, et al.
Published: (2024)
by: Liu, Ollie, et al.
Published: (2024)
IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations
by: Fu, Deqing, et al.
Published: (2024)
by: Fu, Deqing, et al.
Published: (2024)
Resa: Transparent Reasoning Models via SAEs
by: Wang, Shangshang, et al.
Published: (2025)
by: Wang, Shangshang, et al.
Published: (2025)
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
by: Zhang, Jiarui, et al.
Published: (2024)
by: Zhang, Jiarui, et al.
Published: (2024)
Tina: Tiny Reasoning Models via LoRA
by: Wang, Shangshang, et al.
Published: (2025)
by: Wang, Shangshang, et al.
Published: (2025)
Pre-trained Large Language Models Use Fourier Features to Compute Addition
by: Zhou, Tianyi, et al.
Published: (2024)
by: Zhou, Tianyi, et al.
Published: (2024)
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
by: Devic, Siddartha, et al.
Published: (2025)
by: Devic, Siddartha, et al.
Published: (2025)
LLM Unlearning Without an Expert Curated Dataset
by: Zhu, Xiaoyuan, et al.
Published: (2025)
by: Zhu, Xiaoyuan, et al.
Published: (2025)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
by: Fu, Deqing, et al.
Published: (2023)
by: Fu, Deqing, et al.
Published: (2023)
FoNE: Precise Single-Token Number Embeddings via Fourier Features
by: Zhou, Tianyi, et al.
Published: (2025)
by: Zhou, Tianyi, et al.
Published: (2025)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
by: Fu, Deqing, et al.
Published: (2026)
by: Fu, Deqing, et al.
Published: (2026)
Transductive Learning Is Compact
by: Asilis, Julian, et al.
Published: (2024)
by: Asilis, Julian, et al.
Published: (2024)
Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
by: Gan, Woody Haosheng, et al.
Published: (2026)
by: Gan, Woody Haosheng, et al.
Published: (2026)
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
by: Ye, Qilin, et al.
Published: (2025)
by: Ye, Qilin, et al.
Published: (2025)
The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval
by: Chiang, Ting-Rui, et al.
Published: (2025)
by: Chiang, Ting-Rui, et al.
Published: (2025)
Pelican Soup Framework: A Theoretical Framework for Language Model Capabilities
by: Chiang, Ting-Rui, et al.
Published: (2024)
by: Chiang, Ting-Rui, et al.
Published: (2024)
On Retrieval Augmentation and the Limitations of Language Model Training
by: Chiang, Ting-Rui, et al.
Published: (2023)
by: Chiang, Ting-Rui, et al.
Published: (2023)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
Transformers Learn Low Sensitivity Functions: Investigations and Implications
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring
by: Liu, Ollie, et al.
Published: (2025)
by: Liu, Ollie, et al.
Published: (2025)
Limitations on Accurate, Trusted, Human-level Reasoning
by: Panigrahy, Rina, et al.
Published: (2025)
by: Panigrahy, Rina, et al.
Published: (2025)
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
by: Fu, Deqing, et al.
Published: (2023)
by: Fu, Deqing, et al.
Published: (2023)
Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics
by: Lee, Isabelle, et al.
Published: (2024)
by: Lee, Isabelle, et al.
Published: (2024)
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning
by: Akgül, Ömer Faruk, et al.
Published: (2026)
by: Akgül, Ömer Faruk, et al.
Published: (2026)
Proper Learnability and the Role of Unlabeled Data
by: Asilis, Julian, et al.
Published: (2025)
by: Asilis, Julian, et al.
Published: (2025)
Regularization and Optimal Multiclass Learning
by: Asilis, Julian, et al.
Published: (2023)
by: Asilis, Julian, et al.
Published: (2023)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026)
by: Vasudeva, Bhavya, et al.
Published: (2026)
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
by: Manakul, Potsawee, et al.
Published: (2026)
by: Manakul, Potsawee, et al.
Published: (2026)
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
by: Zhu, Wang Bill, et al.
Published: (2026)
by: Zhu, Wang Bill, et al.
Published: (2026)
The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities
by: Wu, Zhaofeng, et al.
Published: (2024)
by: Wu, Zhaofeng, et al.
Published: (2024)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
by: Manakul, Potsawee, et al.
Published: (2025)
by: Manakul, Potsawee, et al.
Published: (2025)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
by: Gordon, Brian, et al.
Published: (2023)
by: Gordon, Brian, et al.
Published: (2023)
LocateBench: Evaluating the Locating Ability of Vision Language Models
by: Chiang, Ting-Rui, et al.
Published: (2024)
by: Chiang, Ting-Rui, et al.
Published: (2024)
Simultaneous Swap Regret Minimization via KL-Calibration
by: Luo, Haipeng, et al.
Published: (2025)
by: Luo, Haipeng, et al.
Published: (2025)
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability
by: Khan, Mohammad Aflah, et al.
Published: (2025)
by: Khan, Mohammad Aflah, et al.
Published: (2025)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
by: Liu, Zeyu, et al.
Published: (2026)
by: Liu, Zeyu, et al.
Published: (2026)
Supercharging Simulation-Based Inference for Bayesian Optimal Experimental Design
by: Klein, Samuel, et al.
Published: (2026)
by: Klein, Samuel, et al.
Published: (2026)
A Unified Approach to Memory-Sample Tradeoffs for Detecting Planted Structures
by: Garg, Sumegha, et al.
Published: (2026)
by: Garg, Sumegha, et al.
Published: (2026)
Similar Items
-
DeLLMa: Decision Making Under Uncertainty with Large Language Models
by: Liu, Ollie, et al.
Published: (2024) -
IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations
by: Fu, Deqing, et al.
Published: (2024) -
Resa: Transparent Reasoning Models via SAEs
by: Wang, Shangshang, et al.
Published: (2025) -
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
by: Zhang, Jiarui, et al.
Published: (2024) -
Tina: Tiny Reasoning Models via LoRA
by: Wang, Shangshang, et al.
Published: (2025)