Learning to Steer: Input-dependent Steering for Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Parekh, Jayneel, Khayatan, Pegah, Shukor, Mustafa, Dapogny, Arnaud, Newson, Alasdair, Cord, Matthieu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
by: Khayatan, Pegah, et al.
Published: (2025)
by: Khayatan, Pegah, et al.
Published: (2025)
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
by: Khayatan, Pegah, et al.
Published: (2026)
by: Khayatan, Pegah, et al.
Published: (2026)
A Concept-Based Explainability Framework for Large Multimodal Models
by: Parekh, Jayneel, et al.
Published: (2024)
by: Parekh, Jayneel, et al.
Published: (2024)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Skipping Computations in Multimodal LLMs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
What Makes Multimodal In-Context Learning Work?
by: Baldassini, Folco Bertini, et al.
Published: (2024)
by: Baldassini, Folco Bertini, et al.
Published: (2024)
Restyling Unsupervised Concept Based Interpretable Networks with Generative Models
by: Parekh, Jayneel, et al.
Published: (2024)
by: Parekh, Jayneel, et al.
Published: (2024)
Steering the Verifiability of Multimodal AI Hallucinations
by: Pang, Jianhong, et al.
Published: (2026)
by: Pang, Jianhong, et al.
Published: (2026)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
by: Nguyen, Duy, et al.
Published: (2025)
by: Nguyen, Duy, et al.
Published: (2025)
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
by: Vallaeys, Théophane, et al.
Published: (2024)
by: Vallaeys, Théophane, et al.
Published: (2024)
Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
by: Shukor, Mustafa, et al.
Published: (2023)
by: Shukor, Mustafa, et al.
Published: (2023)
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
by: Caselles-Dupré, Hugo, et al.
Published: (2026)
by: Caselles-Dupré, Hugo, et al.
Published: (2026)
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
by: Bhaskar, Paramananda, et al.
Published: (2026)
by: Bhaskar, Paramananda, et al.
Published: (2026)
ReGentS: Real-World Safety-Critical Driving Scenario Generation Made Stable
by: Yin, Yuan, et al.
Published: (2024)
by: Yin, Yuan, et al.
Published: (2024)
Zero-Shot Refinement of Buildings' Segmentation Models using SAM
by: Mayladan, Ali, et al.
Published: (2023)
by: Mayladan, Ali, et al.
Published: (2023)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
by: Hua, Zhenglin, et al.
Published: (2025)
by: Hua, Zhenglin, et al.
Published: (2025)
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
by: Balakrishnan, Ravikumar, et al.
Published: (2025)
by: Balakrishnan, Ravikumar, et al.
Published: (2025)
VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models
by: Phute, Mansi, et al.
Published: (2025)
by: Phute, Mansi, et al.
Published: (2025)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
by: Liu, Zeyu, et al.
Published: (2026)
by: Liu, Zeyu, et al.
Published: (2026)
Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
by: Wang, Mengru, et al.
Published: (2025)
by: Wang, Mengru, et al.
Published: (2025)
MedSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
by: Pham, Trong-Thang, et al.
Published: (2026)
by: Pham, Trong-Thang, et al.
Published: (2026)
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
by: Hinojosa, Carlos, et al.
Published: (2026)
by: Hinojosa, Carlos, et al.
Published: (2026)
Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs
by: Fu, Xingyu, et al.
Published: (2025)
by: Fu, Xingyu, et al.
Published: (2025)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
by: Yin, Jianghao, et al.
Published: (2026)
by: Yin, Jianghao, et al.
Published: (2026)
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
by: Li, Yunxin, et al.
Published: (2023)
by: Li, Yunxin, et al.
Published: (2023)
Scientific Reasoning: Assessment of Multimodal Generative LLMs
by: Dreyer, Florian, et al.
Published: (2025)
by: Dreyer, Florian, et al.
Published: (2025)
List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
by: Yan, An, et al.
Published: (2024)
by: Yan, An, et al.
Published: (2024)
Inline Critic Steers Image Editing
by: Kang, Weitai, et al.
Published: (2026)
by: Kang, Weitai, et al.
Published: (2026)
EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness
by: Sun, Yueru, et al.
Published: (2026)
by: Sun, Yueru, et al.
Published: (2026)
DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut
by: Couairon, Paul, et al.
Published: (2024)
by: Couairon, Paul, et al.
Published: (2024)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
by: Zhao, Xiangyu, et al.
Published: (2023)
by: Zhao, Xiangyu, et al.
Published: (2023)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
Moment Sampling in Video LLMs for Long-Form Video QA
by: Chasmai, Mustafa, et al.
Published: (2025)
by: Chasmai, Mustafa, et al.
Published: (2025)
Decoding Vision Transformers: the Diffusion Steering Lens
by: Takatsuki, Ryota, et al.
Published: (2025)
by: Takatsuki, Ryota, et al.
Published: (2025)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
by: Sarto, Sara, et al.
Published: (2025)
by: Sarto, Sara, et al.
Published: (2025)
Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring
by: Onsu, Murat Arda, et al.
Published: (2025)
by: Onsu, Murat Arda, et al.
Published: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
Similar Items
-
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
by: Khayatan, Pegah, et al.
Published: (2025) -
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
by: Khayatan, Pegah, et al.
Published: (2026) -
A Concept-Based Explainability Framework for Large Multimodal Models
by: Parekh, Jayneel, et al.
Published: (2024) -
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024) -
Skipping Computations in Multimodal LLMs
by: Shukor, Mustafa, et al.
Published: (2024)