VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Phute, Mansi, Balakrishnan, Ravikumar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
von: Balakrishnan, Ravikumar, et al.
Veröffentlicht: (2025)
von: Balakrishnan, Ravikumar, et al.
Veröffentlicht: (2025)
VISOR: VIsual Spatial Object Reasoning for Language-driven Object Navigation
von: Taioli, Francesco, et al.
Veröffentlicht: (2026)
von: Taioli, Francesco, et al.
Veröffentlicht: (2026)
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
von: Shen, Yucheng, et al.
Veröffentlicht: (2026)
von: Shen, Yucheng, et al.
Veröffentlicht: (2026)
Language Models Can Explain Visual Features via Steering
von: Ferrando, Javier, et al.
Veröffentlicht: (2026)
von: Ferrando, Javier, et al.
Veröffentlicht: (2026)
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
von: Zhang, Yuanhong, et al.
Veröffentlicht: (2026)
von: Zhang, Yuanhong, et al.
Veröffentlicht: (2026)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
von: Yin, Jianghao, et al.
Veröffentlicht: (2026)
von: Yin, Jianghao, et al.
Veröffentlicht: (2026)
SynthVision -- Harnessing Minimal Input for Maximal Output in Computer Vision Models using Synthetic Image data
von: Kularathne, Yudara, et al.
Veröffentlicht: (2024)
von: Kularathne, Yudara, et al.
Veröffentlicht: (2024)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
von: Wu, Sihao, et al.
Veröffentlicht: (2025)
von: Wu, Sihao, et al.
Veröffentlicht: (2025)
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2026)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
Mitigating Hallucination in Vision-Language Models through Barrier-Regulated Adaptive Closed-form Steering
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Single-Input Multi-Output Model Merging: Leveraging Foundation Models for Dense Multi-Task Learning
von: Giraldo, Juan Garcia, et al.
Veröffentlicht: (2025)
von: Giraldo, Juan Garcia, et al.
Veröffentlicht: (2025)
3D Gaussian and Diffusion-Based Gaze Redirection
von: Panchalingam, Abiram, et al.
Veröffentlicht: (2025)
von: Panchalingam, Abiram, et al.
Veröffentlicht: (2025)
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
VFM-VLM: Vision Foundation Model and Vision Language Model based Visual Comparison for 3D Pose Estimation
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2025)
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2025)
Learning to Steer: Input-dependent Steering for Multimodal LLMs
von: Parekh, Jayneel, et al.
Veröffentlicht: (2025)
von: Parekh, Jayneel, et al.
Veröffentlicht: (2025)
Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models
von: Balakrishnan, Ravikumar, et al.
Veröffentlicht: (2026)
von: Balakrishnan, Ravikumar, et al.
Veröffentlicht: (2026)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
von: Liu, Zeyu, et al.
Veröffentlicht: (2026)
von: Liu, Zeyu, et al.
Veröffentlicht: (2026)
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
Skill-Conditioned Visual Geolocation for Vision-Language Models
von: Yang, Chenjie, et al.
Veröffentlicht: (2026)
von: Yang, Chenjie, et al.
Veröffentlicht: (2026)
Generative Visual Communication in the Era of Vision-Language Models
von: Vinker, Yael
Veröffentlicht: (2024)
von: Vinker, Yael
Veröffentlicht: (2024)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
von: Tao, Hongyuan, et al.
Veröffentlicht: (2025)
von: Tao, Hongyuan, et al.
Veröffentlicht: (2025)
Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
von: Chae, Hyunsik, et al.
Veröffentlicht: (2025)
von: Chae, Hyunsik, et al.
Veröffentlicht: (2025)
Decoding Vision Transformers: the Diffusion Steering Lens
von: Takatsuki, Ryota, et al.
Veröffentlicht: (2025)
von: Takatsuki, Ryota, et al.
Veröffentlicht: (2025)
Fine-Tuning Vision-Language Models for Visual Navigation Assistance
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
The Effects of Visual Priming on Cooperative Behavior in Vision-Language Models
von: Ong, Kenneth J. K.
Veröffentlicht: (2026)
von: Ong, Kenneth J. K.
Veröffentlicht: (2026)
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
von: Song, Python, et al.
Veröffentlicht: (2025)
von: Song, Python, et al.
Veröffentlicht: (2025)
TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
von: YU, Mark, et al.
Veröffentlicht: (2025)
von: YU, Mark, et al.
Veröffentlicht: (2025)
Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language Models
von: Babaiee, Zahra, et al.
Veröffentlicht: (2025)
von: Babaiee, Zahra, et al.
Veröffentlicht: (2025)
When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations
von: Lalai, Harsh Nishant, et al.
Veröffentlicht: (2026)
von: Lalai, Harsh Nishant, et al.
Veröffentlicht: (2026)
Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering
von: Liu, Shuliang, et al.
Veröffentlicht: (2026)
von: Liu, Shuliang, et al.
Veröffentlicht: (2026)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
von: Kim, Sohee, et al.
Veröffentlicht: (2025)
von: Kim, Sohee, et al.
Veröffentlicht: (2025)
Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models
von: Góral, Gracjan, et al.
Veröffentlicht: (2025)
von: Góral, Gracjan, et al.
Veröffentlicht: (2025)
Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
von: Ernhofer, Benjamin Raphael, et al.
Veröffentlicht: (2025)
von: Ernhofer, Benjamin Raphael, et al.
Veröffentlicht: (2025)
QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models
von: Zhang, Yudong, et al.
Veröffentlicht: (2025)
von: Zhang, Yudong, et al.
Veröffentlicht: (2025)
Medical Large Vision Language Models with Multi-Image Visual Ability
von: Yang, Xikai, et al.
Veröffentlicht: (2025)
von: Yang, Xikai, et al.
Veröffentlicht: (2025)
Adapting Lightweight Vision Language Models for Radiological Visual Question Answering
von: Shourya, Aditya, et al.
Veröffentlicht: (2025)
von: Shourya, Aditya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
von: Balakrishnan, Ravikumar, et al.
Veröffentlicht: (2025) -
VISOR: VIsual Spatial Object Reasoning for Language-driven Object Navigation
von: Taioli, Francesco, et al.
Veröffentlicht: (2026) -
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
von: Shen, Yucheng, et al.
Veröffentlicht: (2026) -
Language Models Can Explain Visual Features via Steering
von: Ferrando, Javier, et al.
Veröffentlicht: (2026) -
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
von: Zhang, Yuanhong, et al.
Veröffentlicht: (2026)