VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Phute, Mansi, Balakrishnan, Ravikumar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
di: Balakrishnan, Ravikumar, et al.
Pubblicazione: (2025)
di: Balakrishnan, Ravikumar, et al.
Pubblicazione: (2025)
VISOR: VIsual Spatial Object Reasoning for Language-driven Object Navigation
di: Taioli, Francesco, et al.
Pubblicazione: (2026)
di: Taioli, Francesco, et al.
Pubblicazione: (2026)
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
di: Shen, Yucheng, et al.
Pubblicazione: (2026)
di: Shen, Yucheng, et al.
Pubblicazione: (2026)
Language Models Can Explain Visual Features via Steering
di: Ferrando, Javier, et al.
Pubblicazione: (2026)
di: Ferrando, Javier, et al.
Pubblicazione: (2026)
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
di: Zhang, Yuanhong, et al.
Pubblicazione: (2026)
di: Zhang, Yuanhong, et al.
Pubblicazione: (2026)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
SynthVision -- Harnessing Minimal Input for Maximal Output in Computer Vision Models using Synthetic Image data
di: Kularathne, Yudara, et al.
Pubblicazione: (2024)
di: Kularathne, Yudara, et al.
Pubblicazione: (2024)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
di: Yang, Jiaxi, et al.
Pubblicazione: (2026)
di: Yang, Jiaxi, et al.
Pubblicazione: (2026)
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
di: Wu, Sihao, et al.
Pubblicazione: (2025)
di: Wu, Sihao, et al.
Pubblicazione: (2025)
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
di: Rashid, Md Rafi Ur, et al.
Pubblicazione: (2026)
di: Rashid, Md Rafi Ur, et al.
Pubblicazione: (2026)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
di: Zou, Zhengtao, et al.
Pubblicazione: (2025)
di: Zou, Zhengtao, et al.
Pubblicazione: (2025)
Mitigating Hallucination in Vision-Language Models through Barrier-Regulated Adaptive Closed-form Steering
di: Jana, Soumyadeep, et al.
Pubblicazione: (2026)
di: Jana, Soumyadeep, et al.
Pubblicazione: (2026)
Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
Single-Input Multi-Output Model Merging: Leveraging Foundation Models for Dense Multi-Task Learning
di: Giraldo, Juan Garcia, et al.
Pubblicazione: (2025)
di: Giraldo, Juan Garcia, et al.
Pubblicazione: (2025)
3D Gaussian and Diffusion-Based Gaze Redirection
di: Panchalingam, Abiram, et al.
Pubblicazione: (2025)
di: Panchalingam, Abiram, et al.
Pubblicazione: (2025)
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
VFM-VLM: Vision Foundation Model and Vision Language Model based Visual Comparison for 3D Pose Estimation
di: Sarowar, Md Selim, et al.
Pubblicazione: (2025)
di: Sarowar, Md Selim, et al.
Pubblicazione: (2025)
Learning to Steer: Input-dependent Steering for Multimodal LLMs
di: Parekh, Jayneel, et al.
Pubblicazione: (2025)
di: Parekh, Jayneel, et al.
Pubblicazione: (2025)
Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models
di: Balakrishnan, Ravikumar, et al.
Pubblicazione: (2026)
di: Balakrishnan, Ravikumar, et al.
Pubblicazione: (2026)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
di: Liu, Zeyu, et al.
Pubblicazione: (2026)
di: Liu, Zeyu, et al.
Pubblicazione: (2026)
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
di: Maeda, Koki, et al.
Pubblicazione: (2024)
di: Maeda, Koki, et al.
Pubblicazione: (2024)
Skill-Conditioned Visual Geolocation for Vision-Language Models
di: Yang, Chenjie, et al.
Pubblicazione: (2026)
di: Yang, Chenjie, et al.
Pubblicazione: (2026)
Generative Visual Communication in the Era of Vision-Language Models
di: Vinker, Yael
Pubblicazione: (2024)
di: Vinker, Yael
Pubblicazione: (2024)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
di: Tao, Hongyuan, et al.
Pubblicazione: (2025)
di: Tao, Hongyuan, et al.
Pubblicazione: (2025)
Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
di: Chae, Hyunsik, et al.
Pubblicazione: (2025)
di: Chae, Hyunsik, et al.
Pubblicazione: (2025)
Decoding Vision Transformers: the Diffusion Steering Lens
di: Takatsuki, Ryota, et al.
Pubblicazione: (2025)
di: Takatsuki, Ryota, et al.
Pubblicazione: (2025)
Fine-Tuning Vision-Language Models for Visual Navigation Assistance
di: Li, Xiao, et al.
Pubblicazione: (2025)
di: Li, Xiao, et al.
Pubblicazione: (2025)
The Effects of Visual Priming on Cooperative Behavior in Vision-Language Models
di: Ong, Kenneth J. K.
Pubblicazione: (2026)
di: Ong, Kenneth J. K.
Pubblicazione: (2026)
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
di: Song, Python, et al.
Pubblicazione: (2025)
di: Song, Python, et al.
Pubblicazione: (2025)
TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
di: YU, Mark, et al.
Pubblicazione: (2025)
di: YU, Mark, et al.
Pubblicazione: (2025)
Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language Models
di: Babaiee, Zahra, et al.
Pubblicazione: (2025)
di: Babaiee, Zahra, et al.
Pubblicazione: (2025)
When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations
di: Lalai, Harsh Nishant, et al.
Pubblicazione: (2026)
di: Lalai, Harsh Nishant, et al.
Pubblicazione: (2026)
Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering
di: Liu, Shuliang, et al.
Pubblicazione: (2026)
di: Liu, Shuliang, et al.
Pubblicazione: (2026)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
di: Ren, Yiming, et al.
Pubblicazione: (2026)
di: Ren, Yiming, et al.
Pubblicazione: (2026)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
di: Kim, Sohee, et al.
Pubblicazione: (2025)
di: Kim, Sohee, et al.
Pubblicazione: (2025)
Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models
di: Góral, Gracjan, et al.
Pubblicazione: (2025)
di: Góral, Gracjan, et al.
Pubblicazione: (2025)
Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
di: Ernhofer, Benjamin Raphael, et al.
Pubblicazione: (2025)
di: Ernhofer, Benjamin Raphael, et al.
Pubblicazione: (2025)
QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models
di: Zhang, Yudong, et al.
Pubblicazione: (2025)
di: Zhang, Yudong, et al.
Pubblicazione: (2025)
Medical Large Vision Language Models with Multi-Image Visual Ability
di: Yang, Xikai, et al.
Pubblicazione: (2025)
di: Yang, Xikai, et al.
Pubblicazione: (2025)
Adapting Lightweight Vision Language Models for Radiological Visual Question Answering
di: Shourya, Aditya, et al.
Pubblicazione: (2025)
di: Shourya, Aditya, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
di: Balakrishnan, Ravikumar, et al.
Pubblicazione: (2025) -
VISOR: VIsual Spatial Object Reasoning for Language-driven Object Navigation
di: Taioli, Francesco, et al.
Pubblicazione: (2026) -
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
di: Shen, Yucheng, et al.
Pubblicazione: (2026) -
Language Models Can Explain Visual Features via Steering
di: Ferrando, Javier, et al.
Pubblicazione: (2026) -
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
di: Zhang, Yuanhong, et al.
Pubblicazione: (2026)