Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Dafnis, Konstantinos M., Metaxas, Dimitris N. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
di: Sui, Elaine, et al.
Pubblicazione: (2024)
di: Sui, Elaine, et al.
Pubblicazione: (2024)
Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
di: Gu, Difei, et al.
Pubblicazione: (2025)
di: Gu, Difei, et al.
Pubblicazione: (2025)
How to Trace Latent Generative Model Generated Images without Artificial Watermark?
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
di: Li, Zhuowei, et al.
Pubblicazione: (2025)
RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
di: Gu, Difei, et al.
Pubblicazione: (2025)
di: Gu, Difei, et al.
Pubblicazione: (2025)
Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2025)
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2025)
Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection
di: Aqeel, Muhammad, et al.
Pubblicazione: (2026)
di: Aqeel, Muhammad, et al.
Pubblicazione: (2026)
Reducing Hallucinations in Vision-Language Models via Latent Space Steering
di: Liu, Sheng, et al.
Pubblicazione: (2024)
di: Liu, Sheng, et al.
Pubblicazione: (2024)
Fusing Domain-Specific Content from Large Language Models into Knowledge Graphs for Enhanced Zero Shot Object State Classification
di: Gouidis, Filippos, et al.
Pubblicazione: (2024)
di: Gouidis, Filippos, et al.
Pubblicazione: (2024)
AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models
di: Metzen, Jan Hendrik, et al.
Pubblicazione: (2023)
di: Metzen, Jan Hendrik, et al.
Pubblicazione: (2023)
Steering Rectified Flow Models in the Vector Field for Controlled Image Generation
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
di: Yang, Jiaxi, et al.
Pubblicazione: (2026)
di: Yang, Jiaxi, et al.
Pubblicazione: (2026)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization
di: Nguyen, Son, et al.
Pubblicazione: (2025)
di: Nguyen, Son, et al.
Pubblicazione: (2025)
ChatGPT Encounters Morphing Attack Detection: Zero-Shot MAD with Multi-Modal Large Language Models and General Vision Models
di: Zhang, Haoyu, et al.
Pubblicazione: (2025)
di: Zhang, Haoyu, et al.
Pubblicazione: (2025)
Visual Language Models as Zero-Shot Deepfake Detectors
di: Pirogov, Viacheslav
Pubblicazione: (2025)
di: Pirogov, Viacheslav
Pubblicazione: (2025)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
di: Nagar, Aishik, et al.
Pubblicazione: (2024)
di: Nagar, Aishik, et al.
Pubblicazione: (2024)
VisionTS: Visual Masked Autoencoders Are Free-Lunch Zero-Shot Time Series Forecasters
di: Chen, Mouxiang, et al.
Pubblicazione: (2024)
di: Chen, Mouxiang, et al.
Pubblicazione: (2024)
ATHENA: Adaptive Test-Time Steering for Improving Count Fidelity in Diffusion Models
di: Sepehri, Mohammad Shahab, et al.
Pubblicazione: (2026)
di: Sepehri, Mohammad Shahab, et al.
Pubblicazione: (2026)
Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models
di: Song, Fei, et al.
Pubblicazione: (2025)
di: Song, Fei, et al.
Pubblicazione: (2025)
Zero-Shot Generalization of Vision-Based RL Without Data Augmentation
di: Batra, Sumeet, et al.
Pubblicazione: (2024)
di: Batra, Sumeet, et al.
Pubblicazione: (2024)
Steering Sparse Autoencoder Latents to Control Dynamic Head Pruning in Vision Transformers (Student Abstract)
di: Lee, Yousung, et al.
Pubblicazione: (2026)
di: Lee, Yousung, et al.
Pubblicazione: (2026)
VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
di: Yu, Xinlei, et al.
Pubblicazione: (2025)
di: Yu, Xinlei, et al.
Pubblicazione: (2025)
Zero-Shot Action Generalization with Limited Observations
di: Alchihabi, Abdullah, et al.
Pubblicazione: (2025)
di: Alchihabi, Abdullah, et al.
Pubblicazione: (2025)
Vision-Language Models Unlock Task-Centric Latent Actions
di: Nikulin, Alexander, et al.
Pubblicazione: (2026)
di: Nikulin, Alexander, et al.
Pubblicazione: (2026)
Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model
di: Ma, Huan, et al.
Pubblicazione: (2024)
di: Ma, Huan, et al.
Pubblicazione: (2024)
Language-Driven Anchors for Zero-Shot Adversarial Robustness
di: Li, Xiao, et al.
Pubblicazione: (2023)
di: Li, Xiao, et al.
Pubblicazione: (2023)
Efficient Few-Shot Learning in Remote Sensing: Fusing Vision and Vision-Language Models
di: Chua, Jia Yun, et al.
Pubblicazione: (2025)
di: Chua, Jia Yun, et al.
Pubblicazione: (2025)
Entropy-Aware Structural Alignment for Zero-Shot Handwritten Chinese Character Recognition
di: Luo, Qiuming, et al.
Pubblicazione: (2026)
di: Luo, Qiuming, et al.
Pubblicazione: (2026)
ProtoCLIP: Prototype-Aligned Latent Refinement for Robust Zero-Shot Chest X-Ray Classification
di: Kittler, Florian, et al.
Pubblicazione: (2026)
di: Kittler, Florian, et al.
Pubblicazione: (2026)
MotionCraft: Physics-based Zero-Shot Video Generation
di: Aira, Luca Savant, et al.
Pubblicazione: (2024)
di: Aira, Luca Savant, et al.
Pubblicazione: (2024)
AgentDrug: Utilizing Large Language Models in An Agentic Workflow for Zero-Shot Molecular Editing
di: Le, Khiem, et al.
Pubblicazione: (2024)
di: Le, Khiem, et al.
Pubblicazione: (2024)
PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting
di: Maqbool, Danyal, et al.
Pubblicazione: (2025)
di: Maqbool, Danyal, et al.
Pubblicazione: (2025)
A Recipe for Improving Remote Sensing VLM Zero Shot Generalization
di: Barzilai, Aviad, et al.
Pubblicazione: (2025)
di: Barzilai, Aviad, et al.
Pubblicazione: (2025)
Bayesian Modeling of Zero-Shot Classifications for Urban Flood Detection
di: Franchi, Matt, et al.
Pubblicazione: (2025)
di: Franchi, Matt, et al.
Pubblicazione: (2025)
Zero-Shot Decentralized Federated Learning
di: Masano, Alessio, et al.
Pubblicazione: (2025)
di: Masano, Alessio, et al.
Pubblicazione: (2025)
Synthetic Data is Sufficient for Zero-Shot Visual Generalization from Offline Data
di: Güzel, Ahmet H., et al.
Pubblicazione: (2025)
di: Güzel, Ahmet H., et al.
Pubblicazione: (2025)
Spectrum-Aware Parameter Efficient Fine-Tuning for Diffusion Models
di: Zhang, Xinxi, et al.
Pubblicazione: (2024)
di: Zhang, Xinxi, et al.
Pubblicazione: (2024)
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
di: Hinojosa, Carlos, et al.
Pubblicazione: (2026)
di: Hinojosa, Carlos, et al.
Pubblicazione: (2026)
ProKeR: A Kernel Perspective on Few-Shot Adaptation of Large Vision-Language Models
di: Bendou, Yassir, et al.
Pubblicazione: (2025)
di: Bendou, Yassir, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
di: Sui, Elaine, et al.
Pubblicazione: (2024) -
Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
di: Gu, Difei, et al.
Pubblicazione: (2025) -
How to Trace Latent Generative Model Generated Images without Artificial Watermark?
di: Wang, Zhenting, et al.
Pubblicazione: (2024) -
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
di: Li, Zhuowei, et al.
Pubblicazione: (2025) -
RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
di: Gu, Difei, et al.
Pubblicazione: (2025)