Bridging Hidden States in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fein-Ashley, Benjamin, Fein-Ashley, Jacob |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion Models with Anisotropic Gaussian Splatting for Image Inpainting
by: Fein-Ashley, Jacob, et al.
Published: (2024)
by: Fein-Ashley, Jacob, et al.
Published: (2024)
HVT: A Comprehensive Vision Framework for Learning in Non-Euclidean Space
by: Fein-Ashley, Jacob, et al.
Published: (2024)
by: Fein-Ashley, Jacob, et al.
Published: (2024)
Studying the Effects of Self-Attention on SAR Automatic Target Recognition
by: Fein-Ashley, Jacob, et al.
Published: (2024)
by: Fein-Ashley, Jacob, et al.
Published: (2024)
Latent Denoising Improves Visual Alignment in Large Multimodal Models
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
A Single Graph Convolution Is All You Need: Efficient Grayscale Image Classification
by: Fein-Ashley, Jacob, et al.
Published: (2024)
by: Fein-Ashley, Jacob, et al.
Published: (2024)
Benchmarking Deep Learning Classifiers for SAR Automatic Target Recognition
by: Fein-Ashley, Jacob, et al.
Published: (2023)
by: Fein-Ashley, Jacob, et al.
Published: (2023)
A Comparison of Traditional and Deep Learning Methods for Parameter Estimation of the Ornstein-Uhlenbeck Process
by: Fein-Ashley, Jacob
Published: (2024)
by: Fein-Ashley, Jacob
Published: (2024)
Solve the Loop: Attractor Models for Language and Reasoning
by: Fein-Ashley, Jacob, et al.
Published: (2026)
by: Fein-Ashley, Jacob, et al.
Published: (2026)
Iterate to Accelerate: A Unified Framework for Iterative Reasoning and Feedback Convergence
by: Fein-Ashley, Jacob
Published: (2025)
by: Fein-Ashley, Jacob
Published: (2025)
Flowing Through Layers: A Continuous Dynamical Systems Perspective on Transformers
by: Fein-Ashley, Jacob
Published: (2025)
by: Fein-Ashley, Jacob
Published: (2025)
Linear Diffusion Networks
by: Fein-Ashley, Jacob
Published: (2025)
by: Fein-Ashley, Jacob
Published: (2025)
LouvreSAE: Sparse Autoencoders for Interpretable and Controllable Style Transfer
by: Panda, Raina, et al.
Published: (2025)
by: Panda, Raina, et al.
Published: (2025)
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
by: Truong, Thanh-Dat, et al.
Published: (2025)
by: Truong, Thanh-Dat, et al.
Published: (2025)
Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models
by: Sharma, Pranav, et al.
Published: (2025)
by: Sharma, Pranav, et al.
Published: (2025)
Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
by: Hemmat, Arshia, et al.
Published: (2024)
by: Hemmat, Arshia, et al.
Published: (2024)
Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives
by: Park, Ji-jun, et al.
Published: (2024)
by: Park, Ji-jun, et al.
Published: (2024)
HARMONY: Hidden Activation Representations and Model Output-Aware Uncertainty Estimation for Vision-Language Models
by: Mushtaq, Erum, et al.
Published: (2025)
by: Mushtaq, Erum, et al.
Published: (2025)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024)
by: Zhao, Shihao, et al.
Published: (2024)
PolarVLM: Bridging the Semantic-Physical Gap in Vision-Language Models
by: Li, Yuliang, et al.
Published: (2026)
by: Li, Yuliang, et al.
Published: (2026)
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2024)
by: Zhan, Yufei, et al.
Published: (2024)
Leveraging Vision-Language Foundation Models to Reveal Hidden Image-Attribute Relationships in Medical Imaging
by: Kumar, Amar, et al.
Published: (2025)
by: Kumar, Amar, et al.
Published: (2025)
TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos
by: Yu, Yakun, et al.
Published: (2026)
by: Yu, Yakun, et al.
Published: (2026)
EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic Segmentation
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
by: Zhang, Mingyuan, et al.
Published: (2025)
by: Zhang, Mingyuan, et al.
Published: (2025)
UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition
by: Zhang, Zhenrong, et al.
Published: (2024)
by: Zhang, Zhenrong, et al.
Published: (2024)
The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?
by: Zhao, Qinyu, et al.
Published: (2024)
by: Zhao, Qinyu, et al.
Published: (2024)
WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
by: Yang, Jian, et al.
Published: (2025)
by: Yang, Jian, et al.
Published: (2025)
ParGo: Bridging Vision-Language with Partial and Global Views
by: Wang, An-Lan, et al.
Published: (2024)
by: Wang, An-Lan, et al.
Published: (2024)
Optimizing Vision-Language Interactions Through Decoder-Only Models
by: Tanaka, Kaito, et al.
Published: (2024)
by: Tanaka, Kaito, et al.
Published: (2024)
Hidden Clones: Exposing and Fixing Family Bias in Vision-Language Model Ensembles
by: Bugaud, Zacharie
Published: (2026)
by: Bugaud, Zacharie
Published: (2026)
Selective State Space Memory for Large Vision-Language Models
by: Ng, Chee, et al.
Published: (2024)
by: Ng, Chee, et al.
Published: (2024)
Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging
by: Li, Yitong, et al.
Published: (2025)
by: Li, Yitong, et al.
Published: (2025)
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
by: Han, Yuhang, et al.
Published: (2026)
by: Han, Yuhang, et al.
Published: (2026)
StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning
by: Sun, Xiaowen, et al.
Published: (2026)
by: Sun, Xiaowen, et al.
Published: (2026)
Visual Cues of Gender and Race are Associated with Stereotyping in Vision-Language Models
by: Lee, Messi H. J., et al.
Published: (2025)
by: Lee, Messi H. J., et al.
Published: (2025)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
Bridge the Modality and Capability Gaps in Vision-Language Model Selection
by: Yi, Chao, et al.
Published: (2024)
by: Yi, Chao, et al.
Published: (2024)
Application Of Vision-Language Models For Assessing Osteoarthritis Disease Severity
by: Felfeliyan, Banafshe, et al.
Published: (2024)
by: Felfeliyan, Banafshe, et al.
Published: (2024)
Similar Items
-
Diffusion Models with Anisotropic Gaussian Splatting for Image Inpainting
by: Fein-Ashley, Jacob, et al.
Published: (2024) -
HVT: A Comprehensive Vision Framework for Learning in Non-Euclidean Space
by: Fein-Ashley, Jacob, et al.
Published: (2024) -
Studying the Effects of Self-Attention on SAR Automatic Target Recognition
by: Fein-Ashley, Jacob, et al.
Published: (2024) -
Latent Denoising Improves Visual Alignment in Large Multimodal Models
by: Parikh, Dhruv, et al.
Published: (2026) -
ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning
by: Parikh, Dhruv, et al.
Published: (2025)