Line of Sight: On Linear Representations in VLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Rajaram, Achyuta, Schwettmann, Sarah, Andreas, Jacob, Conmy, Arthur |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatic Discovery of Visual Circuits
by: Rajaram, Achyuta, et al.
Published: (2024)
by: Rajaram, Achyuta, et al.
Published: (2024)
A Multimodal Automated Interpretability Agent
by: Shaham, Tamar Rott, et al.
Published: (2024)
by: Shaham, Tamar Rott, et al.
Published: (2024)
Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
by: Ali, Muhammad, et al.
Published: (2025)
by: Ali, Muhammad, et al.
Published: (2025)
Nearest Neighbor Normalization Improves Multimodal Retrieval
by: Chowdhury, Neil, et al.
Published: (2024)
by: Chowdhury, Neil, et al.
Published: (2024)
CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task Solvers
by: Mallis, Dimitrios, et al.
Published: (2024)
by: Mallis, Dimitrios, et al.
Published: (2024)
mmWave Radar-Based Non-Line-of-Sight Pedestrian Localization at T-Junctions Utilizing Road Layout Extraction via Camera
by: Park, Byeonggyu, et al.
Published: (2025)
by: Park, Byeonggyu, et al.
Published: (2025)
NIGHT -- Non-Line-of-Sight Imaging from Indirect Time of Flight Data
by: Caligiuri, Matteo, et al.
Published: (2024)
by: Caligiuri, Matteo, et al.
Published: (2024)
WebSight: A Vision-First Architecture for Robust Web Agents
by: Bhathal, Tanvir, et al.
Published: (2025)
by: Bhathal, Tanvir, et al.
Published: (2025)
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
by: Chen, Kaijin, et al.
Published: (2026)
by: Chen, Kaijin, et al.
Published: (2026)
Interpreting CLIP's Image Representation via Text-Based Decomposition
by: Gandelsman, Yossi, et al.
Published: (2023)
by: Gandelsman, Yossi, et al.
Published: (2023)
ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
OnSight Pathology: A real-time platform-agnostic computational pathology companion for histopathology
by: Hu, Jinzhen, et al.
Published: (2025)
by: Hu, Jinzhen, et al.
Published: (2025)
InkSight: Offline-to-Online Handwriting Conversion by Teaching Vision-Language Models to Read and Write
by: Mitrevski, Blagoj, et al.
Published: (2024)
by: Mitrevski, Blagoj, et al.
Published: (2024)
Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs
by: Zaazou, Youssef, et al.
Published: (2026)
by: Zaazou, Youssef, et al.
Published: (2026)
EndoSight AI: Deep Learning-Driven Real-Time Gastrointestinal Polyp Detection and Segmentation for Enhanced Endoscopic Diagnostics
by: Cavadia, Daniel
Published: (2025)
by: Cavadia, Daniel
Published: (2025)
FinSight-Net:A Physics-Aware Decoupled Network with Frequency-Domain Compensation for Underwater Fish Detection in Smart Aquaculture
by: Yang, Jinsong, et al.
Published: (2026)
by: Yang, Jinsong, et al.
Published: (2026)
CAPRMIL: Context-Aware Patch Representations for Multiple Instance Learning
by: Lolos, Andreas, et al.
Published: (2025)
by: Lolos, Andreas, et al.
Published: (2025)
From Linear Probing to Joint-Weighted Token Hierarchy: A Foundation Model Bridging Global and Cellular Representations in Biomarker Detection
by: Liu, Jingsong, et al.
Published: (2025)
by: Liu, Jingsong, et al.
Published: (2025)
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
by: Grigore, Diana-Nicoleta, et al.
Published: (2025)
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning
by: Gan, Rui, et al.
Published: (2026)
by: Gan, Rui, et al.
Published: (2026)
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
by: Zhao, Yaqi, et al.
Published: (2024)
by: Zhao, Yaqi, et al.
Published: (2024)
Hidden in Plain Sight: Undetectable Adversarial Bias Attacks on Vulnerable Patient Populations
by: Kulkarni, Pranav, et al.
Published: (2024)
by: Kulkarni, Pranav, et al.
Published: (2024)
DeepSight: An All-in-One LM Safety Toolkit
by: Zhang, Bo, et al.
Published: (2026)
by: Zhang, Bo, et al.
Published: (2026)
Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition
by: Wang, Zhifeng, et al.
Published: (2025)
by: Wang, Zhifeng, et al.
Published: (2025)
i-MAE: Are Latent Representations in Masked Autoencoders Linearly Separable?
by: Zhang, Kevin, et al.
Published: (2022)
by: Zhang, Kevin, et al.
Published: (2022)
ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision Models
by: Wei, Guoyizhe, et al.
Published: (2025)
by: Wei, Guoyizhe, et al.
Published: (2025)
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
by: Liao, Bencheng, et al.
Published: (2024)
by: Liao, Bencheng, et al.
Published: (2024)
No Alignment Needed for Generation: Learning Linearly Separable Representations in Diffusion Models
by: Yun, Junno, et al.
Published: (2025)
by: Yun, Junno, et al.
Published: (2025)
OOSTraj: Out-of-Sight Trajectory Prediction With Vision-Positioning Denoising
by: Zhang, Haichao, et al.
Published: (2024)
by: Zhang, Haichao, et al.
Published: (2024)
Steerable Visual Representations
by: Ruthardt, Jona, et al.
Published: (2026)
by: Ruthardt, Jona, et al.
Published: (2026)
Training-Free Representation Guidance for Diffusion Models with a Representation Alignment Projector
by: Zu, Wenqiang, et al.
Published: (2026)
by: Zu, Wenqiang, et al.
Published: (2026)
The Linear Attention Resurrection in Vision Transformer
by: Zheng, Chuanyang
Published: (2025)
by: Zheng, Chuanyang
Published: (2025)
Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions
by: Kang, Wan Ju, et al.
Published: (2025)
by: Kang, Wan Ju, et al.
Published: (2025)
I Walk the Line: Examining the Role of Gestalt Continuity in Object Binding for Vision Transformers
by: Tartaglini, Alexa R., et al.
Published: (2026)
by: Tartaglini, Alexa R., et al.
Published: (2026)
Driving scenario generation and evaluation using a structured layer representation and foundational models
by: Hubert, Arthur, et al.
Published: (2025)
by: Hubert, Arthur, et al.
Published: (2025)
On the Temporality for Sketch Representation Learning
by: Junior, Marcelo Isaias de Moraes, et al.
Published: (2025)
by: Junior, Marcelo Isaias de Moraes, et al.
Published: (2025)
Sparse Representation Learning for Vessels
by: Prabhakar, Chinmay, et al.
Published: (2026)
by: Prabhakar, Chinmay, et al.
Published: (2026)
Rethink MAE with Linear Time-Invariant Dynamics
by: Wang, Zice
Published: (2026)
by: Wang, Zice
Published: (2026)
Foreign-Object Detection in High-Voltage Transmission Line Based on Improved YOLOv8m
by: Wang, Zhenyue, et al.
Published: (2025)
by: Wang, Zhenyue, et al.
Published: (2025)
Similar Items
-
Automatic Discovery of Visual Circuits
by: Rajaram, Achyuta, et al.
Published: (2024) -
A Multimodal Automated Interpretability Agent
by: Shaham, Tamar Rott, et al.
Published: (2024) -
Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
by: Ali, Muhammad, et al.
Published: (2025) -
Nearest Neighbor Normalization Improves Multimodal Retrieval
by: Chowdhury, Neil, et al.
Published: (2024) -
CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task Solvers
by: Mallis, Dimitrios, et al.
Published: (2024)