Interpreting the linear structure of vision-language model embedding spaces
Fuente:
arXiv
Saved in:
| Main Authors: | Papadimitriou, Isabel, Su, Huangyuan, Fel, Thomas, Kakade, Sham, Gil, Stephanie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
by: Shih, Yu-Fei, et al.
Published: (2025)
by: Shih, Yu-Fei, et al.
Published: (2025)
Mano Technical Report
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
by: Wang, Bing, et al.
Published: (2025)
by: Wang, Bing, et al.
Published: (2025)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos
by: Sukhani, Siddhant, et al.
Published: (2025)
by: Sukhani, Siddhant, et al.
Published: (2025)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
by: Wang, Kangsheng, et al.
Published: (2025)
by: Wang, Kangsheng, et al.
Published: (2025)
Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model
by: Cuong, Dinh Viet, et al.
Published: (2025)
by: Cuong, Dinh Viet, et al.
Published: (2025)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
by: Jiang, Jingjing, et al.
Published: (2025)
by: Jiang, Jingjing, et al.
Published: (2025)
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
by: Fazli, Mehrdad, et al.
Published: (2025)
by: Fazli, Mehrdad, et al.
Published: (2025)
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
by: An, Wenbin, et al.
Published: (2025)
by: An, Wenbin, et al.
Published: (2025)
ChronusOmni: Improving Time Awareness of Omni Large Language Models
by: Chen, Yijing, et al.
Published: (2025)
by: Chen, Yijing, et al.
Published: (2025)
HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
by: Yang, Yiqing, et al.
Published: (2025)
by: Yang, Yiqing, et al.
Published: (2025)
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
by: Ghaleb, Esam, et al.
Published: (2025)
by: Ghaleb, Esam, et al.
Published: (2025)
Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
by: Stamatakis, Markos, et al.
Published: (2025)
by: Stamatakis, Markos, et al.
Published: (2025)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
by: Wu, Jiaying, et al.
Published: (2025)
by: Wu, Jiaying, et al.
Published: (2025)
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
by: Gan, Chengguang, et al.
Published: (2025)
by: Gan, Chengguang, et al.
Published: (2025)
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
by: An, Zhaoyi, et al.
Published: (2025)
by: An, Zhaoyi, et al.
Published: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
by: Zhang, Xueqiao, et al.
Published: (2025)
by: Zhang, Xueqiao, et al.
Published: (2025)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
by: Masumura, Ryo, et al.
Published: (2025)
by: Masumura, Ryo, et al.
Published: (2025)
How Far Are We from Generating Missing Modalities with Foundation Models?
by: Ke, Guanzhou, et al.
Published: (2025)
by: Ke, Guanzhou, et al.
Published: (2025)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
by: Zhao, Zhixian, et al.
Published: (2026)
by: Zhao, Zhixian, et al.
Published: (2026)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
by: Jiang, Chaoya, et al.
Published: (2024)
by: Jiang, Chaoya, et al.
Published: (2024)
LocoMotion: Learning Motion-Focused Video-Language Representations
by: Doughty, Hazel, et al.
Published: (2024)
by: Doughty, Hazel, et al.
Published: (2024)
DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling
by: Lan, Jing, et al.
Published: (2026)
by: Lan, Jing, et al.
Published: (2026)
Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection
by: Nishimura, Taichi, et al.
Published: (2024)
by: Nishimura, Taichi, et al.
Published: (2024)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
by: Li, Jinyuan, et al.
Published: (2024)
by: Li, Jinyuan, et al.
Published: (2024)
BioVL-QR: Egocentric Biochemical Vision-and-Language Dataset Using Micro QR Codes
by: Nishimoto, Tomohiro, et al.
Published: (2024)
by: Nishimoto, Tomohiro, et al.
Published: (2024)
MLANet: Multi-Level Attention Network with Sub-instruction for Continuous Vision-and-Language Navigation
by: He, Zongtao, et al.
Published: (2023)
by: He, Zongtao, et al.
Published: (2023)
DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter
by: Dong, Ziyi, et al.
Published: (2022)
by: Dong, Ziyi, et al.
Published: (2022)
Recipe Generation from Unsegmented Cooking Videos
by: Nishimura, Taichi, et al.
Published: (2022)
by: Nishimura, Taichi, et al.
Published: (2022)
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
by: Xu, Jiaqi, et al.
Published: (2023)
by: Xu, Jiaqi, et al.
Published: (2023)
M$^3$Face: A Unified Multi-Modal Multilingual Framework for Human Face Generation and Editing
by: Mofayezi, Mohammadreza, et al.
Published: (2024)
by: Mofayezi, Mohammadreza, et al.
Published: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
by: Wang, Jiapeng, et al.
Published: (2024)
by: Wang, Jiapeng, et al.
Published: (2024)
See or Guess: Counterfactually Regularized Image Captioning
by: Cao, Qian, et al.
Published: (2024)
by: Cao, Qian, et al.
Published: (2024)
Multi-task Prompt Words Learning for Social Media Content Generation
by: Xue, Haochen, et al.
Published: (2024)
by: Xue, Haochen, et al.
Published: (2024)
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
by: Maeda, Koki, et al.
Published: (2024)
by: Maeda, Koki, et al.
Published: (2024)
Similar Items
-
Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions
by: Zhang, Xiao, et al.
Published: (2025) -
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
by: Wang, Xiao, et al.
Published: (2025) -
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
by: Shih, Yu-Fei, et al.
Published: (2025) -
Mano Technical Report
by: Fu, Tianyu, et al.
Published: (2025) -
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
by: Wang, Bing, et al.
Published: (2025)