LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Krojer, Benno, Nayak, Shravan, Mañas, Oscar, Adlakha, Vaibhav, Elliott, Desmond, Reddy, Siva, Mosbach, Marius |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Improving Automatic VQA Evaluation Using Large Language Models
di: Mañas, Oscar, et al.
Pubblicazione: (2023)
di: Mañas, Oscar, et al.
Pubblicazione: (2023)
Understanding the Influence of Synthetic Data for Text Embedders
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2025)
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2025)
Learning Action and Reasoning-Centric Image Editing from Videos and Simulations
di: Krojer, Benno, et al.
Pubblicazione: (2024)
di: Krojer, Benno, et al.
Pubblicazione: (2024)
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
di: BehnamGhader, Parishad, et al.
Pubblicazione: (2024)
di: BehnamGhader, Parishad, et al.
Pubblicazione: (2024)
LLM2Vec-Gen: Generative Embeddings from Large Language Models
di: BehnamGhader, Parishad, et al.
Pubblicazione: (2026)
di: BehnamGhader, Parishad, et al.
Pubblicazione: (2026)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
di: Li, Hongliang, et al.
Pubblicazione: (2025)
di: Li, Hongliang, et al.
Pubblicazione: (2025)
Forecasting Downstream Performance of LLMs With Proxy Metrics
di: Patel, Arkil, et al.
Pubblicazione: (2026)
di: Patel, Arkil, et al.
Pubblicazione: (2026)
Value Drifts: Tracing Value Alignment During LLM Post-Training
di: Bhatia, Mehar, et al.
Pubblicazione: (2025)
di: Bhatia, Mehar, et al.
Pubblicazione: (2025)
The Promise of RL for Autoregressive Image Editing
di: Ahmadi, Saba, et al.
Pubblicazione: (2025)
di: Ahmadi, Saba, et al.
Pubblicazione: (2025)
Not All Data Are Unlearned Equally
di: Krishnan, Aravind, et al.
Pubblicazione: (2025)
di: Krishnan, Aravind, et al.
Pubblicazione: (2025)
Benchmarking Vision Language Models for Cultural Understanding
di: Nayak, Shravan, et al.
Pubblicazione: (2024)
di: Nayak, Shravan, et al.
Pubblicazione: (2024)
Interpretability Needs a New Paradigm
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
di: Liu, Jinming, et al.
Pubblicazione: (2025)
di: Liu, Jinming, et al.
Pubblicazione: (2025)
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
di: Salazar, Israfel, et al.
Pubblicazione: (2025)
di: Salazar, Israfel, et al.
Pubblicazione: (2025)
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
di: Krojer, Benno, et al.
Pubblicazione: (2025)
di: Krojer, Benno, et al.
Pubblicazione: (2025)
Med-SegLens: Latent-Level Model Diffing for Interpretable Medical Image Segmentation
di: Ahmed, Salma J., et al.
Pubblicazione: (2026)
di: Ahmed, Salma J., et al.
Pubblicazione: (2026)
Seeing What Tastes Good: Revisiting Multimodal Distributional Semantics in the Billion Parameter Era
di: Oneata, Dan, et al.
Pubblicazione: (2025)
di: Oneata, Dan, et al.
Pubblicazione: (2025)
Multimodal LLM Augmented Reasoning for Interpretable Visual Perception Analysis
di: Chaudhari, Shravan, et al.
Pubblicazione: (2025)
di: Chaudhari, Shravan, et al.
Pubblicazione: (2025)
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
di: Sun, Zhichao, et al.
Pubblicazione: (2026)
di: Sun, Zhichao, et al.
Pubblicazione: (2026)
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
di: Wang, Dongsheng, et al.
Pubblicazione: (2024)
di: Wang, Dongsheng, et al.
Pubblicazione: (2024)
Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement
di: Nzoyem, Roussel Desmond, et al.
Pubblicazione: (2026)
di: Nzoyem, Roussel Desmond, et al.
Pubblicazione: (2026)
ULTra: Unveiling Latent Token Interpretability in Transformer-Based Understanding and Segmentation
di: Hosseini, Hesam, et al.
Pubblicazione: (2024)
di: Hosseini, Hesam, et al.
Pubblicazione: (2024)
Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
di: Yang, Zeyuan, et al.
Pubblicazione: (2025)
di: Yang, Zeyuan, et al.
Pubblicazione: (2025)
Efficient Test-Time Scaling for Small Vision-Language Models
di: Kaya, Mehmet Onurcan, et al.
Pubblicazione: (2025)
di: Kaya, Mehmet Onurcan, et al.
Pubblicazione: (2025)
RL-AD-Net: Reinforcement Learning Guided Adaptive Displacement in Latent Space for Refined Point Cloud Completion
di: Paregi, Bhanu Pratap, et al.
Pubblicazione: (2025)
di: Paregi, Bhanu Pratap, et al.
Pubblicazione: (2025)
Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs
di: Zheng, Yiren, et al.
Pubblicazione: (2026)
di: Zheng, Yiren, et al.
Pubblicazione: (2026)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
di: Villegas, Danae Sánchez, et al.
Pubblicazione: (2025)
di: Villegas, Danae Sánchez, et al.
Pubblicazione: (2025)
Latent Denoising Makes Good Tokenizers
di: Yang, Jiawei, et al.
Pubblicazione: (2025)
di: Yang, Jiawei, et al.
Pubblicazione: (2025)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
di: Li, Yan, et al.
Pubblicazione: (2026)
di: Li, Yan, et al.
Pubblicazione: (2026)
FeatureLens: A Highly Generalizable and Interpretable Framework for Detecting Adversarial Examples Based on Image Features
di: Yang, Zhigang, et al.
Pubblicazione: (2025)
di: Yang, Zhigang, et al.
Pubblicazione: (2025)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
di: Zhu, Jiaying, et al.
Pubblicazione: (2025)
di: Zhu, Jiaying, et al.
Pubblicazione: (2025)
EchoingPixels: Cross-Modal Adaptive Token Reduction for Efficient Audio-Visual LLMs
di: Gong, Chao, et al.
Pubblicazione: (2025)
di: Gong, Chao, et al.
Pubblicazione: (2025)
VisualLens: Personalization through Task-Agnostic Visual History
di: Zhu, Wang Bill, et al.
Pubblicazione: (2024)
di: Zhu, Wang Bill, et al.
Pubblicazione: (2024)
Build the web for agents, not agents for the web
di: Lù, Xing Han, et al.
Pubblicazione: (2025)
di: Lù, Xing Han, et al.
Pubblicazione: (2025)
Token Activation Map to Visually Explain Multimodal LLMs
di: Li, Yi, et al.
Pubblicazione: (2025)
di: Li, Yi, et al.
Pubblicazione: (2025)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2025)
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2025)
What explains the success of cross-modal fine-tuning with ORCA?
di: García-de-Herreros, Paloma, et al.
Pubblicazione: (2024)
di: García-de-Herreros, Paloma, et al.
Pubblicazione: (2024)
Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens
di: Xie, Qingsong, et al.
Pubblicazione: (2025)
di: Xie, Qingsong, et al.
Pubblicazione: (2025)
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
Discovering Failure Modes in Vision-Language Models using RL
di: Jain, Kanishk, et al.
Pubblicazione: (2026)
di: Jain, Kanishk, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Improving Automatic VQA Evaluation Using Large Language Models
di: Mañas, Oscar, et al.
Pubblicazione: (2023) -
Understanding the Influence of Synthetic Data for Text Embedders
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2025) -
Learning Action and Reasoning-Centric Image Editing from Videos and Simulations
di: Krojer, Benno, et al.
Pubblicazione: (2024) -
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
di: BehnamGhader, Parishad, et al.
Pubblicazione: (2024) -
LLM2Vec-Gen: Generative Embeddings from Large Language Models
di: BehnamGhader, Parishad, et al.
Pubblicazione: (2026)