Visual Representations inside the Language Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Benlin, Kamath, Amita, Grunde-McLaughlin, Madeleine, Han, Winson, Krishna, Ranjay |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Hard Positive Truth about Vision-Language Compositionality
di: Kamath, Amita, et al.
Pubblicazione: (2024)
di: Kamath, Amita, et al.
Pubblicazione: (2024)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
di: Kamath, Amita, et al.
Pubblicazione: (2026)
di: Kamath, Amita, et al.
Pubblicazione: (2026)
Seeking and Updating with Live Visual Knowledge
di: Fu, Mingyang, et al.
Pubblicazione: (2025)
di: Fu, Mingyang, et al.
Pubblicazione: (2025)
Designing LLM Chains by Adapting Techniques from Crowdsourcing Workflows
di: Grunde-McLaughlin, Madeleine, et al.
Pubblicazione: (2023)
di: Grunde-McLaughlin, Madeleine, et al.
Pubblicazione: (2023)
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
di: Hu, Yushi, et al.
Pubblicazione: (2024)
di: Hu, Yushi, et al.
Pubblicazione: (2024)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
di: Hu, Yushi, et al.
Pubblicazione: (2023)
di: Hu, Yushi, et al.
Pubblicazione: (2023)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
di: Su, Xia, et al.
Pubblicazione: (2026)
di: Su, Xia, et al.
Pubblicazione: (2026)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
di: Maharana, Adyasha, et al.
Pubblicazione: (2023)
di: Maharana, Adyasha, et al.
Pubblicazione: (2023)
Matryoshka Query Transformer for Large Vision-Language Models
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
di: Kamath, Amita, et al.
Pubblicazione: (2025)
di: Kamath, Amita, et al.
Pubblicazione: (2025)
Reinforced Visual Perception with Tools
di: Zhou, Zetong, et al.
Pubblicazione: (2025)
di: Zhou, Zetong, et al.
Pubblicazione: (2025)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
di: Che, Liwei, et al.
Pubblicazione: (2026)
di: Che, Liwei, et al.
Pubblicazione: (2026)
Efficient Inference of Vision Instruction-Following Models with Elastic Cache
di: Liu, Zuyan, et al.
Pubblicazione: (2024)
di: Liu, Zuyan, et al.
Pubblicazione: (2024)
Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
di: Fei, Yang, et al.
Pubblicazione: (2025)
di: Fei, Yang, et al.
Pubblicazione: (2025)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
di: Ma, Zixian, et al.
Pubblicazione: (2024)
di: Ma, Zixian, et al.
Pubblicazione: (2024)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
di: Li, Baiqi, et al.
Pubblicazione: (2024)
di: Li, Baiqi, et al.
Pubblicazione: (2024)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
di: Yang, Yue, et al.
Pubblicazione: (2023)
di: Yang, Yue, et al.
Pubblicazione: (2023)
Video-Based Reward Modeling for Computer-Use Agents
di: Song, Linxin, et al.
Pubblicazione: (2026)
di: Song, Linxin, et al.
Pubblicazione: (2026)
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
di: Li, Shaoxuan, et al.
Pubblicazione: (2026)
di: Li, Shaoxuan, et al.
Pubblicazione: (2026)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
Refining Skewed Perceptions in Vision-Language Contrastive Models through Visual Representations
di: Dai, Haocheng, et al.
Pubblicazione: (2024)
di: Dai, Haocheng, et al.
Pubblicazione: (2024)
Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
di: Liu, Benlin, et al.
Pubblicazione: (2024)
di: Liu, Benlin, et al.
Pubblicazione: (2024)
Semantic and Expressive Variation in Image Captions Across Languages
di: Ye, Andre, et al.
Pubblicazione: (2023)
di: Ye, Andre, et al.
Pubblicazione: (2023)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
di: Liu, Fanfan, et al.
Pubblicazione: (2024)
di: Liu, Fanfan, et al.
Pubblicazione: (2024)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
di: Hsieh, Cheng-Yu, et al.
Pubblicazione: (2025)
di: Hsieh, Cheng-Yu, et al.
Pubblicazione: (2025)
BLINK: Multimodal Large Language Models Can See but Not Perceive
di: Fu, Xingyu, et al.
Pubblicazione: (2024)
di: Fu, Xingyu, et al.
Pubblicazione: (2024)
Iterated Learning Improves Compositionality in Large Vision-Language Models
di: Zheng, Chenhao, et al.
Pubblicazione: (2024)
di: Zheng, Chenhao, et al.
Pubblicazione: (2024)
VividMed: Vision Language Model with Versatile Visual Grounding for Medicine
di: Luo, Lingxiao, et al.
Pubblicazione: (2024)
di: Luo, Lingxiao, et al.
Pubblicazione: (2024)
Posterior Augmented Flow Matching
di: Stoica, George, et al.
Pubblicazione: (2026)
di: Stoica, George, et al.
Pubblicazione: (2026)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
di: Song, Mingyang, et al.
Pubblicazione: (2026)
di: Song, Mingyang, et al.
Pubblicazione: (2026)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
di: Li, Zhuowan, et al.
Pubblicazione: (2022)
di: Li, Zhuowan, et al.
Pubblicazione: (2022)
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
Exploring the Design Space of Visual Context Representation in Video MLLMs
di: Du, Yifan, et al.
Pubblicazione: (2024)
di: Du, Yifan, et al.
Pubblicazione: (2024)
Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings
di: Huemann, Zachary, et al.
Pubblicazione: (2025)
di: Huemann, Zachary, et al.
Pubblicazione: (2025)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
di: Chen, Yukang, et al.
Pubblicazione: (2024)
di: Chen, Yukang, et al.
Pubblicazione: (2024)
Do Vision-Language Models Really Understand Visual Language?
di: Hou, Yifan, et al.
Pubblicazione: (2024)
di: Hou, Yifan, et al.
Pubblicazione: (2024)
SCoPE VLM: Selective Context Processing for Efficient Document Navigation in Vision-Language Models
di: Lim, Gyubeum, et al.
Pubblicazione: (2025)
di: Lim, Gyubeum, et al.
Pubblicazione: (2025)
Clean Evaluations on Contaminated Visual Language Models
di: Lu, Hongyuan, et al.
Pubblicazione: (2024)
di: Lu, Hongyuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Hard Positive Truth about Vision-Language Compositionality
di: Kamath, Amita, et al.
Pubblicazione: (2024) -
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
di: Kamath, Amita, et al.
Pubblicazione: (2026) -
Seeking and Updating with Live Visual Knowledge
di: Fu, Mingyang, et al.
Pubblicazione: (2025) -
Designing LLM Chains by Adapting Techniques from Crowdsourcing Workflows
di: Grunde-McLaughlin, Madeleine, et al.
Pubblicazione: (2023) -
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
di: Hu, Yushi, et al.
Pubblicazione: (2024)