Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bigverdi, Mahtab, Luo, Zelun, Hsieh, Cheng-Yu, Shen, Ethan, Chen, Dongping, Shapiro, Linda G., Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
von: Li, Linjie, et al.
Veröffentlicht: (2025)
von: Li, Linjie, et al.
Veröffentlicht: (2025)
MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2025)
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2025)
MIMIC: Masked Image Modeling with Image Correspondences
von: Marathe, Kalyani, et al.
Veröffentlicht: (2023)
von: Marathe, Kalyani, et al.
Veröffentlicht: (2023)
Data Alignment for Zero-Shot Concept Generation in Dermatology AI
von: Gadgil, Soham, et al.
Veröffentlicht: (2024)
von: Gadgil, Soham, et al.
Veröffentlicht: (2024)
Reinforced Visual Perception with Tools
von: Zhou, Zetong, et al.
Veröffentlicht: (2025)
von: Zhou, Zetong, et al.
Veröffentlicht: (2025)
The Hard Positive Truth about Vision-Language Compositionality
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
Introducing Visual Perception Token into Multimodal Large Language Model
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
von: Yu, Runpeng, et al.
Veröffentlicht: (2025)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
von: Seyfioglu, Mehmet Saygin, et al.
Veröffentlicht: (2023)
von: Seyfioglu, Mehmet Saygin, et al.
Veröffentlicht: (2023)
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
von: Ikezogwo, Wisdom O., et al.
Veröffentlicht: (2025)
von: Ikezogwo, Wisdom O., et al.
Veröffentlicht: (2025)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
von: Hu, Yushi, et al.
Veröffentlicht: (2023)
von: Hu, Yushi, et al.
Veröffentlicht: (2023)
Spotlight on Token Perception for Multimodal Reinforcement Learning
von: Huang, Siyuan, et al.
Veröffentlicht: (2025)
von: Huang, Siyuan, et al.
Veröffentlicht: (2025)
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
von: Hu, Yushi, et al.
Veröffentlicht: (2024)
von: Hu, Yushi, et al.
Veröffentlicht: (2024)
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
Visual Representations inside the Language Model
von: Liu, Benlin, et al.
Veröffentlicht: (2025)
von: Liu, Benlin, et al.
Veröffentlicht: (2025)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
Seeking and Updating with Live Visual Knowledge
von: Fu, Mingyang, et al.
Veröffentlicht: (2025)
von: Fu, Mingyang, et al.
Veröffentlicht: (2025)
Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
von: Liu, Benlin, et al.
Veröffentlicht: (2024)
von: Liu, Benlin, et al.
Veröffentlicht: (2024)
Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
von: Yang, Zeyuan, et al.
Veröffentlicht: (2025)
von: Yang, Zeyuan, et al.
Veröffentlicht: (2025)
MultiRef: Controllable Image Generation with Multiple Visual References
von: Chen, Ruoxi, et al.
Veröffentlicht: (2025)
von: Chen, Ruoxi, et al.
Veröffentlicht: (2025)
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
von: Hsieh, Yu-Guan, et al.
Veröffentlicht: (2024)
von: Hsieh, Yu-Guan, et al.
Veröffentlicht: (2024)
Iterated Learning Improves Compositionality in Large Vision-Language Models
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
von: Fan, Xiang, et al.
Veröffentlicht: (2026)
von: Fan, Xiang, et al.
Veröffentlicht: (2026)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
von: Che, Liwei, et al.
Veröffentlicht: (2026)
von: Che, Liwei, et al.
Veröffentlicht: (2026)
Gene-Level Representation Learning via Interventional Style Transfer in Optical Pooled Screening
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2024)
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2024)
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models
von: Tan, Xudong, et al.
Veröffentlicht: (2025)
von: Tan, Xudong, et al.
Veröffentlicht: (2025)
REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations
von: Sushko, Peter, et al.
Veröffentlicht: (2025)
von: Sushko, Peter, et al.
Veröffentlicht: (2025)
Synthetic Visual Genome
von: Park, Jae Sung, et al.
Veröffentlicht: (2025)
von: Park, Jae Sung, et al.
Veröffentlicht: (2025)
RefTok: Reference-Based Tokenization for Video Generation
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning
von: Zafar, Anas, et al.
Veröffentlicht: (2026)
von: Zafar, Anas, et al.
Veröffentlicht: (2026)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
Enhancing Vision Language Models with Logic Reasoning for Situational Awareness
von: Pradeep, Pavana, et al.
Veröffentlicht: (2026)
von: Pradeep, Pavana, et al.
Veröffentlicht: (2026)
Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
von: Zhang, Wenqi, et al.
Veröffentlicht: (2024)
von: Zhang, Wenqi, et al.
Veröffentlicht: (2024)
The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better
von: Geng, Scott, et al.
Veröffentlicht: (2024)
von: Geng, Scott, et al.
Veröffentlicht: (2024)
SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
von: Shen, Chuming, et al.
Veröffentlicht: (2025)
von: Shen, Chuming, et al.
Veröffentlicht: (2025)
Same or Not? Enhancing Visual Perception in Vision-Language Models
von: Marsili, Damiano, et al.
Veröffentlicht: (2025)
von: Marsili, Damiano, et al.
Veröffentlicht: (2025)
NVILA: Efficient Frontier Visual Language Models
von: Liu, Zhijian, et al.
Veröffentlicht: (2024)
von: Liu, Zhijian, et al.
Veröffentlicht: (2024)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
von: Fan, Yingqi, et al.
Veröffentlicht: (2026)
von: Fan, Yingqi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026) -
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
von: Li, Linjie, et al.
Veröffentlicht: (2025) -
MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2025) -
MIMIC: Masked Image Modeling with Image Correspondences
von: Marathe, Kalyani, et al.
Veröffentlicht: (2023) -
Data Alignment for Zero-Shot Concept Generation in Dermatology AI
von: Gadgil, Soham, et al.
Veröffentlicht: (2024)