Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Yongyi, Zhang, Haojie, Li, Shijie, Liu, Nanqing, Liao, Jingyi, Pan, Junyi, Liu, Yuan, Xing, Xiaofen, Sun, Chong, Li, Chen, Chen, Nancy F., Yan, Shuicheng, Yang, Xulei, Xu, Xun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-View Industrial Anomaly Detection with Epipolar Constrained Cross-View Fusion
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
PointSAM: Pointly-Supervised Segment Anything Model for Remote Sensing Images
by: Liu, Nanqing, et al.
Published: (2024)
by: Liu, Nanqing, et al.
Published: (2024)
On the Adversarial Risk of Test Time Adaptation: An Investigation into Realistic Test-Time Data Poisoning
by: Su, Yongyi, et al.
Published: (2024)
by: Su, Yongyi, et al.
Published: (2024)
Thinking Ahead: Foresight Intelligence in MLLMs and World Models
by: Gong, Zhantao, et al.
Published: (2025)
by: Gong, Zhantao, et al.
Published: (2025)
Robust Distribution Alignment for Industrial Anomaly Detection under Distribution Shift
by: Liao, Jingyi, et al.
Published: (2025)
by: Liao, Jingyi, et al.
Published: (2025)
HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
by: Nie, Sihang, et al.
Published: (2025)
by: Nie, Sihang, et al.
Published: (2025)
CLIP-Guided Source-Free Object Detection in Aerial Images
by: Liu, Nanqing, et al.
Published: (2024)
by: Liu, Nanqing, et al.
Published: (2024)
SODA: Out-of-Distribution Detection in Domain-Shifted Point Clouds via Neighborhood Propagation
by: Goodge, Adam, et al.
Published: (2025)
by: Goodge, Adam, et al.
Published: (2025)
Exploring Human-in-the-Loop Test-Time Adaptation by Synergizing Active Learning and Model Selection
by: Li, Yushu, et al.
Published: (2024)
by: Li, Yushu, et al.
Published: (2024)
AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization
by: Liao, Jingyi, et al.
Published: (2025)
by: Liao, Jingyi, et al.
Published: (2025)
Future-Aware Interaction Network For Motion Forecasting
by: Li, Shijie, et al.
Published: (2025)
by: Li, Shijie, et al.
Published: (2025)
MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances
by: Wang, Qirui, et al.
Published: (2026)
by: Wang, Qirui, et al.
Published: (2026)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
by: Li, Hongliang, et al.
Published: (2025)
by: Li, Hongliang, et al.
Published: (2025)
Zero-Shot 3D Visual Grounding from Vision-Language Models
by: Li, Rong, et al.
Published: (2025)
by: Li, Rong, et al.
Published: (2025)
Improving the Generalization of Segmentation Foundation Model under Distribution Shift via Weakly Supervised Adaptation
by: Zhang, Haojie, et al.
Published: (2023)
by: Zhang, Haojie, et al.
Published: (2023)
Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks
by: Leong, Mei Chee, et al.
Published: (2026)
by: Leong, Mei Chee, et al.
Published: (2026)
PsyDT: Using LLMs to Construct the Digital Twin of Psychological Counselor with Personalized Counseling Style for Psychological Counseling
by: Xie, Haojie, et al.
Published: (2024)
by: Xie, Haojie, et al.
Published: (2024)
One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs Hallucination
by: Fa, Zhan, et al.
Published: (2026)
by: Fa, Zhan, et al.
Published: (2026)
VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization
by: Li, Mingxiao, et al.
Published: (2025)
by: Li, Mingxiao, et al.
Published: (2025)
Sparsity Forcing: Reinforcing Token Sparsity of MLLMs
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs
by: Chen, Tianle, et al.
Published: (2025)
by: Chen, Tianle, et al.
Published: (2025)
InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs
by: Tang, Lv, et al.
Published: (2026)
by: Tang, Lv, et al.
Published: (2026)
Global-Aware Monocular Semantic Scene Completion with State Space Models
by: Li, Shijie, et al.
Published: (2025)
by: Li, Shijie, et al.
Published: (2025)
MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any Resolution
by: Liu, Wenzhuo, et al.
Published: (2024)
by: Liu, Wenzhuo, et al.
Published: (2024)
UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph Generation
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
Towards Annotation-Free Validation of MLLMs: A Vision-Language Logical Consistency Metric
by: Gu, Ying, et al.
Published: (2026)
by: Gu, Ying, et al.
Published: (2026)
Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score Propagation
by: Chen, Tiankai, et al.
Published: (2025)
by: Chen, Tiankai, et al.
Published: (2025)
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images
by: Wang, Qirui, et al.
Published: (2025)
by: Wang, Qirui, et al.
Published: (2025)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
by: Chen, Siyi, et al.
Published: (2026)
by: Chen, Siyi, et al.
Published: (2026)
MLLMs are Deeply Affected by Modality Bias
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
AToken: A Unified Tokenizer for Vision
by: Lu, Jiasen, et al.
Published: (2025)
by: Lu, Jiasen, et al.
Published: (2025)
Training-Free Multi-Step Audio Source Separation
by: Zang, Yongyi, et al.
Published: (2025)
by: Zang, Yongyi, et al.
Published: (2025)
Majorization-Guided Test-Time Adaptation for Vision-Language Models under Modality-Specific Shift
by: Chen, Lixian, et al.
Published: (2026)
by: Chen, Lixian, et al.
Published: (2026)
An Empirical Study of Speculative Decoding on Software Engineering Tasks
by: Li, Yijia, et al.
Published: (2026)
by: Li, Yijia, et al.
Published: (2026)
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
by: Jin, Yang, et al.
Published: (2023)
by: Jin, Yang, et al.
Published: (2023)
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
by: Yu, Hanxun, et al.
Published: (2026)
by: Yu, Hanxun, et al.
Published: (2026)
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
by: Wang, Zitian, et al.
Published: (2025)
by: Wang, Zitian, et al.
Published: (2025)
Similar Items
-
Multi-View Industrial Anomaly Detection with Epipolar Constrained Cross-View Fusion
by: Liu, Yifan, et al.
Published: (2025) -
PointSAM: Pointly-Supervised Segment Anything Model for Remote Sensing Images
by: Liu, Nanqing, et al.
Published: (2024) -
On the Adversarial Risk of Test Time Adaptation: An Investigation into Realistic Test-Time Data Poisoning
by: Su, Yongyi, et al.
Published: (2024) -
Thinking Ahead: Foresight Intelligence in MLLMs and World Models
by: Gong, Zhantao, et al.
Published: (2025) -
Robust Distribution Alignment for Industrial Anomaly Detection under Distribution Shift
by: Liao, Jingyi, et al.
Published: (2025)