Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mao, Junyuan, Li, Qiankun, Meng, Linghao, He, Zhicheng, Zhou, Xinliang, Wang, Kun, Liu, Yang, Jin, Yueming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
IPCV: Information-Preserving Compression for MLLM Visual Encoders
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
Multi-scale Temporal Prediction via Incremental Generation and Multi-agent Collaboration
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding
von: Wang, Nan, et al.
Veröffentlicht: (2026)
von: Wang, Nan, et al.
Veröffentlicht: (2026)
Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach
von: Liu, Feiyang, et al.
Veröffentlicht: (2024)
von: Liu, Feiyang, et al.
Veröffentlicht: (2024)
Semantic Similarity Score for Measuring Visual Similarity at Semantic Level
von: Fan, Senran, et al.
Veröffentlicht: (2024)
von: Fan, Senran, et al.
Veröffentlicht: (2024)
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
von: Fang, Rongyao, et al.
Veröffentlicht: (2024)
von: Fang, Rongyao, et al.
Veröffentlicht: (2024)
MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
von: Ran, Ran, et al.
Veröffentlicht: (2026)
von: Ran, Ran, et al.
Veröffentlicht: (2026)
GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
von: He, Feng, et al.
Veröffentlicht: (2025)
von: He, Feng, et al.
Veröffentlicht: (2025)
InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perception
von: Li, Haijie, et al.
Veröffentlicht: (2024)
von: Li, Haijie, et al.
Veröffentlicht: (2024)
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
PixelLM: Pixel Reasoning with Large Multimodal Model
von: Ren, Zhongwei, et al.
Veröffentlicht: (2023)
von: Ren, Zhongwei, et al.
Veröffentlicht: (2023)
CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder
von: Ma, Lichen, et al.
Veröffentlicht: (2024)
von: Ma, Lichen, et al.
Veröffentlicht: (2024)
Tuning a SAM-Based Model with Multi-Cognitive Visual Adapter to Remote Sensing Instance Segmentation
von: Zheng, Linghao, et al.
Veröffentlicht: (2024)
von: Zheng, Linghao, et al.
Veröffentlicht: (2024)
Efficient Reinforcement Learning Through Adaptively Pretrained Visual Encoder
von: Zhang, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhan, et al.
Veröffentlicht: (2025)
SSA-Seg: Semantic and Spatial Adaptive Pixel-level Classifier for Semantic Segmentation
von: Ma, Xiaowen, et al.
Veröffentlicht: (2024)
von: Ma, Xiaowen, et al.
Veröffentlicht: (2024)
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities
von: Liu, Jing, et al.
Veröffentlicht: (2025)
von: Liu, Jing, et al.
Veröffentlicht: (2025)
Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
von: Li, Qiankun, et al.
Veröffentlicht: (2025)
von: Li, Qiankun, et al.
Veröffentlicht: (2025)
Robust Domain Adaptive Object Detection with Unified Multi-Granularity Alignment
von: Zhang, Libo, et al.
Veröffentlicht: (2023)
von: Zhang, Libo, et al.
Veröffentlicht: (2023)
Scribble Hides Class: Promoting Scribble-Based Weakly-Supervised Semantic Segmentation with Its Class Label
von: Zhang, Xinliang, et al.
Veröffentlicht: (2024)
von: Zhang, Xinliang, et al.
Veröffentlicht: (2024)
Weakly Supervised Pixel-Level Annotation with Visual Interpretability
von: Nasir, Basma, et al.
Veröffentlicht: (2025)
von: Nasir, Basma, et al.
Veröffentlicht: (2025)
Semantic Granularity Navigation in Image Editing
von: Lu, Liangsi, et al.
Veröffentlicht: (2026)
von: Lu, Liangsi, et al.
Veröffentlicht: (2026)
Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Visual Neural Decoding via Improved Visual-EEG Semantic Consistency
von: Chen, Hongzhou, et al.
Veröffentlicht: (2024)
von: Chen, Hongzhou, et al.
Veröffentlicht: (2024)
UPOCR: Towards Unified Pixel-Level OCR Interface
von: Peng, Dezhi, et al.
Veröffentlicht: (2023)
von: Peng, Dezhi, et al.
Veröffentlicht: (2023)
Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
von: Liu, Jinming, et al.
Veröffentlicht: (2025)
von: Liu, Jinming, et al.
Veröffentlicht: (2025)
From Semantics to Pixels: Coarse-to-Fine Masked Autoencoders for Hierarchical Visual Understanding
von: Xiang, Wenzhao, et al.
Veröffentlicht: (2026)
von: Xiang, Wenzhao, et al.
Veröffentlicht: (2026)
Robust MLLM Unlearning via Visual Knowledge Distillation
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
von: Wang, Yuhang, et al.
Veröffentlicht: (2025)
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
von: Cai, Dexian, et al.
Veröffentlicht: (2025)
von: Cai, Dexian, et al.
Veröffentlicht: (2025)
FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
von: Liu, Yaoli, et al.
Veröffentlicht: (2025)
von: Liu, Yaoli, et al.
Veröffentlicht: (2025)
MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
von: Tong, Qinyue, et al.
Veröffentlicht: (2025)
von: Tong, Qinyue, et al.
Veröffentlicht: (2025)
Reducing Semantic Ambiguity In Domain Adaptive Semantic Segmentation Via Probabilistic Prototypical Pixel Contrast
von: Hao, Xiaoke, et al.
Veröffentlicht: (2024)
von: Hao, Xiaoke, et al.
Veröffentlicht: (2024)
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
von: You, Zuyao, et al.
Veröffentlicht: (2025)
von: You, Zuyao, et al.
Veröffentlicht: (2025)
InstructX: Towards Unified Visual Editing with MLLM Guidance
von: Mou, Chong, et al.
Veröffentlicht: (2025)
von: Mou, Chong, et al.
Veröffentlicht: (2025)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
von: Chang, Boyu, et al.
Veröffentlicht: (2026)
von: Chang, Boyu, et al.
Veröffentlicht: (2026)
Enhancing Lip Reading with Multi-Scale Video and Multi-Encoder
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
von: Jiang, Dongsheng, et al.
Veröffentlicht: (2023)
von: Jiang, Dongsheng, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
IPCV: Information-Preserving Compression for MLLM Visual Encoders
von: Chen, Yuan, et al.
Veröffentlicht: (2025) -
Multi-scale Temporal Prediction via Incremental Generation and Multi-agent Collaboration
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025) -
PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding
von: Wang, Nan, et al.
Veröffentlicht: (2026) -
Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach
von: Liu, Feiyang, et al.
Veröffentlicht: (2024) -
Semantic Similarity Score for Measuring Visual Similarity at Semantic Level
von: Fan, Senran, et al.
Veröffentlicht: (2024)