Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Yuheng, Pei, Xiaohuan, Dong, Minjing, Xu, Chang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models
by: Shi, Yuheng, et al.
Published: (2026)
by: Shi, Yuheng, et al.
Published: (2026)
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation
by: Shi, Yuheng, et al.
Published: (2024)
by: Shi, Yuheng, et al.
Published: (2024)
Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space Model
by: Shi, Yuheng, et al.
Published: (2024)
by: Shi, Yuheng, et al.
Published: (2024)
Multi-RoI Human Mesh Recovery with Camera Consistency and Contrastive Losses
by: Nie, Yongwei, et al.
Published: (2024)
by: Nie, Yongwei, et al.
Published: (2024)
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
by: Zhang, Shilong, et al.
Published: (2023)
by: Zhang, Shilong, et al.
Published: (2023)
VSSD: Vision Mamba with Non-Causal State Space Duality
by: Shi, Yuheng, et al.
Published: (2024)
by: Shi, Yuheng, et al.
Published: (2024)
DetVPCC: RoI-based Point Cloud Sequence Compression for 3D Object Detection
by: Yan, Mingxuan, et al.
Published: (2025)
by: Yan, Mingxuan, et al.
Published: (2025)
Cross-Self KV Cache Pruning for Efficient Vision-Language Inference
by: Pei, Xiaohuan, et al.
Published: (2024)
by: Pei, Xiaohuan, et al.
Published: (2024)
Efficient Image-to-Image Diffusion Classifier for Adversarial Robustness
by: Mei, Hefei, et al.
Published: (2024)
by: Mei, Hefei, et al.
Published: (2024)
FiGKD: Fine-Grained Knowledge Distillation via High-Frequency Detail Transfer
by: Kim, Seonghak
Published: (2025)
by: Kim, Seonghak
Published: (2025)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
EfficientVMamba: Atrous Selective Scan for Light Weight Visual Mamba
by: Pei, Xiaohuan, et al.
Published: (2024)
by: Pei, Xiaohuan, et al.
Published: (2024)
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
by: Gaur, Manu, et al.
Published: (2024)
by: Gaur, Manu, et al.
Published: (2024)
Feature Clipping for Uncertainty Calibration
by: Tao, Linwei, et al.
Published: (2024)
by: Tao, Linwei, et al.
Published: (2024)
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
by: Zhong, Liangyu, et al.
Published: (2025)
by: Zhong, Liangyu, et al.
Published: (2025)
Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
by: Huang, Qihan, et al.
Published: (2025)
by: Huang, Qihan, et al.
Published: (2025)
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
by: Zhao, Peisen, et al.
Published: (2026)
by: Zhao, Peisen, et al.
Published: (2026)
FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition
by: Yu, Enhui, et al.
Published: (2026)
by: Yu, Enhui, et al.
Published: (2026)
VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
by: Mei, Hefei, et al.
Published: (2025)
by: Mei, Hefei, et al.
Published: (2025)
PA-Attack: Guiding Gray-Box Attacks on LVLM Vision Encoders with Prototypes and Attention
by: Mei, Hefei, et al.
Published: (2026)
by: Mei, Hefei, et al.
Published: (2026)
DEIG: Detail-Enhanced Instance Generation with Fine-Grained Semantic Control
by: Du, Shiyan, et al.
Published: (2026)
by: Du, Shiyan, et al.
Published: (2026)
PP-SSL : Priority-Perception Self-Supervised Learning for Fine-Grained Recognition
by: Li, ShuaiHeng, et al.
Published: (2024)
by: Li, ShuaiHeng, et al.
Published: (2024)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
by: Zhu, Younan, et al.
Published: (2025)
by: Zhu, Younan, et al.
Published: (2025)
Beyond One-Hot Labels: Semantic Mixing for Model Calibration
by: Luo, Haoyang, et al.
Published: (2025)
by: Luo, Haoyang, et al.
Published: (2025)
Diversifying Counterattacks: Orthogonal Exploration for Robust CLIP Inference
by: Jiang, Chengze, et al.
Published: (2025)
by: Jiang, Chengze, et al.
Published: (2025)
Cross-Level Multi-Instance Distillation for Self-Supervised Fine-Grained Visual Categorization
by: Bi, Qi, et al.
Published: (2024)
by: Bi, Qi, et al.
Published: (2024)
Micro-Expression Recognition via Fine-Grained Dynamic Perception
by: Shao, Zhiwen, et al.
Published: (2025)
by: Shao, Zhiwen, et al.
Published: (2025)
D3FNet: A Differential Attention Fusion Network for Fine-Grained Road Structure Extraction in Remote Perception Systems
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Rethinking Causal Mask Attention for Vision-Language Inference
by: Pei, Xiaohuan, et al.
Published: (2025)
by: Pei, Xiaohuan, et al.
Published: (2025)
Beyond Illumination: Fine-Grained Detail Preservation in Extreme Dark Image Restoration
by: Zhang, Tongshun, et al.
Published: (2025)
by: Zhang, Tongshun, et al.
Published: (2025)
Rotation Augmented Distillation for Exemplar-Free Class Incremental Learning with Detailed Analysis
by: Chen, Xiuwei, et al.
Published: (2023)
by: Chen, Xiuwei, et al.
Published: (2023)
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
by: Wu, Xueqing, et al.
Published: (2024)
by: Wu, Xueqing, et al.
Published: (2024)
New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration
by: Yang, Xuzheng, et al.
Published: (2025)
by: Yang, Xuzheng, et al.
Published: (2025)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
by: Wei, Lai, et al.
Published: (2026)
by: Wei, Lai, et al.
Published: (2026)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
by: Xu, Binqian, et al.
Published: (2024)
by: Xu, Binqian, et al.
Published: (2024)
EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
by: Shan, Haozhe, et al.
Published: (2026)
by: Shan, Haozhe, et al.
Published: (2026)
DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
Backdooring Self-Supervised Contrastive Learning by Noisy Alignment
by: Chen, Tuo, et al.
Published: (2025)
by: Chen, Tuo, et al.
Published: (2025)
Fine-Grained Prototypes Distillation for Few-Shot Object Detection
by: Wang, Zichen, et al.
Published: (2024)
by: Wang, Zichen, et al.
Published: (2024)
Self-Improving 4D Perception via Self-Distillation
by: Huang, Nan, et al.
Published: (2026)
by: Huang, Nan, et al.
Published: (2026)
Similar Items
-
Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models
by: Shi, Yuheng, et al.
Published: (2026) -
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation
by: Shi, Yuheng, et al.
Published: (2024) -
Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space Model
by: Shi, Yuheng, et al.
Published: (2024) -
Multi-RoI Human Mesh Recovery with Camera Consistency and Contrastive Losses
by: Nie, Yongwei, et al.
Published: (2024) -
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
by: Zhang, Shilong, et al.
Published: (2023)