Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Mingxin, Liu, Yuliang, Liang, Dingkang, Jin, Lianwen, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Progressive Evolution from Single-Point to Polygon for Scene Text
by: Deng, Linger, et al.
Published: (2023)
by: Deng, Linger, et al.
Published: (2023)
VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
by: Liu, Yuliang, et al.
Published: (2024)
by: Liu, Yuliang, et al.
Published: (2024)
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
by: Peng, Dezhi, et al.
Published: (2023)
by: Peng, Dezhi, et al.
Published: (2023)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
by: Liu, Yuliang, et al.
Published: (2023)
by: Liu, Yuliang, et al.
Published: (2023)
A Unified Image-Dense Annotation Generation Model for Underwater Scenes
by: Lin, Hongkai, et al.
Published: (2025)
by: Lin, Hongkai, et al.
Published: (2025)
More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
by: Lin, Hongkai, et al.
Published: (2025)
by: Lin, Hongkai, et al.
Published: (2025)
Puzzle Pieces Picker: Deciphering Ancient Chinese Characters with Radical Reconstruction
by: Wang, Pengjie, et al.
Published: (2024)
by: Wang, Pengjie, et al.
Published: (2024)
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
by: Liu, Yuliang, et al.
Published: (2024)
by: Liu, Yuliang, et al.
Published: (2024)
MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks
by: Zhu, Xingkui, et al.
Published: (2024)
by: Zhu, Xingkui, et al.
Published: (2024)
SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
by: Luo, Dongliang, et al.
Published: (2025)
by: Luo, Dongliang, et al.
Published: (2025)
Hierarchical Side-Tuning for Vision Transformers
by: Lin, Weifeng, et al.
Published: (2023)
by: Lin, Weifeng, et al.
Published: (2023)
Privacy-Preserving Biometric Verification with Handwritten Random Digit String
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
by: Li, Zhang, et al.
Published: (2023)
by: Li, Zhang, et al.
Published: (2023)
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
by: Li, Zhang, et al.
Published: (2025)
by: Li, Zhang, et al.
Published: (2025)
Deciphering Oracle Bone Language with Diffusion Models
by: Guan, Haisu, et al.
Published: (2024)
by: Guan, Haisu, et al.
Published: (2024)
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs
by: Li, Hongliang, et al.
Published: (2025)
by: Li, Hongliang, et al.
Published: (2025)
MINIMA: Modality Invariant Image Matching
by: Ren, Jiangwei, et al.
Published: (2024)
by: Ren, Jiangwei, et al.
Published: (2024)
LEGO: Self-Supervised Representation Learning for Scene Text Images
by: Ren, Yujin, et al.
Published: (2024)
by: Ren, Yujin, et al.
Published: (2024)
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
by: Yin, Weijie, et al.
Published: (2025)
by: Yin, Weijie, et al.
Published: (2025)
Online Writer Retrieval with Chinese Handwritten Phrases: A Synergistic Temporal-Frequency Representation Learning Approach
by: Zhang, Peirong, et al.
Published: (2024)
by: Zhang, Peirong, et al.
Published: (2024)
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
by: Ou, Linyu, et al.
Published: (2025)
by: Ou, Linyu, et al.
Published: (2025)
Webly-Supervised Image Manipulation Localization via Category-Aware Auto-Annotation
by: Qu, Chenfan, et al.
Published: (2025)
by: Qu, Chenfan, et al.
Published: (2025)
An open dataset for oracle bone script recognition and decipherment
by: Wang, Pengjie, et al.
Published: (2024)
by: Wang, Pengjie, et al.
Published: (2024)
Anomaly Detection by Adapting a pre-trained Vision Language Model
by: Cai, Yuxuan, et al.
Published: (2024)
by: Cai, Yuxuan, et al.
Published: (2024)
PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding
by: Liu, Siyuan, et al.
Published: (2026)
by: Liu, Siyuan, et al.
Published: (2026)
TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering
by: Zhu, Hanshen, et al.
Published: (2026)
by: Zhu, Hanshen, et al.
Published: (2026)
VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
by: Zhang, Ruifei, et al.
Published: (2025)
by: Zhang, Ruifei, et al.
Published: (2025)
Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
by: Han, Minghao, et al.
Published: (2025)
by: Han, Minghao, et al.
Published: (2025)
DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
by: Yin, Liang, et al.
Published: (2025)
by: Yin, Liang, et al.
Published: (2025)
Make Your ViT-based Multi-view 3D Detectors Faster via Token Compression
by: Zhang, Dingyuan, et al.
Published: (2024)
by: Zhang, Dingyuan, et al.
Published: (2024)
DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation
by: Wang, Jiapeng, et al.
Published: (2024)
by: Wang, Jiapeng, et al.
Published: (2024)
MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs
by: Liu, Jiyao, et al.
Published: (2025)
by: Liu, Jiyao, et al.
Published: (2025)
Training-free Geometric Image Editing on Diffusion Models
by: Zhu, Hanshen, et al.
Published: (2025)
by: Zhu, Hanshen, et al.
Published: (2025)
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning in Spectral Domain for Point Cloud Learning
by: Liang, Dingkang, et al.
Published: (2024)
by: Liang, Dingkang, et al.
Published: (2024)
Similar Items
-
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024) -
Progressive Evolution from Single-Point to Polygon for Scene Text
by: Deng, Linger, et al.
Published: (2023) -
VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
by: Liu, Yuliang, et al.
Published: (2024) -
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
by: Huang, Mingxin, et al.
Published: (2024) -
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
by: Peng, Dezhi, et al.
Published: (2023)