FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Hao, Jing, Zhao, Yuxiang, Chen, Song, Sun, Yanpeng, Chen, Qiang, Zhang, Gang, Yao, Kun, Ding, Errui, Wang, Jingdong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VRP-SAM: SAM with Visual Reference Prompt
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
by: Zhang, Guosheng, et al.
Published: (2025)
by: Zhang, Guosheng, et al.
Published: (2025)
OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
MS-DETR: Efficient DETR Training with Mixed Supervision
by: Zhao, Chuyang, et al.
Published: (2024)
by: Zhao, Chuyang, et al.
Published: (2024)
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
by: Lyu, Pengyuan, et al.
Published: (2024)
by: Lyu, Pengyuan, et al.
Published: (2024)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Skim then Focus: Integrating Contextual and Fine-grained Views for Repetitive Action Counting
by: Zhao, Zhengqi, et al.
Published: (2024)
by: Zhao, Zhengqi, et al.
Published: (2024)
Add-SD: Rational Generation without Manual Reference
by: Yang, Lingfeng, et al.
Published: (2024)
by: Yang, Lingfeng, et al.
Published: (2024)
Exploring Effective Factors for Improving Visual In-Context Learning
by: Sun, Yanpeng, et al.
Published: (2023)
by: Sun, Yanpeng, et al.
Published: (2023)
MonoFormer: One Transformer for Both Diffusion and Autoregression
by: Zhao, Chuyang, et al.
Published: (2024)
by: Zhao, Chuyang, et al.
Published: (2024)
Continual SFT Matches Multimodal RLHF with Negative Supervision
by: Zhu, Ke, et al.
Published: (2024)
by: Zhu, Ke, et al.
Published: (2024)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Towards Unified Multi-granularity Text Detection with Interactive Attention
by: Wan, Xingyu, et al.
Published: (2024)
by: Wan, Xingyu, et al.
Published: (2024)
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024)
by: Zhang, Mengxi, et al.
Published: (2024)
LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection
by: Chen, Qiang, et al.
Published: (2024)
by: Chen, Qiang, et al.
Published: (2024)
XLD: A Cross-Lane Dataset for Benchmarking Novel Driving View Synthesis
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Dense Connector for MLLMs
by: Yao, Huanjin, et al.
Published: (2024)
by: Yao, Huanjin, et al.
Published: (2024)
MedQ-Engine: A Closed-Loop Data Engine for Evolving MLLMs in Medical Image Quality Assessment
by: Liu, Jiyao, et al.
Published: (2026)
by: Liu, Jiyao, et al.
Published: (2026)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
by: Ouyang, Kun, et al.
Published: (2024)
by: Ouyang, Kun, et al.
Published: (2024)
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
by: Lu, Junxin, et al.
Published: (2026)
by: Lu, Junxin, et al.
Published: (2026)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
by: Liu, Huan, et al.
Published: (2024)
by: Liu, Huan, et al.
Published: (2024)
GGRt: Towards Pose-free Generalizable 3D Gaussian Splatting in Real-time
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
VDG: Vision-Only Dynamic Gaussian for Driving Simulation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map Construction
by: Yang, Jing, et al.
Published: (2024)
by: Yang, Jing, et al.
Published: (2024)
GVA: Reconstructing Vivid 3D Gaussian Avatars from Monocular Videos
by: Liu, Xinqi, et al.
Published: (2024)
by: Liu, Xinqi, et al.
Published: (2024)
Splatter-360: Generalizable 360$^{\circ}$ Gaussian Splatting for Wide-baseline Panoramic Images
by: Chen, Zheng, et al.
Published: (2024)
by: Chen, Zheng, et al.
Published: (2024)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
by: Li, Qiaoru, et al.
Published: (2026)
by: Li, Qiaoru, et al.
Published: (2026)
LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction
by: Du, Penghui, et al.
Published: (2024)
by: Du, Penghui, et al.
Published: (2024)
Anno-incomplete Multi-dataset Detection
by: Xu, Yiran, et al.
Published: (2024)
by: Xu, Yiran, et al.
Published: (2024)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
by: Sun, Yanpeng, et al.
Published: (2025)
by: Sun, Yanpeng, et al.
Published: (2025)
TexRO: Generating Delicate Textures of 3D Models by Recursive Optimization
by: Wu, Jinbo, et al.
Published: (2024)
by: Wu, Jinbo, et al.
Published: (2024)
GIR: 3D Gaussian Inverse Rendering for Relightable Scene Factorization
by: Shi, Yahao, et al.
Published: (2023)
by: Shi, Yahao, et al.
Published: (2023)
ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
by: Zhang, Shilan, et al.
Published: (2025)
by: Zhang, Shilan, et al.
Published: (2025)
CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval
by: Wang, Haoran, et al.
Published: (2022)
by: Wang, Haoran, et al.
Published: (2022)
TopoSD: Topology-Enhanced Lane Segment Perception with SDMap Prior
by: Yang, Sen, et al.
Published: (2024)
by: Yang, Sen, et al.
Published: (2024)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation
by: Chen, Bohong, et al.
Published: (2024)
by: Chen, Bohong, et al.
Published: (2024)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
ForgeryVCR: Visual-Centric Reasoning via Efficient Forensic Tools in MLLMs for Image Forgery Detection and Localization
by: Wang, Youqi, et al.
Published: (2026)
by: Wang, Youqi, et al.
Published: (2026)
Similar Items
-
VRP-SAM: SAM with Visual Reference Prompt
by: Sun, Yanpeng, et al.
Published: (2024) -
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
by: Zhang, Guosheng, et al.
Published: (2025) -
OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer
by: Wang, Yu, et al.
Published: (2024) -
MS-DETR: Efficient DETR Training with Mixed Supervision
by: Zhao, Chuyang, et al.
Published: (2024) -
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
by: Lyu, Pengyuan, et al.
Published: (2024)