FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Hao, Jing, Zhao, Yuxiang, Chen, Song, Sun, Yanpeng, Chen, Qiang, Zhang, Gang, Yao, Kun, Ding, Errui, Wang, Jingdong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VRP-SAM: SAM with Visual Reference Prompt
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
di: Zhang, Guosheng, et al.
Pubblicazione: (2025)
di: Zhang, Guosheng, et al.
Pubblicazione: (2025)
OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer
di: Wang, Yu, et al.
Pubblicazione: (2024)
di: Wang, Yu, et al.
Pubblicazione: (2024)
MS-DETR: Efficient DETR Training with Mixed Supervision
di: Zhao, Chuyang, et al.
Pubblicazione: (2024)
di: Zhao, Chuyang, et al.
Pubblicazione: (2024)
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
di: Lyu, Pengyuan, et al.
Pubblicazione: (2024)
di: Lyu, Pengyuan, et al.
Pubblicazione: (2024)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
Skim then Focus: Integrating Contextual and Fine-grained Views for Repetitive Action Counting
di: Zhao, Zhengqi, et al.
Pubblicazione: (2024)
di: Zhao, Zhengqi, et al.
Pubblicazione: (2024)
Add-SD: Rational Generation without Manual Reference
di: Yang, Lingfeng, et al.
Pubblicazione: (2024)
di: Yang, Lingfeng, et al.
Pubblicazione: (2024)
Exploring Effective Factors for Improving Visual In-Context Learning
di: Sun, Yanpeng, et al.
Pubblicazione: (2023)
di: Sun, Yanpeng, et al.
Pubblicazione: (2023)
MonoFormer: One Transformer for Both Diffusion and Autoregression
di: Zhao, Chuyang, et al.
Pubblicazione: (2024)
di: Zhao, Chuyang, et al.
Pubblicazione: (2024)
Continual SFT Matches Multimodal RLHF with Negative Supervision
di: Zhu, Ke, et al.
Pubblicazione: (2024)
di: Zhu, Ke, et al.
Pubblicazione: (2024)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
Towards Unified Multi-granularity Text Detection with Interactive Attention
di: Wan, Xingyu, et al.
Pubblicazione: (2024)
di: Wan, Xingyu, et al.
Pubblicazione: (2024)
Automated Multi-level Preference for MLLMs
di: Zhang, Mengxi, et al.
Pubblicazione: (2024)
di: Zhang, Mengxi, et al.
Pubblicazione: (2024)
LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection
di: Chen, Qiang, et al.
Pubblicazione: (2024)
di: Chen, Qiang, et al.
Pubblicazione: (2024)
XLD: A Cross-Lane Dataset for Benchmarking Novel Driving View Synthesis
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
Dense Connector for MLLMs
di: Yao, Huanjin, et al.
Pubblicazione: (2024)
di: Yao, Huanjin, et al.
Pubblicazione: (2024)
MedQ-Engine: A Closed-Loop Data Engine for Evolving MLLMs in Medical Image Quality Assessment
di: Liu, Jiyao, et al.
Pubblicazione: (2026)
di: Liu, Jiyao, et al.
Pubblicazione: (2026)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
di: Ouyang, Kun, et al.
Pubblicazione: (2024)
di: Ouyang, Kun, et al.
Pubblicazione: (2024)
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
di: Lu, Junxin, et al.
Pubblicazione: (2026)
di: Lu, Junxin, et al.
Pubblicazione: (2026)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
di: Zhang, Tao, et al.
Pubblicazione: (2025)
di: Zhang, Tao, et al.
Pubblicazione: (2025)
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
di: Liu, Huan, et al.
Pubblicazione: (2024)
di: Liu, Huan, et al.
Pubblicazione: (2024)
GGRt: Towards Pose-free Generalizable 3D Gaussian Splatting in Real-time
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
VDG: Vision-Only Dynamic Gaussian for Driving Simulation
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map Construction
di: Yang, Jing, et al.
Pubblicazione: (2024)
di: Yang, Jing, et al.
Pubblicazione: (2024)
GVA: Reconstructing Vivid 3D Gaussian Avatars from Monocular Videos
di: Liu, Xinqi, et al.
Pubblicazione: (2024)
di: Liu, Xinqi, et al.
Pubblicazione: (2024)
Splatter-360: Generalizable 360$^{\circ}$ Gaussian Splatting for Wide-baseline Panoramic Images
di: Chen, Zheng, et al.
Pubblicazione: (2024)
di: Chen, Zheng, et al.
Pubblicazione: (2024)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
di: Li, Qiaoru, et al.
Pubblicazione: (2026)
di: Li, Qiaoru, et al.
Pubblicazione: (2026)
LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction
di: Du, Penghui, et al.
Pubblicazione: (2024)
di: Du, Penghui, et al.
Pubblicazione: (2024)
Anno-incomplete Multi-dataset Detection
di: Xu, Yiran, et al.
Pubblicazione: (2024)
di: Xu, Yiran, et al.
Pubblicazione: (2024)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
di: Sun, Yanpeng, et al.
Pubblicazione: (2025)
di: Sun, Yanpeng, et al.
Pubblicazione: (2025)
TexRO: Generating Delicate Textures of 3D Models by Recursive Optimization
di: Wu, Jinbo, et al.
Pubblicazione: (2024)
di: Wu, Jinbo, et al.
Pubblicazione: (2024)
GIR: 3D Gaussian Inverse Rendering for Relightable Scene Factorization
di: Shi, Yahao, et al.
Pubblicazione: (2023)
di: Shi, Yahao, et al.
Pubblicazione: (2023)
ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
di: Zhang, Shilan, et al.
Pubblicazione: (2025)
di: Zhang, Shilan, et al.
Pubblicazione: (2025)
CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval
di: Wang, Haoran, et al.
Pubblicazione: (2022)
di: Wang, Haoran, et al.
Pubblicazione: (2022)
TopoSD: Topology-Enhanced Lane Segment Perception with SDMap Prior
di: Yang, Sen, et al.
Pubblicazione: (2024)
di: Yang, Sen, et al.
Pubblicazione: (2024)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
di: Zhang, Shan, et al.
Pubblicazione: (2025)
di: Zhang, Shan, et al.
Pubblicazione: (2025)
Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation
di: Chen, Bohong, et al.
Pubblicazione: (2024)
di: Chen, Bohong, et al.
Pubblicazione: (2024)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
di: Miao, Ziqi, et al.
Pubblicazione: (2025)
di: Miao, Ziqi, et al.
Pubblicazione: (2025)
ForgeryVCR: Visual-Centric Reasoning via Efficient Forensic Tools in MLLMs for Image Forgery Detection and Localization
di: Wang, Youqi, et al.
Pubblicazione: (2026)
di: Wang, Youqi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
VRP-SAM: SAM with Visual Reference Prompt
di: Sun, Yanpeng, et al.
Pubblicazione: (2024) -
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
di: Zhang, Guosheng, et al.
Pubblicazione: (2025) -
OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer
di: Wang, Yu, et al.
Pubblicazione: (2024) -
MS-DETR: Efficient DETR Training with Mixed Supervision
di: Zhao, Chuyang, et al.
Pubblicazione: (2024) -
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
di: Lyu, Pengyuan, et al.
Pubblicazione: (2024)