Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Han, Huo, Yuqi, Zhao, Zijia, Lu, Haoyu, Wu, Shu, Wang, Bingning, Liu, Qiang, Chen, Weipeng, Wang, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Motion-Aware Video MLLM
by: Zhao, Zijia, et al.
Published: (2025)
by: Zhao, Zijia, et al.
Published: (2025)
Exploring the Design Space of Visual Context Representation in Video MLLMs
by: Du, Yifan, et al.
Published: (2024)
by: Du, Yifan, et al.
Published: (2024)
Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
by: Zhao, Zijia, et al.
Published: (2024)
by: Zhao, Zijia, et al.
Published: (2024)
Towards Event-oriented Long Video Understanding
by: Du, Yifan, et al.
Published: (2024)
by: Du, Yifan, et al.
Published: (2024)
Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
by: Du, Yifan, et al.
Published: (2025)
by: Du, Yifan, et al.
Published: (2025)
MetaGPT: Merging Large Language Models Using Model Exclusive Task Arithmetic
by: Zhou, Yuyan, et al.
Published: (2024)
by: Zhou, Yuyan, et al.
Published: (2024)
Checkpoint Merging via Bayesian Optimization in LLM Pretraining
by: Liu, Deyuan, et al.
Published: (2024)
by: Liu, Deyuan, et al.
Published: (2024)
KV Shifting Attention Enhances Language Modeling
by: Xu, Mingyu, et al.
Published: (2024)
by: Xu, Mingyu, et al.
Published: (2024)
Full-ECE: A Metric For Token-level Calibration on Large Language Models
by: Liu, Han, et al.
Published: (2024)
by: Liu, Han, et al.
Published: (2024)
LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation
by: Dong, Zican, et al.
Published: (2025)
by: Dong, Zican, et al.
Published: (2025)
HRM-Text: Efficient Pretraining Beyond Scaling
by: Wang, Guan, et al.
Published: (2026)
by: Wang, Guan, et al.
Published: (2026)
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
by: Men, Xin, et al.
Published: (2024)
by: Men, Xin, et al.
Published: (2024)
Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters
by: Wang, Weizhi, et al.
Published: (2024)
by: Wang, Weizhi, et al.
Published: (2024)
Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language Models
by: Chen, Zhipeng, et al.
Published: (2024)
by: Chen, Zhipeng, et al.
Published: (2024)
Base of RoPE Bounds Context Length
by: Men, Xin, et al.
Published: (2024)
by: Men, Xin, et al.
Published: (2024)
Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models
by: Chen, Jie, et al.
Published: (2024)
by: Chen, Jie, et al.
Published: (2024)
StyleMamba : State Space Model for Efficient Text-driven Image Style Transfer
by: Wang, Zijia, et al.
Published: (2024)
by: Wang, Zijia, et al.
Published: (2024)
Exploring the Low-Pass Filtering Behavior in Image Super-Resolution
by: Deng, Haoyu, et al.
Published: (2024)
by: Deng, Haoyu, et al.
Published: (2024)
Text-Guided Molecule Generation with Diffusion Language Model
by: Gong, Haisong, et al.
Published: (2024)
by: Gong, Haisong, et al.
Published: (2024)
Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
by: Han, Yuhang, et al.
Published: (2024)
by: Han, Yuhang, et al.
Published: (2024)
Revisiting MLLM Based Image Quality Assessment: Errors and Remedy
by: Tang, Zhenchen, et al.
Published: (2025)
by: Tang, Zhenchen, et al.
Published: (2025)
Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation
by: He, Yiguo, et al.
Published: (2025)
by: He, Yiguo, et al.
Published: (2025)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
Learned Image Compression with Text Quality Enhancement
by: Lai, Chih-Yu, et al.
Published: (2024)
by: Lai, Chih-Yu, et al.
Published: (2024)
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
by: Su, Tongtong, et al.
Published: (2025)
by: Su, Tongtong, et al.
Published: (2025)
Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting
by: Zhao, Haoyu, et al.
Published: (2024)
by: Zhao, Haoyu, et al.
Published: (2024)
UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification
by: Fan, Qihang, et al.
Published: (2026)
by: Fan, Qihang, et al.
Published: (2026)
MLLM-as-a-Judge for Image Safety without Human Labeling
by: Wang, Zhenting, et al.
Published: (2024)
by: Wang, Zhenting, et al.
Published: (2024)
Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval
by: Liu, Delong, et al.
Published: (2023)
by: Liu, Delong, et al.
Published: (2023)
Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
by: Negoita, Vlad, et al.
Published: (2025)
by: Negoita, Vlad, et al.
Published: (2025)
Towards Any-Quality Image Segmentation via Generative and Adaptive Latent Space Enhancement
by: Guo, Guangqian, et al.
Published: (2026)
by: Guo, Guangqian, et al.
Published: (2026)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
Text-only Synthesis for Image Captioning
by: Zhou, Qing, et al.
Published: (2024)
by: Zhou, Qing, et al.
Published: (2024)
Multi-MLLM Knowledge Distillation for Out-of-Context News Detection
by: Gu, Yimeng, et al.
Published: (2025)
by: Gu, Yimeng, et al.
Published: (2025)
Personalized Text Generation with Contrastive Activation Steering
by: Zhang, Jinghao, et al.
Published: (2025)
by: Zhang, Jinghao, et al.
Published: (2025)
MT$^{3}$: Scaling MLLM-based Text Image Machine Translation via Multi-Task Reinforcement Learning
by: Feng, Zhaopeng, et al.
Published: (2025)
by: Feng, Zhaopeng, et al.
Published: (2025)
TIPS: Text-Image Pretraining with Spatial awareness
by: Maninis, Kevis-Kokitsi, et al.
Published: (2024)
by: Maninis, Kevis-Kokitsi, et al.
Published: (2024)
NAG: A Unified Native Architecture for Encoder-free Text-Graph Modeling in Language Models
by: Gong, Haisong, et al.
Published: (2026)
by: Gong, Haisong, et al.
Published: (2026)
Elysium: Exploring Object-level Perception in Videos via MLLM
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Mask-ControlNet: Higher-Quality Image Generation with An Additional Mask Prompt
by: Huang, Zhiqi, et al.
Published: (2024)
by: Huang, Zhiqi, et al.
Published: (2024)
Similar Items
-
Efficient Motion-Aware Video MLLM
by: Zhao, Zijia, et al.
Published: (2025) -
Exploring the Design Space of Visual Context Representation in Video MLLMs
by: Du, Yifan, et al.
Published: (2024) -
Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
by: Zhao, Zijia, et al.
Published: (2024) -
Towards Event-oriented Long Video Understanding
by: Du, Yifan, et al.
Published: (2024) -
Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
by: Du, Yifan, et al.
Published: (2025)