Train a Unified Multimodal Data Quality Classifier with Synthetic Data
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Weizhi, Lin, Rongmei, Li, Shiyang, Lockard, Colin, Sarkhel, Ritesh, Lokegaonkar, Sanket, Shang, Jingbo, Yan, Xifeng, Zalmout, Nasser, Li, Xian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer
by: Deng, Chunyuan, et al.
Published: (2026)
by: Deng, Chunyuan, et al.
Published: (2026)
DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities
by: Lee, Jing Yang, et al.
Published: (2025)
by: Lee, Jing Yang, et al.
Published: (2025)
Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters
by: Wang, Weizhi, et al.
Published: (2024)
by: Wang, Weizhi, et al.
Published: (2024)
Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
by: Wang, Weizhi, et al.
Published: (2025)
by: Wang, Weizhi, et al.
Published: (2025)
Adaptive Layer-skipping in Pre-trained LLMs
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
Direct Multi-Token Decoding
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
Incubating Text Classifiers Following User Instruction with Nothing but LLM
by: Peng, Letian, et al.
Published: (2024)
by: Peng, Letian, et al.
Published: (2024)
Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training
by: Zhuang, Yuchen, et al.
Published: (2025)
by: Zhuang, Yuchen, et al.
Published: (2025)
From Measurement Instruments to Data: Leveraging Theory-Driven Synthetic Training Data for Classifying Social Constructs
by: Birkenmaier, Lukas, et al.
Published: (2024)
by: Birkenmaier, Lukas, et al.
Published: (2024)
Bot or Human? Detecting ChatGPT Imposters with A Single Question
by: Wang, Hong, et al.
Published: (2023)
by: Wang, Hong, et al.
Published: (2023)
Smaller Language Models are capable of selecting Instruction-Tuning Training Data for Larger Language Models
by: Mekala, Dheeraj, et al.
Published: (2024)
by: Mekala, Dheeraj, et al.
Published: (2024)
Analysis of Classifier Training on Synthetic Data for Cross-Domain Datasets
by: Cortés, Andoni, et al.
Published: (2024)
by: Cortés, Andoni, et al.
Published: (2024)
Multimodal Misinformation Detection by Learning from Synthetic Data with Multimodal LLMs
by: Zeng, Fengzhu, et al.
Published: (2024)
by: Zeng, Fengzhu, et al.
Published: (2024)
DOCMASTER: A Unified Platform for Annotation, Training, & Inference in Document Question-Answering
by: Nguyen, Alex, et al.
Published: (2024)
by: Nguyen, Alex, et al.
Published: (2024)
Multimodal Policy Internalization for Conversational Agents
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
Unifying Structured Data as Graph for Data-to-Text Pre-Training
by: Li, Shujie, et al.
Published: (2024)
by: Li, Shujie, et al.
Published: (2024)
Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data
by: Li, Haoxin, et al.
Published: (2025)
by: Li, Haoxin, et al.
Published: (2025)
Multimodal Language Models with Modality-Specific Experts for Financial Forecasting from Interleaved Sequences of Text and Time Series
by: Koval, Ross, et al.
Published: (2025)
by: Koval, Ross, et al.
Published: (2025)
Controllable Data Augmentation for Few-Shot Text Mining with Chain-of-Thought Attribute Manipulation
by: Peng, Letian, et al.
Published: (2023)
by: Peng, Letian, et al.
Published: (2023)
A Novel Taxonomy for Navigating and Classifying Synthetic Data in Healthcare Applications
by: van Dijk, Bram, et al.
Published: (2024)
by: van Dijk, Bram, et al.
Published: (2024)
AlpaGasus: Training A Better Alpaca with Fewer Data
by: Chen, Lichang, et al.
Published: (2023)
by: Chen, Lichang, et al.
Published: (2023)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
by: Rahman, Tanzila, et al.
Published: (2026)
by: Rahman, Tanzila, et al.
Published: (2026)
PIGUIQA: A Physical Imaging Guided Perceptual Framework for Underwater Image Quality Assessment
by: Xian, Weizhi, et al.
Published: (2024)
by: Xian, Weizhi, et al.
Published: (2024)
Noise-Aware Training of Layout-Aware Language Models
by: Sarkhel, Ritesh, et al.
Published: (2024)
by: Sarkhel, Ritesh, et al.
Published: (2024)
Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
MDSF: Context-Aware Multi-Dimensional Data Storytelling Framework based on Large language Model
by: Zhang, Chengze, et al.
Published: (2025)
by: Zhang, Chengze, et al.
Published: (2025)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
by: Zheng, Ruobing, et al.
Published: (2026)
by: Zheng, Ruobing, et al.
Published: (2026)
Data Diversity Matters for Robust Instruction Tuning
by: Bukharin, Alexander, et al.
Published: (2023)
by: Bukharin, Alexander, et al.
Published: (2023)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
Data Debugging is NP-hard for Classifiers Trained with SGD
by: Guo, Zizheng, et al.
Published: (2024)
by: Guo, Zizheng, et al.
Published: (2024)
On the Diversity of Synthetic Data and its Impact on Training Large Language Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
RNR: Teaching Large Language Models to Follow Roles and Rules
by: Wang, Kuan, et al.
Published: (2024)
by: Wang, Kuan, et al.
Published: (2024)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
Disentangling Fine-Tuning from Pre-Training in Visual Captioning with Hybrid Markov Logic
by: Shah, Monika, et al.
Published: (2025)
by: Shah, Monika, et al.
Published: (2025)
Nimbus: A Unified Embodied Synthetic Data Generation Framework
by: He, Zeyu, et al.
Published: (2026)
by: He, Zeyu, et al.
Published: (2026)
What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?
by: Yan, David, et al.
Published: (2025)
by: Yan, David, et al.
Published: (2025)
MEMORYLLM: Towards Self-Updatable Large Language Models
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Training Language Models to Generate Quality Code with Program Analysis Feedback
by: Yao, Feng, et al.
Published: (2025)
by: Yao, Feng, et al.
Published: (2025)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
SynPlanResearch-R1: Encouraging Tool Exploration for Deep Research with Synthetic Plans
by: Zeng, Hansi, et al.
Published: (2026)
by: Zeng, Hansi, et al.
Published: (2026)
Similar Items
-
ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer
by: Deng, Chunyuan, et al.
Published: (2026) -
DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities
by: Lee, Jing Yang, et al.
Published: (2025) -
Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters
by: Wang, Weizhi, et al.
Published: (2024) -
Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
by: Wang, Weizhi, et al.
Published: (2025) -
Adaptive Layer-skipping in Pre-trained LLMs
by: Luo, Xuan, et al.
Published: (2025)