Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Daoyuan, Wang, Haibin, Huang, Yilun, Ge, Ce, Li, Yaliang, Ding, Bolin, Zhou, Jingren |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
di: Jiao, Qirui, et al.
Pubblicazione: (2024)
di: Jiao, Qirui, et al.
Pubblicazione: (2024)
The Synergy between Data and Multi-Modal Large Language Models: A Survey from Co-Development Perspective
di: Qin, Zhen, et al.
Pubblicazione: (2024)
di: Qin, Zhen, et al.
Pubblicazione: (2024)
Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models
di: Chen, Daoyuan, et al.
Pubblicazione: (2024)
di: Chen, Daoyuan, et al.
Pubblicazione: (2024)
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
di: Zhou, Ting, et al.
Pubblicazione: (2024)
di: Zhou, Ting, et al.
Pubblicazione: (2024)
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information
di: Jiao, Qirui, et al.
Pubblicazione: (2024)
di: Jiao, Qirui, et al.
Pubblicazione: (2024)
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
di: Jiao, Qirui, et al.
Pubblicazione: (2025)
di: Jiao, Qirui, et al.
Pubblicazione: (2025)
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
di: Ge, Ce, et al.
Pubblicazione: (2024)
di: Ge, Ce, et al.
Pubblicazione: (2024)
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
di: Li, Yuyi, et al.
Pubblicazione: (2025)
di: Li, Yuyi, et al.
Pubblicazione: (2025)
AttriCtrl: Fine-Grained Control of Aesthetic Attribute Intensity in Diffusion Models
di: Chen, Die, et al.
Pubblicazione: (2025)
di: Chen, Die, et al.
Pubblicazione: (2025)
AutoLoRA: Automatic LoRA Retrieval and Fine-Grained Gated Fusion for Text-to-Image Generation
di: Li, Zhiwen, et al.
Pubblicazione: (2025)
di: Li, Zhiwen, et al.
Pubblicazione: (2025)
VIRAL: Visual In-Context Reasoning via Analogy in Diffusion Transformers
di: Li, Zhiwen, et al.
Pubblicazione: (2026)
di: Li, Zhiwen, et al.
Pubblicazione: (2026)
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
di: Duan, Zhongjie, et al.
Pubblicazione: (2024)
di: Duan, Zhongjie, et al.
Pubblicazione: (2024)
Data Processing Techniques for Modern Multimodal Models
di: Li, Yinheng, et al.
Pubblicazione: (2024)
di: Li, Yinheng, et al.
Pubblicazione: (2024)
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
di: Xu, Zhe, et al.
Pubblicazione: (2025)
di: Xu, Zhe, et al.
Pubblicazione: (2025)
Spectral Evolution Search: Efficient Inference-Time Scaling for Reward-Aligned Image Generation
di: Ye, Jinyan, et al.
Pubblicazione: (2026)
di: Ye, Jinyan, et al.
Pubblicazione: (2026)
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers
di: Chen, Dongping, et al.
Pubblicazione: (2026)
di: Chen, Dongping, et al.
Pubblicazione: (2026)
Modeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving
di: Zhou, Yuhan, et al.
Pubblicazione: (2026)
di: Zhou, Yuhan, et al.
Pubblicazione: (2026)
HoloDx: Knowledge- and Data-Driven Multimodal Diagnosis of Alzheimer's Disease
di: Chen, Qiuhui, et al.
Pubblicazione: (2025)
di: Chen, Qiuhui, et al.
Pubblicazione: (2025)
BOTS: A Unified Framework for Bayesian Online Task Selection in LLM Reinforcement Finetuning
di: Shen, Qianli, et al.
Pubblicazione: (2025)
di: Shen, Qianli, et al.
Pubblicazione: (2025)
ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning
di: Duan, Zhongjie, et al.
Pubblicazione: (2024)
di: Duan, Zhongjie, et al.
Pubblicazione: (2024)
PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection
di: Sun, Haibin, et al.
Pubblicazione: (2025)
di: Sun, Haibin, et al.
Pubblicazione: (2025)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
di: Wen, Xin, et al.
Pubblicazione: (2025)
di: Wen, Xin, et al.
Pubblicazione: (2025)
Backdooring Vision-Language Models with Out-Of-Distribution Data
di: Lyu, Weimin, et al.
Pubblicazione: (2024)
di: Lyu, Weimin, et al.
Pubblicazione: (2024)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
di: Song, Tingyu, et al.
Pubblicazione: (2025)
di: Song, Tingyu, et al.
Pubblicazione: (2025)
RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data
di: Chen, Harold Haodong, et al.
Pubblicazione: (2026)
di: Chen, Harold Haodong, et al.
Pubblicazione: (2026)
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
di: Xu, Guixian, et al.
Pubblicazione: (2026)
di: Xu, Guixian, et al.
Pubblicazione: (2026)
LGTM: Local-to-Global Text-Driven Human Motion Diffusion Model
di: Sun, Haowen, et al.
Pubblicazione: (2024)
di: Sun, Haowen, et al.
Pubblicazione: (2024)
Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding
di: Zhang, Lina, et al.
Pubblicazione: (2026)
di: Zhang, Lina, et al.
Pubblicazione: (2026)
Less is More: High-value Data Selection for Visual Instruction Tuning
di: Liu, Zikang, et al.
Pubblicazione: (2024)
di: Liu, Zikang, et al.
Pubblicazione: (2024)
WorldMark: A Unified Benchmark Suite for Interactive Video World Models
di: Xu, Xiaojie, et al.
Pubblicazione: (2026)
di: Xu, Xiaojie, et al.
Pubblicazione: (2026)
CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models
di: Zhou, Nan, et al.
Pubblicazione: (2026)
di: Zhou, Nan, et al.
Pubblicazione: (2026)
Growth Inhibitors for Suppressing Inappropriate Image Concepts in Diffusion Models
di: Chen, Die, et al.
Pubblicazione: (2024)
di: Chen, Die, et al.
Pubblicazione: (2024)
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
di: Huang, Ziqi, et al.
Pubblicazione: (2024)
di: Huang, Ziqi, et al.
Pubblicazione: (2024)
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
di: Ge, Yuying, et al.
Pubblicazione: (2024)
di: Ge, Yuying, et al.
Pubblicazione: (2024)
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
di: Ye, Junyan, et al.
Pubblicazione: (2024)
di: Ye, Junyan, et al.
Pubblicazione: (2024)
C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
di: Chen, Xiuwei, et al.
Pubblicazione: (2025)
di: Chen, Xiuwei, et al.
Pubblicazione: (2025)
ICANet: A Method of Short Video Emotion Recognition Driven by Multimodal Data
di: Wu, Xuecheng, et al.
Pubblicazione: (2022)
di: Wu, Xuecheng, et al.
Pubblicazione: (2022)
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
di: Chen, Zhe, et al.
Pubblicazione: (2024)
di: Chen, Zhe, et al.
Pubblicazione: (2024)
Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
di: Xia, Jiaer, et al.
Pubblicazione: (2025)
di: Xia, Jiaer, et al.
Pubblicazione: (2025)
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
di: Chen, Zeyu, et al.
Pubblicazione: (2026)
di: Chen, Zeyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
di: Jiao, Qirui, et al.
Pubblicazione: (2024) -
The Synergy between Data and Multi-Modal Large Language Models: A Survey from Co-Development Perspective
di: Qin, Zhen, et al.
Pubblicazione: (2024) -
Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models
di: Chen, Daoyuan, et al.
Pubblicazione: (2024) -
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
di: Zhou, Ting, et al.
Pubblicazione: (2024) -
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information
di: Jiao, Qirui, et al.
Pubblicazione: (2024)