MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Bingbing, Salekin, Sirajul, Kang, Feiyang, Howe, Bill, Wang, Lucy Lu, Movellan, Javier, Bilkhu, Manjot |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
by: Huang, Tzu-Heng, et al.
Published: (2026)
by: Huang, Tzu-Heng, et al.
Published: (2026)
Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights
by: Huang, Tzu-Heng, et al.
Published: (2025)
by: Huang, Tzu-Heng, et al.
Published: (2025)
Characterizing LLM Abstention Behavior in Science QA with Context Perturbations
by: Wen, Bingbing, et al.
Published: (2024)
by: Wen, Bingbing, et al.
Published: (2024)
STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes
by: Xu, Chenjun, et al.
Published: (2026)
by: Xu, Chenjun, et al.
Published: (2026)
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
by: Xu, Chenjun, et al.
Published: (2025)
by: Xu, Chenjun, et al.
Published: (2025)
Know Your Limits: A Survey of Abstention in Large Language Models
by: Wen, Bingbing, et al.
Published: (2024)
by: Wen, Bingbing, et al.
Published: (2024)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
by: Kang, Feiyang, et al.
Published: (2024)
by: Kang, Feiyang, et al.
Published: (2024)
Global distributions of modern planktic foraminifera abundance and biomass - Gridded data product (NetCDF) - Contribution to the MAREDAT World Ocean Atlas of Plankton Functional Types
by: Schiebel, Ralf, et al.
Published: (2012)
by: Schiebel, Ralf, et al.
Published: (2012)
Midtraining Bridges Pretraining and Posttraining Distributions
by: Liu, Emmy, et al.
Published: (2025)
by: Liu, Emmy, et al.
Published: (2025)
Are Data Experts Buying into Differentially Private Synthetic Data? Gathering Community Perspectives
by: Rosenblatt, Lucas, et al.
Published: (2024)
by: Rosenblatt, Lucas, et al.
Published: (2024)
Model Spec Midtraining: Improving How Alignment Training Generalizes
by: Li, Chloe, et al.
Published: (2026)
by: Li, Chloe, et al.
Published: (2026)
Japón, España y los debates actuales sobre la memoria traumática del siglo XX: una propuesta de análisis histórico comparado
by: Jesús Movellán Haro
Published: (2024)
by: Jesús Movellán Haro
Published: (2024)
Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification
by: Cao, Zongwan, et al.
Published: (2026)
by: Cao, Zongwan, et al.
Published: (2026)
SARN: Structurally-Aware Recurrent Network for Spatio-Temporal Disaggregation
by: Han, Bin, et al.
Published: (2023)
by: Han, Bin, et al.
Published: (2023)
Counter-example to continuity of measure in uncountable unions
by: Bilkhu, Simranjeet, et al.
Published: (2025)
by: Bilkhu, Simranjeet, et al.
Published: (2025)
OPCap:Object-aware Prompting Captioning
by: Huang, Feiyang
Published: (2024)
by: Huang, Feiyang
Published: (2024)
Estorbos a un regidor advenedizo: justicia, facciones y conflicto urbano en la España del siglo XVII
by: Tomás A. Mantecón Movellán
Published: (2020)
by: Tomás A. Mantecón Movellán
Published: (2020)
Cencerradas, cultura moral campesina y disciplinamiento social en la España del Antiguo Régimen
by: Tomás A. Mantecón Movellán
Published: (2013)
by: Tomás A. Mantecón Movellán
Published: (2013)
Epistemic Alignment: A Mediating Framework for User-LLM Knowledge Delivery
by: Clark, Nicholas, et al.
Published: (2025)
by: Clark, Nicholas, et al.
Published: (2025)
Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training
by: Zhang, Mozhi, et al.
Published: (2025)
by: Zhang, Mozhi, et al.
Published: (2025)
ViTOC: Vision Transformer and Object-aware Captioner
by: Huang, Feiyang
Published: (2024)
by: Huang, Feiyang
Published: (2024)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
Development of the African Marine Atlas
by: Scott, Lucy
Published: (2008)
by: Scott, Lucy
Published: (2008)
Can Large Language Models Integrate Spatial Data? Empirical Insights into Reasoning Strengths and Computational Weaknesses
by: Han, Bin, et al.
Published: (2025)
by: Han, Bin, et al.
Published: (2025)
Skin Cancer Images Classification using Transfer Learning Techniques
by: Islam, Md Sirajul, et al.
Published: (2024)
by: Islam, Md Sirajul, et al.
Published: (2024)
Effects of salinity and temperature on the larval development of a sesarmid crab Neosarmatium trispinosum Davie (Crustacea: Brachyura: Sesarmidae) from mangrove swamp in Okinawa Island, Japan
by: Sirajul Islam, Md., et al.
Published: (2003)
by: Sirajul Islam, Md., et al.
Published: (2003)
Reliable, Routable, and Reproducible: Collection of Pedestrian Pathways at Statewide Scale
by: Zhang, Yuxiang, et al.
Published: (2024)
by: Zhang, Yuxiang, et al.
Published: (2024)
ML-EAT: A Multilevel Embedding Association Test for Interpretable and Transparent Social Science
by: Wolfe, Robert, et al.
Published: (2024)
by: Wolfe, Robert, et al.
Published: (2024)
Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?
by: Yang, Yiwei, et al.
Published: (2025)
by: Yang, Yiwei, et al.
Published: (2025)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
by: Ye, Jiasheng, et al.
Published: (2024)
by: Ye, Jiasheng, et al.
Published: (2024)
Optimizing Pretraining Data Mixtures with LLM-Estimated Utility
by: Held, William, et al.
Published: (2025)
by: Held, William, et al.
Published: (2025)
Re-Mix: Optimizing Data Mixtures for Large Scale Imitation Learning
by: Hejna, Joey, et al.
Published: (2024)
by: Hejna, Joey, et al.
Published: (2024)
Expressivity of Spiking Neural Networks
by: Singh, Manjot, et al.
Published: (2023)
by: Singh, Manjot, et al.
Published: (2023)
SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards
by: Zhang, Dengjia, et al.
Published: (2026)
by: Zhang, Dengjia, et al.
Published: (2026)
FedClust: Optimizing Federated Learning on Non-IID Data through Weight-Driven Client Clustering
by: Islam, Md Sirajul, et al.
Published: (2024)
by: Islam, Md Sirajul, et al.
Published: (2024)
CurvFed: Curvature-Aligned Federated Learning for Fairness without Demographics
by: Sharma, Harshit, et al.
Published: (2024)
by: Sharma, Harshit, et al.
Published: (2024)
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
by: Wang, Zhenyu, et al.
Published: (2026)
by: Wang, Zhenyu, et al.
Published: (2026)
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
by: Du, Yiyang, et al.
Published: (2026)
by: Du, Yiyang, et al.
Published: (2026)
DAOpt: Modeling and Evaluation of Data-Driven Optimization under Uncertainty with LLMs
by: Zhu, WenZhuo, et al.
Published: (2025)
by: Zhu, WenZhuo, et al.
Published: (2025)
LaiDA: Linguistics-aware In-context Learning with Data Augmentation for Metaphor Components Identification
by: Liu, Hongde, et al.
Published: (2024)
by: Liu, Hongde, et al.
Published: (2024)
Similar Items
-
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
by: Huang, Tzu-Heng, et al.
Published: (2026) -
Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights
by: Huang, Tzu-Heng, et al.
Published: (2025) -
Characterizing LLM Abstention Behavior in Science QA with Context Perturbations
by: Wen, Bingbing, et al.
Published: (2024) -
STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes
by: Xu, Chenjun, et al.
Published: (2026) -
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
by: Xu, Chenjun, et al.
Published: (2025)