Data Curation Matters: Model Collapse and Spurious Shift Performance Prediction from Training on Uncurated Text Embeddings
Fuente:
arXiv
Guardado en:
| Autores principales: | Mattioli, Lucas, Hadichou, Youness Ait, Chaouche, Sabrina, Gonzalez, Martin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LACON: Training Text-to-Image Model from Uncurated Data
por: Liang, Zhiyang, et al.
Publicado: (2026)
por: Liang, Zhiyang, et al.
Publicado: (2026)
Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations
por: Tamayo-Rousseau, Camilo, et al.
Publicado: (2025)
por: Tamayo-Rousseau, Camilo, et al.
Publicado: (2025)
Training on Plausible Counterfactuals Removes Spurious Correlations
por: Sadiku, Shpresim, et al.
Publicado: (2025)
por: Sadiku, Shpresim, et al.
Publicado: (2025)
Bridging Explainability and Embeddings: BEE Aware of Spuriousness
por: Păduraru, Cristian Daniel, et al.
Publicado: (2024)
por: Păduraru, Cristian Daniel, et al.
Publicado: (2024)
Complexity Matters: Dynamics of Feature Learning in the Presence of Spurious Correlations
por: Qiu, GuanWen, et al.
Publicado: (2024)
por: Qiu, GuanWen, et al.
Publicado: (2024)
Autoguided Online Data Curation for Diffusion Model Training
por: Pais, Valeria, et al.
Publicado: (2025)
por: Pais, Valeria, et al.
Publicado: (2025)
Spurious Rewards: Rethinking Training Signals in RLVR
por: Shao, Rulin, et al.
Publicado: (2025)
por: Shao, Rulin, et al.
Publicado: (2025)
Spurious Correlation-Aware Embedding Regularization for Worst-Group Robustness
por: Park, Subeen, et al.
Publicado: (2025)
por: Park, Subeen, et al.
Publicado: (2025)
How to Synthesize Text Data without Model Collapse?
por: Zhu, Xuekai, et al.
Publicado: (2024)
por: Zhu, Xuekai, et al.
Publicado: (2024)
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
por: Falahati, Ali, et al.
Publicado: (2026)
por: Falahati, Ali, et al.
Publicado: (2026)
Severing Spurious Correlations with Data Pruning
por: Mulchandani, Varun, et al.
Publicado: (2025)
por: Mulchandani, Varun, et al.
Publicado: (2025)
Mitigating Spurious Correlations in LLMs via Causality-Aware Post-Training
por: Gui, Shurui, et al.
Publicado: (2025)
por: Gui, Shurui, et al.
Publicado: (2025)
Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models
por: Na, Byeonghu, et al.
Publicado: (2025)
por: Na, Byeonghu, et al.
Publicado: (2025)
Self-Consuming Generative Models with Adversarially Curated Data
por: Wei, Xiukun, et al.
Publicado: (2025)
por: Wei, Xiukun, et al.
Publicado: (2025)
Monitoring Neural Training with Topology: A Footprint-Predictable Collapse Index
por: Kalinowski, Alexander
Publicado: (2026)
por: Kalinowski, Alexander
Publicado: (2026)
RiverText: A Python Library for Training and Evaluating Incremental Word Embeddings from Text Data Streams
por: Iturra-Bocaz, Gabriel, et al.
Publicado: (2025)
por: Iturra-Bocaz, Gabriel, et al.
Publicado: (2025)
Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models
por: Hu, Zizhao, et al.
Publicado: (2025)
por: Hu, Zizhao, et al.
Publicado: (2025)
Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
por: Wang, Jiachen T., et al.
Publicado: (2025)
por: Wang, Jiachen T., et al.
Publicado: (2025)
Spurious Forgetting in Continual Learning of Language Models
por: Zheng, Junhao, et al.
Publicado: (2025)
por: Zheng, Junhao, et al.
Publicado: (2025)
Curation Leaks: Membership Inference Attacks against Data Curation for Machine Learning
por: Wahdany, Dariush, et al.
Publicado: (2026)
por: Wahdany, Dariush, et al.
Publicado: (2026)
Rate of Model Collapse in Recursive Training
por: Suresh, Ananda Theertha, et al.
Publicado: (2024)
por: Suresh, Ananda Theertha, et al.
Publicado: (2024)
On the Embedding Collapse when Scaling up Recommendation Models
por: Guo, Xingzhuo, et al.
Publicado: (2023)
por: Guo, Xingzhuo, et al.
Publicado: (2023)
The Pragmatic Frames of Spurious Correlations in Machine Learning: Interpreting How and Why They Matter
por: Bell, Samuel J., et al.
Publicado: (2024)
por: Bell, Samuel J., et al.
Publicado: (2024)
Freeze then Train: Towards Provable Representation Learning under Spurious Correlations and Feature Noise
por: Ye, Haotian, et al.
Publicado: (2022)
por: Ye, Haotian, et al.
Publicado: (2022)
LoRA Training in the NTK Regime has No Spurious Local Minima
por: Jang, Uijeong, et al.
Publicado: (2024)
por: Jang, Uijeong, et al.
Publicado: (2024)
Embedding Hidden Adversarial Capabilities in Pre-Trained Diffusion Models
por: Beerens, Lucas, et al.
Publicado: (2025)
por: Beerens, Lucas, et al.
Publicado: (2025)
Scaling with Collapse: Efficient and Predictable Training of LLM Families
por: Bergsma, Shane, et al.
Publicado: (2025)
por: Bergsma, Shane, et al.
Publicado: (2025)
Escaping Collapse: The Strength of Weak Data for Large Language Model Training
por: Amin, Kareem, et al.
Publicado: (2025)
por: Amin, Kareem, et al.
Publicado: (2025)
PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding Projection
por: Molahasani, Mahdiyar, et al.
Publicado: (2025)
por: Molahasani, Mahdiyar, et al.
Publicado: (2025)
TextAge: A Curated and Diverse Text Dataset for Age Classification
por: Cheekati, Shravan, et al.
Publicado: (2024)
por: Cheekati, Shravan, et al.
Publicado: (2024)
TopoCurate:Modeling Interaction Topology for Tool-Use Agent Training
por: Yang, Jinluan, et al.
Publicado: (2026)
por: Yang, Jinluan, et al.
Publicado: (2026)
Fighting Spurious Correlations in Text Classification via a Causal Learning Perspective
por: Zhou, Yuqing, et al.
Publicado: (2024)
por: Zhou, Yuqing, et al.
Publicado: (2024)
SEED: Domain-Specific Data Curation With Large Language Models
por: Chen, Zui, et al.
Publicado: (2023)
por: Chen, Zui, et al.
Publicado: (2023)
Not All Invariants Are Equal: Curating Training Data to Accelerate Program Verification with SLMs
por: Pinto, Ido, et al.
Publicado: (2026)
por: Pinto, Ido, et al.
Publicado: (2026)
Out of Spuriousity: Improving Robustness to Spurious Correlations without Group Annotations
por: Le, Phuong Quynh, et al.
Publicado: (2024)
por: Le, Phuong Quynh, et al.
Publicado: (2024)
CaTE Data Curation for Trustworthy AI
por: Clemens-Sewall, Mary Versa, et al.
Publicado: (2025)
por: Clemens-Sewall, Mary Versa, et al.
Publicado: (2025)
Verified Training for Counterfactual Explanation Robustness under Data Shift
por: Meyer, Anna P., et al.
Publicado: (2024)
por: Meyer, Anna P., et al.
Publicado: (2024)
Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection
por: Fan, Ziqing, et al.
Publicado: (2025)
por: Fan, Ziqing, et al.
Publicado: (2025)
Reassessing the Validity of Spurious Correlations Benchmarks
por: Bell, Samuel J., et al.
Publicado: (2024)
por: Bell, Samuel J., et al.
Publicado: (2024)
Spurious Privacy Leakage in Neural Networks
por: Zhang, Chenxiang, et al.
Publicado: (2025)
por: Zhang, Chenxiang, et al.
Publicado: (2025)
Ejemplares similares
-
LACON: Training Text-to-Image Model from Uncurated Data
por: Liang, Zhiyang, et al.
Publicado: (2026) -
Your Attention Matters: to Improve Model Robustness to Noise and Spurious Correlations
por: Tamayo-Rousseau, Camilo, et al.
Publicado: (2025) -
Training on Plausible Counterfactuals Removes Spurious Correlations
por: Sadiku, Shpresim, et al.
Publicado: (2025) -
Bridging Explainability and Embeddings: BEE Aware of Spuriousness
por: Păduraru, Cristian Daniel, et al.
Publicado: (2024) -
Complexity Matters: Dynamics of Feature Learning in the Presence of Spurious Correlations
por: Qiu, GuanWen, et al.
Publicado: (2024)