Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Djuhera, Aladin, Kadhe, Swanand Ravindra, Zawad, Syed, Ahmed, Farhan, Ludwig, Heiko, Boche, Holger |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
by: Djuhera, Aladin, et al.
Published: (2026)
by: Djuhera, Aladin, et al.
Published: (2026)
SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
Against the Monolithic Wireless World Model: Why NextG Needs Composable and Agentic Intelligence
by: Djuhera, Aladin, et al.
Published: (2026)
by: Djuhera, Aladin, et al.
Published: (2026)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
MambaCSP: Hybrid-Attention State Space Models for Hardware-Efficient Channel State Prediction
by: Djuhera, Aladin, et al.
Published: (2026)
by: Djuhera, Aladin, et al.
Published: (2026)
"Don't Do That!": Guiding Embodied Systems through Large Language Model-based Constraint Generation
by: Seffo, Amin, et al.
Published: (2025)
by: Seffo, Amin, et al.
Published: (2025)
Turning Generative Models Degenerate: The Power of Data Poisoning Attacks
by: Jiang, Shuli, et al.
Published: (2024)
by: Jiang, Shuli, et al.
Published: (2024)
STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs
by: An, Sungeun, et al.
Published: (2026)
by: An, Sungeun, et al.
Published: (2026)
Know When To Fold 'Em: Token-Efficient LLM Synthetic Data Generation via Multi-Stage In-Flight Rejection
by: Chowdhury, Anjir Ahmed, et al.
Published: (2026)
by: Chowdhury, Anjir Ahmed, et al.
Published: (2026)
SCoTT: Strategic Chain-of-Thought Tasking for Wireless-Aware Robot Navigation in Digital Twins
by: Djuhera, Aladin, et al.
Published: (2024)
by: Djuhera, Aladin, et al.
Published: (2024)
R-MTLLMF: Resilient Multi-Task Large Language Model Fusion at the Wireless Edge
by: Djuhera, Aladin, et al.
Published: (2024)
by: Djuhera, Aladin, et al.
Published: (2024)
GneissWeb: Preparing High Quality Data for LLMs at Scale
by: Gohari, Hajar Emami, et al.
Published: (2025)
by: Gohari, Hajar Emami, et al.
Published: (2025)
R-SFLLM: Jamming Resilient Framework for Split Federated Learning with Large Language Models
by: Djuhera, Aladin, et al.
Published: (2024)
by: Djuhera, Aladin, et al.
Published: (2024)
Evaluating the Dynamics of Membership Privacy in Deep Learning
by: Chen, Yuetian, et al.
Published: (2025)
by: Chen, Yuetian, et al.
Published: (2025)
Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle
by: Haider, Emman, et al.
Published: (2024)
by: Haider, Emman, et al.
Published: (2024)
SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
by: Alrashed, Sultan, et al.
Published: (2025)
by: Alrashed, Sultan, et al.
Published: (2025)
PEML: Parameter-efficient Multi-Task Learning with Optimized Continuous Prompts
by: Chowdhury, Anjir Ahmed, et al.
Published: (2026)
by: Chowdhury, Anjir Ahmed, et al.
Published: (2026)
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents
by: Ngong, Ivoline, et al.
Published: (2025)
by: Ngong, Ivoline, et al.
Published: (2025)
Post-Training Language Models for Crosslingual Consistency
by: Liu, Tianyu, et al.
Published: (2026)
by: Liu, Tianyu, et al.
Published: (2026)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
by: Finlayson, Matthew, et al.
Published: (2025)
by: Finlayson, Matthew, et al.
Published: (2025)
Enhancing Depressive Post Detection in Bangla: A Comparative Study of TF-IDF, BERT and FastText Embeddings
by: Sazan, Saad Ahmed, et al.
Published: (2024)
by: Sazan, Saad Ahmed, et al.
Published: (2024)
Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
by: Rachamalla, Neel Prabhanjan, et al.
Published: (2025)
by: Rachamalla, Neel Prabhanjan, et al.
Published: (2025)
Synthesizing Post-Training Data for LLMs through Multi-Agent Simulation
by: Tang, Shuo, et al.
Published: (2024)
by: Tang, Shuo, et al.
Published: (2024)
An Analysis of Capacity-Distortion Trade-Offs in Memoryless ISAC Systems
by: Li, Xinyang, et al.
Published: (2024)
by: Li, Xinyang, et al.
Published: (2024)
The Limits of Preference Data for Post-Training
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
by: Liu, Yixin, et al.
Published: (2026)
by: Liu, Yixin, et al.
Published: (2026)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
by: Yuan, Yurun, et al.
Published: (2026)
by: Yuan, Yurun, et al.
Published: (2026)
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
by: Bobbili, Sarat Chandra, et al.
Published: (2025)
by: Bobbili, Sarat Chandra, et al.
Published: (2025)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
by: Zhou, Jin Peng, et al.
Published: (2025)
by: Zhou, Jin Peng, et al.
Published: (2025)
Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
by: Gan, Zeyu, et al.
Published: (2024)
by: Gan, Zeyu, et al.
Published: (2024)
LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agents
by: Yano, Taro, et al.
Published: (2025)
by: Yano, Taro, et al.
Published: (2025)
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training
by: Dhaliwal, Mehak, et al.
Published: (2026)
by: Dhaliwal, Mehak, et al.
Published: (2026)
Aggressive Post-Training Compression on Extremely Large Language Models
by: Zhang, Zining, et al.
Published: (2024)
by: Zhang, Zining, et al.
Published: (2024)
Assessing Robustness to Spurious Correlations in Post-Training Language Models
by: Shuieh, Julia, et al.
Published: (2025)
by: Shuieh, Julia, et al.
Published: (2025)
BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data
by: Zou, Bo, et al.
Published: (2026)
by: Zou, Bo, et al.
Published: (2026)
What Is The Political Content in LLMs' Pre- and Post-Training Data?
by: Ceron, Tanise, et al.
Published: (2025)
by: Ceron, Tanise, et al.
Published: (2025)
Similar Items
-
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
by: Djuhera, Aladin, et al.
Published: (2025) -
TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
by: Djuhera, Aladin, et al.
Published: (2026) -
SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging
by: Djuhera, Aladin, et al.
Published: (2025) -
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
by: Djuhera, Aladin, et al.
Published: (2025) -
Against the Monolithic Wireless World Model: Why NextG Needs Composable and Agentic Intelligence
by: Djuhera, Aladin, et al.
Published: (2026)