The Price of Format: Diversity Collapse in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Yun, Longfei, An, Chenyang, Wang, Zilong, Peng, Letian, Shang, Jingbo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Tale of LLMs and Induced Small Proxies: Scalable Agents for Knowledge Mining
by: Zhang, Sipeng, et al.
Published: (2025)
by: Zhang, Sipeng, et al.
Published: (2025)
UltraGen: Extremely Fine-grained Controllable Generation via Attribute Reconstruction and Global Preference Optimization
by: Yun, Longfei, et al.
Published: (2025)
by: Yun, Longfei, et al.
Published: (2025)
Configurable Foundation Models: Building LLMs from a Modular Perspective
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
How to Synthesize Text Data without Model Collapse?
by: Zhu, Xuekai, et al.
Published: (2024)
by: Zhu, Xuekai, et al.
Published: (2024)
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Model-diff: A Tool for Comparative Study of Language Models in the Input Space
by: Liu, Weitang, et al.
Published: (2024)
by: Liu, Weitang, et al.
Published: (2024)
A Theory of Time-Sensitive Language Generation: Sparse Hallucination Beats Mode Collapse
by: Ganju, Atul, et al.
Published: (2026)
by: Ganju, Atul, et al.
Published: (2026)
Large Language Models for Time Series: A Survey
by: Zhang, Xiyuan, et al.
Published: (2024)
by: Zhang, Xiyuan, et al.
Published: (2024)
Learning a Decision Tree Algorithm with Transformers
by: Zhuang, Yufan, et al.
Published: (2024)
by: Zhuang, Yufan, et al.
Published: (2024)
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
by: Yang, Zhihe, et al.
Published: (2025)
by: Yang, Zhihe, et al.
Published: (2025)
Not All Layers of LLMs Are Necessary During Inference
by: Fan, Siqi, et al.
Published: (2024)
by: Fan, Siqi, et al.
Published: (2024)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
by: Wang, Chenyang, et al.
Published: (2025)
by: Wang, Chenyang, et al.
Published: (2025)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
by: Sanyal, Sunny, et al.
Published: (2024)
by: Sanyal, Sunny, et al.
Published: (2024)
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
by: Xiao, Boyu, et al.
Published: (2026)
by: Xiao, Boyu, et al.
Published: (2026)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
by: Liu, Wanlong, et al.
Published: (2024)
by: Liu, Wanlong, et al.
Published: (2024)
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
by: Dong, Yihong, et al.
Published: (2025)
by: Dong, Yihong, et al.
Published: (2025)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
by: Yang, Xin, et al.
Published: (2026)
by: Yang, Xin, et al.
Published: (2026)
FedALT: Federated Fine-Tuning through Adaptive Local Training with Rest-of-World LoRA
by: Bian, Jieming, et al.
Published: (2025)
by: Bian, Jieming, et al.
Published: (2025)
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
by: Kang, Yipeng, et al.
Published: (2024)
by: Kang, Yipeng, et al.
Published: (2024)
Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
by: Sahoo, Subramanyam
Published: (2026)
by: Sahoo, Subramanyam
Published: (2026)
HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
by: Li, Qing, et al.
Published: (2025)
by: Li, Qing, et al.
Published: (2025)
TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition
by: Kulkarni, Anay, et al.
Published: (2026)
by: Kulkarni, Anay, et al.
Published: (2026)
Foundations of Large Language Models
by: Xiao, Tong, et al.
Published: (2025)
by: Xiao, Tong, et al.
Published: (2025)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Epistemic Diversity and Knowledge Collapse in Large Language Models
by: Wright, Dustin, et al.
Published: (2025)
by: Wright, Dustin, et al.
Published: (2025)
Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
by: Zou, Heming, et al.
Published: (2025)
by: Zou, Heming, et al.
Published: (2025)
The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
by: Dickson, Craig
Published: (2025)
by: Dickson, Craig
Published: (2025)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step
by: Zhong, Li, et al.
Published: (2024)
by: Zhong, Li, et al.
Published: (2024)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
by: Liu, Mingjie, et al.
Published: (2025)
by: Liu, Mingjie, et al.
Published: (2025)
Correlation and Navigation in the Vocabulary Key Representation Space of Language Models
by: Peng, Letian, et al.
Published: (2024)
by: Peng, Letian, et al.
Published: (2024)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
by: Xiong, Boya, et al.
Published: (2025)
by: Xiong, Boya, et al.
Published: (2025)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
by: Zhou, Sifan, et al.
Published: (2025)
by: Zhou, Sifan, et al.
Published: (2025)
Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data
by: Ling, Zhenqing, et al.
Published: (2025)
by: Ling, Zhenqing, et al.
Published: (2025)
CorrSynth -- A Correlated Sampling Method for Diverse Dataset Generation from LLMs
by: Kowshik, Suhas S, et al.
Published: (2024)
by: Kowshik, Suhas S, et al.
Published: (2024)
RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Using Persuasive Writing Strategies to Explain and Detect Health Misinformation
by: Kamali, Danial, et al.
Published: (2022)
by: Kamali, Danial, et al.
Published: (2022)
Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense
by: Liu, Jiacheng, et al.
Published: (2026)
by: Liu, Jiacheng, et al.
Published: (2026)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
by: Chen, Justin Chih-Yao, et al.
Published: (2023)
by: Chen, Justin Chih-Yao, et al.
Published: (2023)
Similar Items
-
A Tale of LLMs and Induced Small Proxies: Scalable Agents for Knowledge Mining
by: Zhang, Sipeng, et al.
Published: (2025) -
UltraGen: Extremely Fine-grained Controllable Generation via Attribute Reconstruction and Global Preference Optimization
by: Yun, Longfei, et al.
Published: (2025) -
Configurable Foundation Models: Building LLMs from a Modular Perspective
by: Xiao, Chaojun, et al.
Published: (2024) -
How to Synthesize Text Data without Model Collapse?
by: Zhu, Xuekai, et al.
Published: (2024) -
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026)