Beyond Mode Collapse: Distribution Matching for Diverse Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiaozhe, Li, Yang, Fang, Xinyu, Ding, Shengyuan, Li, Peiji, Chen, Yongkang, Ma, Yichuan, Lyu, Tianyi, Li, Linyang, Lin, Dahua, Guo, Qipeng, Liu, Qingwen, Chen, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
by: Ma, Yichuan, et al.
Published: (2026)
by: Ma, Yichuan, et al.
Published: (2026)
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go
by: Ma, Yichuan, et al.
Published: (2026)
by: Ma, Yichuan, et al.
Published: (2026)
NP-Engine: Empowering Optimization Reasoning in Large Language Models with Verifiable Synthetic NP Problems
by: Li, Xiaozhe, et al.
Published: (2025)
by: Li, Xiaozhe, et al.
Published: (2025)
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
by: Li, Peiji, et al.
Published: (2025)
by: Li, Peiji, et al.
Published: (2025)
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
by: Ma, Yichuan, et al.
Published: (2025)
by: Ma, Yichuan, et al.
Published: (2025)
TL-GRPO: Turn-Level RL for Reasoning-Guided Iterative Optimization
by: Li, Peiji, et al.
Published: (2026)
by: Li, Peiji, et al.
Published: (2026)
FastMCTS: A Simple Sampling Strategy for Data Synthesis
by: Li, Peiji, et al.
Published: (2025)
by: Li, Peiji, et al.
Published: (2025)
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
by: Li, Xiaozhe, et al.
Published: (2025)
by: Li, Xiaozhe, et al.
Published: (2025)
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
Case2Code: Scalable Synthetic Data for Code Generation
by: Shao, Yunfan, et al.
Published: (2024)
by: Shao, Yunfan, et al.
Published: (2024)
COINBench: Moving Beyond Individual Perspectives to Collective Intent Understanding
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
Escaping the Context Bottleneck: Active Context Curation for LLM Agents via Reinforcement Learning
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
Beyond Euclidean Proximity: Repairing Latent World Models with Horizon-Matched Trajectory Reachability Metrics
by: Li, Liangyu, et al.
Published: (2026)
by: Li, Liangyu, et al.
Published: (2026)
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
by: Fei, Zhaoye, et al.
Published: (2024)
by: Fei, Zhaoye, et al.
Published: (2024)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
by: Zeng, Zhiyuan, et al.
Published: (2024)
by: Zeng, Zhiyuan, et al.
Published: (2024)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
by: Liu, Henglin, et al.
Published: (2025)
by: Liu, Henglin, et al.
Published: (2025)
Extracellular vesicles—Potential link between periodontal disease and diabetic complications
by: Shengyuan Huang, et al.
Published: (2024)
by: Shengyuan Huang, et al.
Published: (2024)
Anchored Policy Optimization: Mitigating Exploration Collapse Via Support-Constrained Rectification
by: Wang, Tianyi, et al.
Published: (2026)
by: Wang, Tianyi, et al.
Published: (2026)
Visual-ERM: Reward Modeling for Visual Equivalence
by: Liu, Ziyu, et al.
Published: (2026)
by: Liu, Ziyu, et al.
Published: (2026)
Why Reasoning Models Collapse Themselves in Reasoning
by: Zixi, Li
Published: (2025)
by: Zixi, Li
Published: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation
by: Liu, Jinmei, et al.
Published: (2026)
by: Liu, Jinmei, et al.
Published: (2026)
Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learning
by: Chen, Chubin, et al.
Published: (2025)
by: Chen, Chubin, et al.
Published: (2025)
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?
by: Lei, Zhikai, et al.
Published: (2024)
by: Lei, Zhikai, et al.
Published: (2024)
DART-Research/NP-3DP: NP-Slicing— Curved Slicing & Toolpath Generation for 3D Printing
by: Yichuan Li, et al.
Published: (2025)
by: Yichuan Li, et al.
Published: (2025)
Tool-Augmented Agent for Closed-loop Optimization,Simulation,and Modeling Orchestration
by: Deng, Liyuan, et al.
Published: (2026)
by: Deng, Liyuan, et al.
Published: (2026)
COSMO-Agent: Tool-Augmented Agent for Closed-loop Optimization,Simulation,and Modeling Orchestration
by: Deng, Liyuan, et al.
Published: (2026)
by: Deng, Liyuan, et al.
Published: (2026)
Balanced Data Sampling for Language Model Training with Clustering
by: Shao, Yunfan, et al.
Published: (2024)
by: Shao, Yunfan, et al.
Published: (2024)
1.x-Distill: Breaking the Diversity, Quality, and Efficiency Barrier in Distribution Matching Distillation
by: Li, Haoyu, et al.
Published: (2026)
by: Li, Haoyu, et al.
Published: (2026)
Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM
by: Fang, Xinyu, et al.
Published: (2025)
by: Fang, Xinyu, et al.
Published: (2025)
Super-Resolution on Rotationally Scanned Photoacoustic Microscopy Images Incorporating Scanning Prior
by: Pan, Kai, et al.
Published: (2023)
by: Pan, Kai, et al.
Published: (2023)
Exploration by Random Distribution Distillation
by: Fang, Zhirui, et al.
Published: (2025)
by: Fang, Zhirui, et al.
Published: (2025)
Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis
by: Wu, Tianhe, et al.
Published: (2026)
by: Wu, Tianhe, et al.
Published: (2026)
A Survey of Inductive Reasoning for Large Language Models
by: Chen, Kedi, et al.
Published: (2025)
by: Chen, Kedi, et al.
Published: (2025)
MicroDiffuse3D Checkpoints - Pretrain
by: Li, Yongkang
Published: (2026)
by: Li, Yongkang
Published: (2026)
MicroDiffuse3D Checkpoints - Pretrain
by: Li, Yongkang
Published: (2026)
by: Li, Yongkang
Published: (2026)
Similar Items
-
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
by: Ma, Yichuan, et al.
Published: (2026) -
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
by: Li, Xiaozhe, et al.
Published: (2026) -
Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go
by: Ma, Yichuan, et al.
Published: (2026) -
NP-Engine: Empowering Optimization Reasoning in Large Language Models with Verifiable Synthetic NP Problems
by: Li, Xiaozhe, et al.
Published: (2025) -
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
by: Li, Peiji, et al.
Published: (2025)