Break the Block: Dynamic-size Reasoning Blocks for Diffusion Large Language Models via Monotonic Entropy Descent with Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Yan, Qiu, Ruihong, Huang, Zi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models
von: Jiang, Yan, et al.
Veröffentlicht: (2026)
von: Jiang, Yan, et al.
Veröffentlicht: (2026)
When to Commit? Towards Variable-Size Self-Contained Blocks for Discrete Diffusion Language Models
von: Wang, Danny, et al.
Veröffentlicht: (2026)
von: Wang, Danny, et al.
Veröffentlicht: (2026)
TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only
von: Liu, Yilun, et al.
Veröffentlicht: (2026)
von: Liu, Yilun, et al.
Veröffentlicht: (2026)
Does Homophily Help in Robust Test-time Node Classification?
von: Jiang, Yan, et al.
Veröffentlicht: (2025)
von: Jiang, Yan, et al.
Veröffentlicht: (2025)
What Information Matters? Graph Out-of-Distribution Detection via Tri-Component Information Decomposition
von: Wang, Danny, et al.
Veröffentlicht: (2026)
von: Wang, Danny, et al.
Veröffentlicht: (2026)
GCondenser: Benchmarking Graph Condensation
von: Liu, Yilun, et al.
Veröffentlicht: (2024)
von: Liu, Yilun, et al.
Veröffentlicht: (2024)
Save It All: Enabling Full Parameter Tuning for Federated Large Language Models via Cycle Block Gradient Descent
von: Wang, Lin, et al.
Veröffentlicht: (2024)
von: Wang, Lin, et al.
Veröffentlicht: (2024)
Exploiting Block Coordinate Descent for Cost-Effective LLM Model Training
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
GOLD: Graph Out-of-Distribution Detection via Implicit Adversarial Latent Generation
von: Wang, Danny, et al.
Veröffentlicht: (2025)
von: Wang, Danny, et al.
Veröffentlicht: (2025)
LightningRL: Breaking the Accuracy-Parallelism Trade-off of Block-wise dLLMs via Reinforcement Learning
von: Hu, Yanzhe, et al.
Veröffentlicht: (2026)
von: Hu, Yanzhe, et al.
Veröffentlicht: (2026)
Asynchronous Distributed Reinforcement Learning for LQR Control via Zeroth-Order Block Coordinate Descent
von: Jing, Gangshan, et al.
Veröffentlicht: (2021)
von: Jing, Gangshan, et al.
Veröffentlicht: (2021)
GeoBlock: Inferring Block Granularity from Dependency Geometry in Diffusion Language Models
von: Wan, Lipeng, et al.
Veröffentlicht: (2026)
von: Wan, Lipeng, et al.
Veröffentlicht: (2026)
DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
von: Shing, Makoto, et al.
Veröffentlicht: (2025)
von: Shing, Makoto, et al.
Veröffentlicht: (2025)
PUMA: Efficient Continual Graph Learning for Node Classification with Graph Condensation
von: Liu, Yilun, et al.
Veröffentlicht: (2023)
von: Liu, Yilun, et al.
Veröffentlicht: (2023)
ParaBlock: Communication-Computation Parallel Block Coordinate Federated Learning for Large Language Models
von: Wang, Yujia, et al.
Veröffentlicht: (2025)
von: Wang, Yujia, et al.
Veröffentlicht: (2025)
Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
von: Arriola, Marianne, et al.
Veröffentlicht: (2025)
von: Arriola, Marianne, et al.
Veröffentlicht: (2025)
Revisiting Entropy in Reinforcement Learning for Large Reasoning Models
von: Jin, Renren, et al.
Veröffentlicht: (2025)
von: Jin, Renren, et al.
Veröffentlicht: (2025)
Block Circulant Adapter for Large Language Models
von: Ding, Xinyu, et al.
Veröffentlicht: (2025)
von: Ding, Xinyu, et al.
Veröffentlicht: (2025)
SBGD: Improving Graph Diffusion Generative Model via Stochastic Block Diffusion
von: Su, Junwei, et al.
Veröffentlicht: (2025)
von: Su, Junwei, et al.
Veröffentlicht: (2025)
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
von: Zhao, Siyan, et al.
Veröffentlicht: (2025)
von: Zhao, Siyan, et al.
Veröffentlicht: (2025)
Differentially Private Random Block Coordinate Descent
von: Maranjyan, Artavazd, et al.
Veröffentlicht: (2024)
von: Maranjyan, Artavazd, et al.
Veröffentlicht: (2024)
Text Meets Topology: Rethinking Out-of-distribution Detection in Text-Rich Networks
von: Wang, Danny, et al.
Veröffentlicht: (2025)
von: Wang, Danny, et al.
Veröffentlicht: (2025)
From Tokens to Blocks: A Block-Diffusion Perspective on Molecular Generation
von: Yang, Qianwei, et al.
Veröffentlicht: (2026)
von: Yang, Qianwei, et al.
Veröffentlicht: (2026)
CBQ: Cross-Block Quantization for Large Language Models
von: Ding, Xin, et al.
Veröffentlicht: (2023)
von: Ding, Xin, et al.
Veröffentlicht: (2023)
Block Coordinate Descent for Neural Networks Provably Finds Global Minima
von: Akiyama, Shunta
Veröffentlicht: (2025)
von: Akiyama, Shunta
Veröffentlicht: (2025)
TrajDLM: Topology-Aware Block Diffusion Language Model for Trajectory Generation
von: Wongso, Wilson, et al.
Veröffentlicht: (2026)
von: Wongso, Wilson, et al.
Veröffentlicht: (2026)
On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models
von: Wang, Shumin, et al.
Veröffentlicht: (2026)
von: Wang, Shumin, et al.
Veröffentlicht: (2026)
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
ALSA: Anchors in Logit Space for Out-of-Distribution Accuracy Estimation
von: Liu, Chenzhi, et al.
Veröffentlicht: (2025)
von: Liu, Chenzhi, et al.
Veröffentlicht: (2025)
Efficient Large Language Model Inference with Neural Block Linearization
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
FedBCD:Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning
von: Liu, Junkang, et al.
Veröffentlicht: (2026)
von: Liu, Junkang, et al.
Veröffentlicht: (2026)
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
von: Lu, Guanxi, et al.
Veröffentlicht: (2025)
von: Lu, Guanxi, et al.
Veröffentlicht: (2025)
Teaching Large Language Models to Reason with Reinforcement Learning
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
LEGO: Language Model Building Blocks
von: Bhansali, Shrenik, et al.
Veröffentlicht: (2024)
von: Bhansali, Shrenik, et al.
Veröffentlicht: (2024)
Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
von: Park, Jaesung R., et al.
Veröffentlicht: (2025)
von: Park, Jaesung R., et al.
Veröffentlicht: (2025)
Blocked Gibbs meets Diffusion Transformers: Unsupervised Learning for Constraint Optimization
von: Xu, Yudong W., et al.
Veröffentlicht: (2026)
von: Xu, Yudong W., et al.
Veröffentlicht: (2026)
Reasoning in Diffusion Large Language Models is Concentrated in Dynamic Confusion Zones
von: Chen, Ranfei, et al.
Veröffentlicht: (2025)
von: Chen, Ranfei, et al.
Veröffentlicht: (2025)
Hidden State Differential Private Mini-Batch Block Coordinate Descent for Multi-convexity Optimization
von: Chen, Ding, et al.
Veröffentlicht: (2024)
von: Chen, Ding, et al.
Veröffentlicht: (2024)
On Predictability of Reinforcement Learning Dynamics for Large Language Models
von: Cai, Yuchen, et al.
Veröffentlicht: (2025)
von: Cai, Yuchen, et al.
Veröffentlicht: (2025)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models
von: Jiang, Yan, et al.
Veröffentlicht: (2026) -
When to Commit? Towards Variable-Size Self-Contained Blocks for Discrete Diffusion Language Models
von: Wang, Danny, et al.
Veröffentlicht: (2026) -
TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only
von: Liu, Yilun, et al.
Veröffentlicht: (2026) -
Does Homophily Help in Robust Test-time Node Classification?
von: Jiang, Yan, et al.
Veröffentlicht: (2025) -
What Information Matters? Graph Out-of-Distribution Detection via Tri-Component Information Decomposition
von: Wang, Danny, et al.
Veröffentlicht: (2026)