Entropy-guided sequence weighting for efficient exploration in RL-based LLM fine-tuning
Fuente:
arXiv
Saved in:
| Main Author: | Vanlioglu, Abdullah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaGradSelect: An adaptive gradient-guided layer selection method for efficient fine-tuning of SLMs
by: Kumar, Anshul, et al.
Published: (2025)
by: Kumar, Anshul, et al.
Published: (2025)
On Entropy Control in LLM-RL Algorithms
by: Shen, Han
Published: (2025)
by: Shen, Han
Published: (2025)
Open-weight genome language model safeguards: Assessing robustness via adversarial fine-tuning
by: Black, James R. M., et al.
Published: (2025)
by: Black, James R. M., et al.
Published: (2025)
Efficient and Private: Memorisation under differentially private parameter-efficient fine-tuning in language models
by: Ma, Olivia, et al.
Published: (2024)
by: Ma, Olivia, et al.
Published: (2024)
GeoLoRA: Geometric integration for parameter efficient fine-tuning
by: Schotthöfer, Steffen, et al.
Published: (2024)
by: Schotthöfer, Steffen, et al.
Published: (2024)
Epistemically-guided forward-backward exploration
by: Urpí, Núria Armengol, et al.
Published: (2025)
by: Urpí, Núria Armengol, et al.
Published: (2025)
Learning from models beyond fine-tuning
by: Zheng, Hongling, et al.
Published: (2023)
by: Zheng, Hongling, et al.
Published: (2023)
FLoRA: Fused forward-backward adapters for parameter efficient fine-tuning and reducing inference-time latencies of LLMs
by: Gowda, Dhananjaya, et al.
Published: (2025)
by: Gowda, Dhananjaya, et al.
Published: (2025)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
Empirical influence functions to understand the logic of fine-tuning
by: Matelsky, Jordan K., et al.
Published: (2024)
by: Matelsky, Jordan K., et al.
Published: (2024)
Curiosity & Entropy Driven Unsupervised RL in Multiple Environments
by: Dewan, Shaurya, et al.
Published: (2024)
by: Dewan, Shaurya, et al.
Published: (2024)
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
by: Schmied, Thomas, et al.
Published: (2025)
by: Schmied, Thomas, et al.
Published: (2025)
Train on Validation (ToV): Fast data selection with applications to fine-tuning
by: Jain, Ayush, et al.
Published: (2025)
by: Jain, Ayush, et al.
Published: (2025)
Evolutionary fine tuning of quantized convolution-based deep learning models
by: Pietroń, Marcin
Published: (2026)
by: Pietroń, Marcin
Published: (2026)
Novel RL approach for efficient Elevator Group Control Systems
by: Vaartjes, Nathan, et al.
Published: (2025)
by: Vaartjes, Nathan, et al.
Published: (2025)
Token-Efficient RL for LLM Reasoning
by: Lee, Alan, et al.
Published: (2025)
by: Lee, Alan, et al.
Published: (2025)
LIFT: Interpretable truck driving risk prediction with literature-informed fine-tuned LLMs
by: Hu, Xiao, et al.
Published: (2025)
by: Hu, Xiao, et al.
Published: (2025)
Jal Anveshak: Prediction of fishing zones using fine-tuned LlaMa 2
by: Mejari, Arnav, et al.
Published: (2024)
by: Mejari, Arnav, et al.
Published: (2024)
Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective
by: Hu, Xiao, et al.
Published: (2026)
by: Hu, Xiao, et al.
Published: (2026)
Combining LLM decision and RL action selection to improve RL policy for adaptive interventions
by: Karine, Karine, et al.
Published: (2025)
by: Karine, Karine, et al.
Published: (2025)
EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
by: Shi, Jiahe, et al.
Published: (2025)
by: Shi, Jiahe, et al.
Published: (2025)
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
by: Lv, Lei, et al.
Published: (2026)
by: Lv, Lei, et al.
Published: (2026)
Sample-efficient and Scalable Exploration in Continuous-Time RL
by: Iten, Klemens, et al.
Published: (2025)
by: Iten, Klemens, et al.
Published: (2025)
Scores as Actions: a framework of fine-tuning diffusion models by continuous-time reinforcement learning
by: Zhao, Hanyang, et al.
Published: (2024)
by: Zhao, Hanyang, et al.
Published: (2024)
Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuning
by: Tan, Qitao, et al.
Published: (2025)
by: Tan, Qitao, et al.
Published: (2025)
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
by: Sukhija, Bhavya, et al.
Published: (2024)
by: Sukhija, Bhavya, et al.
Published: (2024)
Rethinking harmless refusals when fine-tuning foundation models
by: Pop, Florin, et al.
Published: (2024)
by: Pop, Florin, et al.
Published: (2024)
LaDi-RL: Latent Diffusion Reasoning Prevents Entropy Collapse in Reinforcement Learning
by: Kang, Haoqiang, et al.
Published: (2026)
by: Kang, Haoqiang, et al.
Published: (2026)
FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
by: Brantley, Kianté, et al.
Published: (2025)
by: Brantley, Kianté, et al.
Published: (2025)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
by: Kim, Minseon, et al.
Published: (2025)
by: Kim, Minseon, et al.
Published: (2025)
Investigating the performance of Retrieval-Augmented Generation and fine-tuning for the development of AI-driven knowledge-based systems
by: Lakatos, Robert, et al.
Published: (2024)
by: Lakatos, Robert, et al.
Published: (2024)
Uncertainty quantification in fine-tuned LLMs using LoRA ensembles
by: Balabanov, Oleksandr, et al.
Published: (2024)
by: Balabanov, Oleksandr, et al.
Published: (2024)
Conditional generation of antibody sequences with classifier-guided germline-absorbing discrete diffusion
by: Sanders, Justin, et al.
Published: (2026)
by: Sanders, Justin, et al.
Published: (2026)
Do deep neural networks utilize the weight space efficiently?
by: Koyun, Onur Can, et al.
Published: (2024)
by: Koyun, Onur Can, et al.
Published: (2024)
Uncertainty modeling for fine-tuned implicit functions
by: Susmelj, Anna, et al.
Published: (2024)
by: Susmelj, Anna, et al.
Published: (2024)
Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs
by: Zhao, Kang, et al.
Published: (2024)
by: Zhao, Kang, et al.
Published: (2024)
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
by: Wu, Xian, et al.
Published: (2026)
by: Wu, Xian, et al.
Published: (2026)
Improving embedding with contrastive fine-tuning on small datasets with expert-augmented scores
by: Lu, Jun, et al.
Published: (2024)
by: Lu, Jun, et al.
Published: (2024)
Similar Items
-
AdaGradSelect: An adaptive gradient-guided layer selection method for efficient fine-tuning of SLMs
by: Kumar, Anshul, et al.
Published: (2025) -
On Entropy Control in LLM-RL Algorithms
by: Shen, Han
Published: (2025) -
Open-weight genome language model safeguards: Assessing robustness via adversarial fine-tuning
by: Black, James R. M., et al.
Published: (2025) -
Efficient and Private: Memorisation under differentially private parameter-efficient fine-tuning in language models
by: Ma, Olivia, et al.
Published: (2024) -
GeoLoRA: Geometric integration for parameter efficient fine-tuning
by: Schotthöfer, Steffen, et al.
Published: (2024)