EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Shi, Jiahe, Gao, Zhengqi, Ko, Ching-Yun, Boning, Duane |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
di: Zha, Kaiwen, et al.
Pubblicazione: (2025)
di: Zha, Kaiwen, et al.
Pubblicazione: (2025)
KirchhoffNet: A Scalable Ultra Fast Analog Neural Network
di: Gao, Zhengqi, et al.
Pubblicazione: (2023)
di: Gao, Zhengqi, et al.
Pubblicazione: (2023)
REG: Rectified Gradient Guidance for Conditional Diffusion Models
di: Gao, Zhengqi, et al.
Pubblicazione: (2025)
di: Gao, Zhengqi, et al.
Pubblicazione: (2025)
RDIT: Residual-based Diffusion Implicit Models for Probabilistic Time Series Forecasting
di: Lai, Chih-Yu, et al.
Pubblicazione: (2025)
di: Lai, Chih-Yu, et al.
Pubblicazione: (2025)
OriGen:Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-Reflection
di: Cui, Fan, et al.
Pubblicazione: (2024)
di: Cui, Fan, et al.
Pubblicazione: (2024)
Calibration-Aware Policy Optimization for Reasoning LLMs
di: Wang, Ziqi, et al.
Pubblicazione: (2026)
di: Wang, Ziqi, et al.
Pubblicazione: (2026)
Towards Reliable, Uncertainty-Aware Alignment
di: Banerjee, Debangshu, et al.
Pubblicazione: (2025)
di: Banerjee, Debangshu, et al.
Pubblicazione: (2025)
SP2RINT: Spatially-Decoupled Physics-Inspired Progressive Inverse Optimization for Scalable, PDE-Constrained Meta-Optical Neural Network Training
di: Ma, Pingchuan, et al.
Pubblicazione: (2025)
di: Ma, Pingchuan, et al.
Pubblicazione: (2025)
GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs
di: Zhao, Yiyang, et al.
Pubblicazione: (2025)
di: Zhao, Yiyang, et al.
Pubblicazione: (2025)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
di: Zhang, Yiqi, et al.
Pubblicazione: (2026)
di: Zhang, Yiqi, et al.
Pubblicazione: (2026)
PIC2O-Sim: A Physics-Inspired Causality-Aware Dynamic Convolutional Neural Operator for Ultra-Fast Photonic Device FDTD Simulation
di: Ma, Pingchuan, et al.
Pubblicazione: (2024)
di: Ma, Pingchuan, et al.
Pubblicazione: (2024)
RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs
di: Hu, Ruike, et al.
Pubblicazione: (2025)
di: Hu, Ruike, et al.
Pubblicazione: (2025)
GAC: Stabilizing Asynchronous RL Training for LLMs via Gradient Alignment Control
di: Xu, Haofeng, et al.
Pubblicazione: (2026)
di: Xu, Haofeng, et al.
Pubblicazione: (2026)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
di: Zhan, Guojian, et al.
Pubblicazione: (2025)
di: Zhan, Guojian, et al.
Pubblicazione: (2025)
On Entropy Control in LLM-RL Algorithms
di: Shen, Han
Pubblicazione: (2025)
di: Shen, Han
Pubblicazione: (2025)
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
di: Wang, Yiting, et al.
Pubblicazione: (2025)
di: Wang, Yiting, et al.
Pubblicazione: (2025)
Wrong Code, Right Structure: Learning Netlist Representations from Imperfect LLM-Generated RTL
di: Cai, Siyang, et al.
Pubblicazione: (2026)
di: Cai, Siyang, et al.
Pubblicazione: (2026)
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
di: Zhoubian, Sining, et al.
Pubblicazione: (2025)
di: Zhoubian, Sining, et al.
Pubblicazione: (2025)
Make Every Move Count: LLM-based High-Quality RTL Code Generation Using MCTS
di: DeLorenzo, Matthew, et al.
Pubblicazione: (2024)
di: DeLorenzo, Matthew, et al.
Pubblicazione: (2024)
Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
di: Chen, Xinzhu, et al.
Pubblicazione: (2025)
di: Chen, Xinzhu, et al.
Pubblicazione: (2025)
Curiosity & Entropy Driven Unsupervised RL in Multiple Environments
di: Dewan, Shaurya, et al.
Pubblicazione: (2024)
di: Dewan, Shaurya, et al.
Pubblicazione: (2024)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
di: Muslimani, Calarina, et al.
Pubblicazione: (2025)
di: Muslimani, Calarina, et al.
Pubblicazione: (2025)
Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation
di: Li, Huanyu, et al.
Pubblicazione: (2026)
di: Li, Huanyu, et al.
Pubblicazione: (2026)
Partial Policy Gradients for RL in LLMs
di: Mathur, Puneet, et al.
Pubblicazione: (2026)
di: Mathur, Puneet, et al.
Pubblicazione: (2026)
Entropy Aware Reward Guidance for Diffusion Language Model Alignment
di: Tejaswi, Atula, et al.
Pubblicazione: (2026)
di: Tejaswi, Atula, et al.
Pubblicazione: (2026)
QiMeng-CodeV-SVA: Training Specialized LLMs for Hardware Assertion Generation via RTL-Grounded Bidirectional Data Synthesis
di: Wu, Yutong, et al.
Pubblicazione: (2026)
di: Wu, Yutong, et al.
Pubblicazione: (2026)
A Deep Dive into Scaling RL for Code Generation with Synthetic Data and Curricula
di: Sancaktar, Cansu, et al.
Pubblicazione: (2026)
di: Sancaktar, Cansu, et al.
Pubblicazione: (2026)
RL in Name Only? Analyzing the Structural Assumptions in RL post-training for LLMs
di: Samineni, Soumya Rani, et al.
Pubblicazione: (2025)
di: Samineni, Soumya Rani, et al.
Pubblicazione: (2025)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
di: Hung, Kuo-Han, et al.
Pubblicazione: (2024)
di: Hung, Kuo-Han, et al.
Pubblicazione: (2024)
Multi-Objective Instruction-Aware Representation Learning in Procedural Content Generation RL
di: Kim, Sung-Hyun, et al.
Pubblicazione: (2025)
di: Kim, Sung-Hyun, et al.
Pubblicazione: (2025)
A Systematic Investigation of The RL-Jailbreaker in LLMs
di: Mohammedalamen, Montaser, et al.
Pubblicazione: (2026)
di: Mohammedalamen, Montaser, et al.
Pubblicazione: (2026)
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
di: Cheng, Ching-An, et al.
Pubblicazione: (2024)
di: Cheng, Ching-An, et al.
Pubblicazione: (2024)
Dual Alignment Maximin Optimization for Offline Model-based RL
di: Zhou, Chi, et al.
Pubblicazione: (2025)
di: Zhou, Chi, et al.
Pubblicazione: (2025)
RL-Finetuned LLMs for Privacy-Preserving Synthetic Rewriting
di: Shi, Zhan, et al.
Pubblicazione: (2025)
di: Shi, Zhan, et al.
Pubblicazione: (2025)
A Synthesizable RTL Implementation of Predictive Coding Networks
di: Oh, Timothy
Pubblicazione: (2026)
di: Oh, Timothy
Pubblicazione: (2026)
Are LLMs Better GNN Helpers? Rethinking Robust Graph Learning under Deficiencies with Iterative Refinement
di: Wang, Zhaoyan, et al.
Pubblicazione: (2025)
di: Wang, Zhaoyan, et al.
Pubblicazione: (2025)
Towards Reliable Alignment: Uncertainty-aware RLHF
di: Banerjee, Debangshu, et al.
Pubblicazione: (2024)
di: Banerjee, Debangshu, et al.
Pubblicazione: (2024)
Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment
di: Chen, Tiejin, et al.
Pubblicazione: (2026)
di: Chen, Tiejin, et al.
Pubblicazione: (2026)
RL-GPT: Integrating Reinforcement Learning and Code-as-policy
di: Liu, Shaoteng, et al.
Pubblicazione: (2024)
di: Liu, Shaoteng, et al.
Pubblicazione: (2024)
VeriDispatcher: Multi-Model Dispatching through Pre-Inference Difficulty Prediction for RTL Generation Optimization
di: Wang, Zeng, et al.
Pubblicazione: (2025)
di: Wang, Zeng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
di: Zha, Kaiwen, et al.
Pubblicazione: (2025) -
KirchhoffNet: A Scalable Ultra Fast Analog Neural Network
di: Gao, Zhengqi, et al.
Pubblicazione: (2023) -
REG: Rectified Gradient Guidance for Conditional Diffusion Models
di: Gao, Zhengqi, et al.
Pubblicazione: (2025) -
RDIT: Residual-based Diffusion Implicit Models for Probabilistic Time Series Forecasting
di: Lai, Chih-Yu, et al.
Pubblicazione: (2025) -
OriGen:Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-Reflection
di: Cui, Fan, et al.
Pubblicazione: (2024)