Discovering Reinforcement Learning Interfaces with Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jaswal, Akshat Singh, Baghel, Ashish, Chopra, Paras |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AWE: Adaptive Agents for Dynamic Web Penetration Testing
von: Jaswal, Akshat Singh, et al.
Veröffentlicht: (2026)
von: Jaswal, Akshat Singh, et al.
Veröffentlicht: (2026)
See, Symbolize, Act: Grounding VLMs with Spatial Representations for Better Gameplay
von: Baghel, Ashish, et al.
Veröffentlicht: (2026)
von: Baghel, Ashish, et al.
Veröffentlicht: (2026)
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
von: Sharma, Aman, et al.
Veröffentlicht: (2026)
von: Sharma, Aman, et al.
Veröffentlicht: (2026)
Why LLMs Aren't Scientists Yet: Lessons from Four Autonomous Research Attempts
von: Trehan, Dhruv, et al.
Veröffentlicht: (2026)
von: Trehan, Dhruv, et al.
Veröffentlicht: (2026)
The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
von: Sharma, Aman, et al.
Veröffentlicht: (2025)
von: Sharma, Aman, et al.
Veröffentlicht: (2025)
Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
von: Sharma, Aman, et al.
Veröffentlicht: (2025)
von: Sharma, Aman, et al.
Veröffentlicht: (2025)
Hybrid Neural World Models
von: Lakshmanan, Pranav, et al.
Veröffentlicht: (2026)
von: Lakshmanan, Pranav, et al.
Veröffentlicht: (2026)
A Timeline and Analysis for Representation Plasticity in Large Language Models
von: Kannan, Akshat
Veröffentlicht: (2024)
von: Kannan, Akshat
Veröffentlicht: (2024)
Building Interpretable Models for Moral Decision-Making
von: Goel, Mayank, et al.
Veröffentlicht: (2026)
von: Goel, Mayank, et al.
Veröffentlicht: (2026)
It Takes Two: A Dual Stage Approach for Terminology-Aware Translation
von: Jaswal, Akshat Singh
Veröffentlicht: (2025)
von: Jaswal, Akshat Singh
Veröffentlicht: (2025)
METIS: Mentoring Engine for Thoughtful Inquiry & Solutions
von: Kumar, Abhinav Rajeev, et al.
Veröffentlicht: (2026)
von: Kumar, Abhinav Rajeev, et al.
Veröffentlicht: (2026)
Safe Reinforcement Learning with Learned Non-Markovian Safety Constraints
von: Low, Siow Meng, et al.
Veröffentlicht: (2024)
von: Low, Siow Meng, et al.
Veröffentlicht: (2024)
Offline Safe Reinforcement Learning Using Trajectory Classification
von: Gong, Ze, et al.
Veröffentlicht: (2024)
von: Gong, Ze, et al.
Veröffentlicht: (2024)
Leveraging Constraint Violation Signals For Action-Constrained Reinforcement Learning
von: Brahmanage, Janaka Chathuranga, et al.
Veröffentlicht: (2025)
von: Brahmanage, Janaka Chathuranga, et al.
Veröffentlicht: (2025)
Discovering Temporally-Aware Reinforcement Learning Algorithms
von: Jackson, Matthew Thomas, et al.
Veröffentlicht: (2024)
von: Jackson, Matthew Thomas, et al.
Veröffentlicht: (2024)
Beyond Hard Constraints: Budget-Conditioned Reachability For Safe Offline Reinforcement Learning
von: Brahmanage, Janaka Chathuranga, et al.
Veröffentlicht: (2026)
von: Brahmanage, Janaka Chathuranga, et al.
Veröffentlicht: (2026)
On the Non-Identifiability of Steering Vectors in Large Language Models
von: Venkatesh, Sohan, et al.
Veröffentlicht: (2026)
von: Venkatesh, Sohan, et al.
Veröffentlicht: (2026)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
von: Nirmal, Ayushi, et al.
Veröffentlicht: (2024)
von: Nirmal, Ayushi, et al.
Veröffentlicht: (2024)
Reinforcement Learning with Promising Tokens for Large Language Models
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2026)
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2026)
On Predictability of Reinforcement Learning Dynamics for Large Language Models
von: Cai, Yuchen, et al.
Veröffentlicht: (2025)
von: Cai, Yuchen, et al.
Veröffentlicht: (2025)
Discovering and Reasoning of Causality in the Hidden World with Large Language Models
von: Liu, Chenxi, et al.
Veröffentlicht: (2024)
von: Liu, Chenxi, et al.
Veröffentlicht: (2024)
Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations
von: Singh, Simardeep, et al.
Veröffentlicht: (2026)
von: Singh, Simardeep, et al.
Veröffentlicht: (2026)
Scaling Is All You Need: Autonomous Driving with JAX-Accelerated Reinforcement Learning
von: Harmel, Moritz, et al.
Veröffentlicht: (2023)
von: Harmel, Moritz, et al.
Veröffentlicht: (2023)
Efficient Reinforcement Learning for Large Language Models with Intrinsic Exploration
von: Sun, Yan, et al.
Veröffentlicht: (2025)
von: Sun, Yan, et al.
Veröffentlicht: (2025)
Offline Regularised Reinforcement Learning for Large Language Models Alignment
von: Richemond, Pierre Harvey, et al.
Veröffentlicht: (2024)
von: Richemond, Pierre Harvey, et al.
Veröffentlicht: (2024)
RLAX: Large-Scale, Distributed Reinforcement Learning for Large Language Models on TPUs
von: Zhou, Runlong, et al.
Veröffentlicht: (2025)
von: Zhou, Runlong, et al.
Veröffentlicht: (2025)
FineZip : Pushing the Limits of Large Language Models for Practical Lossless Text Compression
von: Mittu, Fazal, et al.
Veröffentlicht: (2024)
von: Mittu, Fazal, et al.
Veröffentlicht: (2024)
Language Models Entangle Language and Culture
von: Jain, Shourya, et al.
Veröffentlicht: (2026)
von: Jain, Shourya, et al.
Veröffentlicht: (2026)
Dynamic Adversarial Reinforcement Learning for Robust Multimodal Large Language Models
von: Bao, Yicheng, et al.
Veröffentlicht: (2026)
von: Bao, Yicheng, et al.
Veröffentlicht: (2026)
Evolutionary Discovery of Reinforcement Learning Algorithms via Large Language Models
von: Sygkounas, Alkis, et al.
Veröffentlicht: (2026)
von: Sygkounas, Alkis, et al.
Veröffentlicht: (2026)
Discovering What You Can Control: Interventional Boundary Discovery for Reinforcement Learning
von: Liu, Jiaxin, et al.
Veröffentlicht: (2026)
von: Liu, Jiaxin, et al.
Veröffentlicht: (2026)
Internalizing Meta-Experience into Memory for Guided Reinforcement Learning in Large Language Models
von: Huang, Shiting, et al.
Veröffentlicht: (2026)
von: Huang, Shiting, et al.
Veröffentlicht: (2026)
Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2024)
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2024)
CAMEL: Continuous Action Masking Enabled by Large Language Models for Reinforcement Learning
von: Zhao, Yanxiao, et al.
Veröffentlicht: (2025)
von: Zhao, Yanxiao, et al.
Veröffentlicht: (2025)
Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models
von: Balashov, Andrii
Veröffentlicht: (2025)
von: Balashov, Andrii
Veröffentlicht: (2025)
Reward Generation via Large Vision-Language Model in Offline Reinforcement Learning
von: Lee, Younghwan, et al.
Veröffentlicht: (2025)
von: Lee, Younghwan, et al.
Veröffentlicht: (2025)
The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning
von: Xu, Yi, et al.
Veröffentlicht: (2026)
von: Xu, Yi, et al.
Veröffentlicht: (2026)
FinXplore: An Adaptive Deep Reinforcement Learning Framework for Balancing and Discovering Investment Opportunities
von: Choudhary, Himanshu, et al.
Veröffentlicht: (2025)
von: Choudhary, Himanshu, et al.
Veröffentlicht: (2025)
Causal Feature Selection for Responsible Machine Learning
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
Large Language Model-Enhanced Reinforcement Learning for Generic Bus Holding Control Strategies
von: Yu, Jiajie, et al.
Veröffentlicht: (2024)
von: Yu, Jiajie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AWE: Adaptive Agents for Dynamic Web Penetration Testing
von: Jaswal, Akshat Singh, et al.
Veröffentlicht: (2026) -
See, Symbolize, Act: Grounding VLMs with Spatial Representations for Better Gameplay
von: Baghel, Ashish, et al.
Veröffentlicht: (2026) -
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
von: Sharma, Aman, et al.
Veröffentlicht: (2026) -
Why LLMs Aren't Scientists Yet: Lessons from Four Autonomous Research Attempts
von: Trehan, Dhruv, et al.
Veröffentlicht: (2026) -
The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
von: Sharma, Aman, et al.
Veröffentlicht: (2025)