DYSTIL: Dynamic Strategy Induction with Large Language Models for Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Borui, McKeown, Kathleen, Ying, Rex |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
por: Zhang, Yunfan, et al.
Publicado: (2025)
por: Zhang, Yunfan, et al.
Publicado: (2025)
LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News
por: Zhang, Yunfan, et al.
Publicado: (2026)
por: Zhang, Yunfan, et al.
Publicado: (2026)
Guiding LLM Decision-Making with Fairness Reward Models
por: Hall, Zara, et al.
Publicado: (2025)
por: Hall, Zara, et al.
Publicado: (2025)
Getting Serious about Humor: Crafting Humor Datasets with Unfunny Large Language Models
por: Horvitz, Zachary, et al.
Publicado: (2024)
por: Horvitz, Zachary, et al.
Publicado: (2024)
On the Relation between Sensitivity and Accuracy in In-context Learning
por: Chen, Yanda, et al.
Publicado: (2022)
por: Chen, Yanda, et al.
Publicado: (2022)
Parallel Structures in Pre-training Data Yield In-Context Learning
por: Chen, Yanda, et al.
Publicado: (2024)
por: Chen, Yanda, et al.
Publicado: (2024)
Estimating Tail Risks in Language Model Output Distributions
por: Angell, Rico, et al.
Publicado: (2026)
por: Angell, Rico, et al.
Publicado: (2026)
Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
por: Deas, Nicholas, et al.
Publicado: (2025)
por: Deas, Nicholas, et al.
Publicado: (2025)
No Compute Left Behind: Rethinking Reasoning and Sampling with Masked Diffusion Models
por: Horvitz, Zachary, et al.
Publicado: (2025)
por: Horvitz, Zachary, et al.
Publicado: (2025)
Social Orientation: A New Feature for Dialogue Analysis
por: Morrill, Todd, et al.
Publicado: (2024)
por: Morrill, Todd, et al.
Publicado: (2024)
MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models
por: Ha, Hyeonjeong, et al.
Publicado: (2026)
por: Ha, Hyeonjeong, et al.
Publicado: (2026)
Program-Based Strategy Induction for Reinforcement Learning
por: Correa, Carlos G., et al.
Publicado: (2024)
por: Correa, Carlos G., et al.
Publicado: (2024)
A General Framework for Inference-time Scaling and Steering of Diffusion Models
por: Singhal, Raghav, et al.
Publicado: (2025)
por: Singhal, Raghav, et al.
Publicado: (2025)
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
por: Qian, Cheng, et al.
Publicado: (2025)
por: Qian, Cheng, et al.
Publicado: (2025)
StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples
por: Patel, Ajay, et al.
Publicado: (2024)
por: Patel, Ajay, et al.
Publicado: (2024)
Summarization of Opinionated Political Documents with Varied Perspectives
por: Deas, Nicholas, et al.
Publicado: (2024)
por: Deas, Nicholas, et al.
Publicado: (2024)
On Predictability of Reinforcement Learning Dynamics for Large Language Models
por: Cai, Yuchen, et al.
Publicado: (2025)
por: Cai, Yuchen, et al.
Publicado: (2025)
GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Models
por: Chen, Jialin, et al.
Publicado: (2025)
por: Chen, Jialin, et al.
Publicado: (2025)
Purpose in the Machine: Do Traffic Simulators Produce Distributionally Equivalent Outcomes for Reinforcement Learning Applications?
por: Chen, Rex, et al.
Publicado: (2023)
por: Chen, Rex, et al.
Publicado: (2023)
Large Language Model-Enhanced Reinforcement Learning for Generic Bus Holding Control Strategies
por: Yu, Jiajie, et al.
Publicado: (2024)
por: Yu, Jiajie, et al.
Publicado: (2024)
Dynamic Adversarial Reinforcement Learning for Robust Multimodal Large Language Models
por: Bao, Yicheng, et al.
Publicado: (2026)
por: Bao, Yicheng, et al.
Publicado: (2026)
Reading Subtext: Evaluating Large Language Models on Short Story Summarization with Writers
por: Subbiah, Melanie, et al.
Publicado: (2024)
por: Subbiah, Melanie, et al.
Publicado: (2024)
Tele-LLMs: A Series of Specialized Large Language Models for Telecommunications
por: Maatouk, Ali, et al.
Publicado: (2024)
por: Maatouk, Ali, et al.
Publicado: (2024)
Enhancing Multimodal Affective Analysis with Learned Live Comment Features
por: Deng, Zhaoyuan, et al.
Publicado: (2024)
por: Deng, Zhaoyuan, et al.
Publicado: (2024)
Efficient Reinforcement Learning with Large Language Model Priors
por: Yan, Xue, et al.
Publicado: (2024)
por: Yan, Xue, et al.
Publicado: (2024)
AREAL-DTA: Dynamic Tree Attention for Efficient Reinforcement Learning of Large Language Models
por: Zhang, Jiarui, et al.
Publicado: (2026)
por: Zhang, Jiarui, et al.
Publicado: (2026)
Learning Robust Social Strategies with Large Language Models
por: Piche, Dereck, et al.
Publicado: (2025)
por: Piche, Dereck, et al.
Publicado: (2025)
On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models
por: Wang, Shumin, et al.
Publicado: (2026)
por: Wang, Shumin, et al.
Publicado: (2026)
Teaching Large Language Models to Reason with Reinforcement Learning
por: Havrilla, Alex, et al.
Publicado: (2024)
por: Havrilla, Alex, et al.
Publicado: (2024)
Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models
por: Peng, Yifei, et al.
Publicado: (2025)
por: Peng, Yifei, et al.
Publicado: (2025)
Towards Understanding Sensitive and Decisive Patterns in Explainable AI: A Case Study of Model Interpretation in Geometric Deep Learning
por: Zhu, Jiajun, et al.
Publicado: (2024)
por: Zhu, Jiajun, et al.
Publicado: (2024)
Large Language Models as Computable Approximations to Solomonoff Induction
por: Wan, Jun, et al.
Publicado: (2025)
por: Wan, Jun, et al.
Publicado: (2025)
HiLWS: A Human-in-the-Loop Weak Supervision Framework for Curating Clinical and Home Video Data for Neurological Assessment
por: Irani, Atefeh, et al.
Publicado: (2025)
por: Irani, Atefeh, et al.
Publicado: (2025)
The Busemann Process and Steep Highways in Directed First Passage Percolation
por: McKeown, Sam
Publicado: (2025)
por: McKeown, Sam
Publicado: (2025)
Analysis of shape and angular orientation parameters of velocity-dependent dark matter annihilation signals in Galactic Centers for FIRE simulations
por: McKeown, Daniel
Publicado: (2025)
por: McKeown, Daniel
Publicado: (2025)
God's Babies
por: McKeown, John
Publicado: (2018)
por: McKeown, John
Publicado: (2018)
Maryland State Board for Higher Education Operating Budget Guideline Development. A Report to the Joint Chairmen of the Senate Budget and Taxation Committee and House Appropriations Committee, 1982 Session.
por: McKeown, Mary
Publicado: (1982)
por: McKeown, Mary
Publicado: (1982)
Article peer review: Is collegiate cooperation under threat, why and what to do about it?
por: Mick McKeown
Publicado: (2024)
por: Mick McKeown
Publicado: (2024)
The Meldrewfication of Mick
por: Mick McKeown
Publicado: (2024)
por: Mick McKeown
Publicado: (2024)
ImagineBench: Evaluating Reinforcement Learning with Large Language Model Rollouts
por: Pang, Jing-Cheng, et al.
Publicado: (2025)
por: Pang, Jing-Cheng, et al.
Publicado: (2025)
Ejemplares similares
-
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
por: Zhang, Yunfan, et al.
Publicado: (2025) -
LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News
por: Zhang, Yunfan, et al.
Publicado: (2026) -
Guiding LLM Decision-Making with Fairness Reward Models
por: Hall, Zara, et al.
Publicado: (2025) -
Getting Serious about Humor: Crafting Humor Datasets with Unfunny Large Language Models
por: Horvitz, Zachary, et al.
Publicado: (2024) -
On the Relation between Sensitivity and Accuracy in In-context Learning
por: Chen, Yanda, et al.
Publicado: (2022)