Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Putta, Pranav, Mills, Edmund, Garg, Naman, Motwani, Sumeet, Finn, Chelsea, Garg, Divyansh, Rafailov, Rafael |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MALT: Improving Reasoning with Multi-Agent LLM Training
por: Motwani, Sumeet Ramesh, et al.
Publicado: (2024)
por: Motwani, Sumeet Ramesh, et al.
Publicado: (2024)
REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites
por: Garg, Divyansh, et al.
Publicado: (2025)
por: Garg, Divyansh, et al.
Publicado: (2025)
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
por: Xiang, Violet, et al.
Publicado: (2025)
por: Xiang, Violet, et al.
Publicado: (2025)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
por: Qu, Yuxiao, et al.
Publicado: (2024)
por: Qu, Yuxiao, et al.
Publicado: (2024)
Contrastive Preference Learning: Learning from Human Feedback without RL
por: Hejna, Joey, et al.
Publicado: (2023)
por: Hejna, Joey, et al.
Publicado: (2023)
MOTO: Offline Pre-training to Online Fine-tuning for Model-based Robot Learning
por: Rafailov, Rafael, et al.
Publicado: (2024)
por: Rafailov, Rafael, et al.
Publicado: (2024)
From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
por: Rafailov, Rafael, et al.
Publicado: (2024)
por: Rafailov, Rafael, et al.
Publicado: (2024)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
por: Rafailov, Rafael, et al.
Publicado: (2023)
por: Rafailov, Rafael, et al.
Publicado: (2023)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
por: Rafailov, Rafael, et al.
Publicado: (2024)
por: Rafailov, Rafael, et al.
Publicado: (2024)
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
por: Ferrag, Mohamed Amine, et al.
Publicado: (2025)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2025)
h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
por: Motwani, Sumeet Ramesh, et al.
Publicado: (2025)
por: Motwani, Sumeet Ramesh, et al.
Publicado: (2025)
JAF: Judge Agent Forest
por: Garg, Sahil, et al.
Publicado: (2026)
por: Garg, Sahil, et al.
Publicado: (2026)
ROER: Regularized Optimal Experience Replay
por: Li, Changling, et al.
Publicado: (2024)
por: Li, Changling, et al.
Publicado: (2024)
TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse
por: Dong, Perry, et al.
Publicado: (2026)
por: Dong, Perry, et al.
Publicado: (2026)
Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning
por: Wu, Qingyuan, et al.
Publicado: (2025)
por: Wu, Qingyuan, et al.
Publicado: (2025)
Graph Transformers without Positional Encodings
por: Garg, Ayush
Publicado: (2024)
por: Garg, Ayush
Publicado: (2024)
Adaptive Memory Crystallization for Autonomous AI Agent Learning in Dynamic Environments
por: Khanda, Rajat, et al.
Publicado: (2026)
por: Khanda, Rajat, et al.
Publicado: (2026)
STARC: A General Framework For Quantifying Differences Between Reward Functions
por: Skalse, Joar, et al.
Publicado: (2023)
por: Skalse, Joar, et al.
Publicado: (2023)
EXPO: Stable Reinforcement Learning with Expressive Policies
por: Dong, Perry, et al.
Publicado: (2025)
por: Dong, Perry, et al.
Publicado: (2025)
Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
por: Hsu, Sheryl, et al.
Publicado: (2024)
por: Hsu, Sheryl, et al.
Publicado: (2024)
Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised Learning
por: Xie, Johnathan, et al.
Publicado: (2024)
por: Xie, Johnathan, et al.
Publicado: (2024)
Stable Forgetting: Bounded Parameter-Efficient Unlearning in Foundation Models
por: Garg, Arpit, et al.
Publicado: (2025)
por: Garg, Arpit, et al.
Publicado: (2025)
Universal Neural Functionals
por: Zhou, Allan, et al.
Publicado: (2024)
por: Zhou, Allan, et al.
Publicado: (2024)
TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools
por: Gao, Shanghua, et al.
Publicado: (2025)
por: Gao, Shanghua, et al.
Publicado: (2025)
Reinforcement Learning via Implicit Imitation Guidance
por: Dong, Perry, et al.
Publicado: (2025)
por: Dong, Perry, et al.
Publicado: (2025)
The Road of Adaptive AI for Precision in Cybersecurity
por: Garg, Sahil
Publicado: (2025)
por: Garg, Sahil
Publicado: (2025)
Learning Long-Context Diffusion Policies via Past-Token Prediction
por: Torne, Marcel, et al.
Publicado: (2025)
por: Torne, Marcel, et al.
Publicado: (2025)
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
por: Luo, Xufang, et al.
Publicado: (2025)
por: Luo, Xufang, et al.
Publicado: (2025)
Margin-calibrated Classifier Guidance for Property-driven Synthesis Planning
por: Laabid, Najwa, et al.
Publicado: (2026)
por: Laabid, Najwa, et al.
Publicado: (2026)
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
por: Motwani, Sumeet Ramesh, et al.
Publicado: (2026)
por: Motwani, Sumeet Ramesh, et al.
Publicado: (2026)
Self-Initiated Open World Learning for Autonomous AI Agents
por: Liu, Bing, et al.
Publicado: (2021)
por: Liu, Bing, et al.
Publicado: (2021)
Multi-Agent Risks from Advanced AI
por: Hammond, Lewis, et al.
Publicado: (2025)
por: Hammond, Lewis, et al.
Publicado: (2025)
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
por: Gautam, Dhruv, et al.
Publicado: (2025)
por: Gautam, Dhruv, et al.
Publicado: (2025)
Efficient Imitation Learning with Conservative World Models
por: Kolev, Victor, et al.
Publicado: (2024)
por: Kolev, Victor, et al.
Publicado: (2024)
Generative AI Agents in Autonomous Machines: A Safety Perspective
por: Jabbour, Jason, et al.
Publicado: (2024)
por: Jabbour, Jason, et al.
Publicado: (2024)
Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning
por: Ahmed, Nesreen K., et al.
Publicado: (2026)
por: Ahmed, Nesreen K., et al.
Publicado: (2026)
FASTER: Value-Guided Sampling for Fast RL
por: Dong, Perry, et al.
Publicado: (2026)
por: Dong, Perry, et al.
Publicado: (2026)
Polychromic Objectives for Reinforcement Learning
por: Hamid, Jubayer Ibn, et al.
Publicado: (2025)
por: Hamid, Jubayer Ibn, et al.
Publicado: (2025)
ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
por: Liu, Zexi, et al.
Publicado: (2025)
por: Liu, Zexi, et al.
Publicado: (2025)
Reinforcement Learning with Quasi-Hyperbolic Discounting
por: Eshwar, S. R., et al.
Publicado: (2024)
por: Eshwar, S. R., et al.
Publicado: (2024)
Ejemplares similares
-
MALT: Improving Reasoning with Multi-Agent LLM Training
por: Motwani, Sumeet Ramesh, et al.
Publicado: (2024) -
REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites
por: Garg, Divyansh, et al.
Publicado: (2025) -
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
por: Xiang, Violet, et al.
Publicado: (2025) -
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
por: Qu, Yuxiao, et al.
Publicado: (2024) -
Contrastive Preference Learning: Learning from Human Feedback without RL
por: Hejna, Joey, et al.
Publicado: (2023)