Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yongjin, Kang, Sinjae, Lee, Juyong, Lee, Dongjun, Yun, Se-Young, Lee, Kimin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automated Filtering of Human Feedback Data for Aligning Text-to-Image Diffusion Models
by: Yang, Yongjin, et al.
Published: (2024)
by: Yang, Yongjin, et al.
Published: (2024)
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
by: Yang, Yongjin, et al.
Published: (2025)
by: Yang, Yongjin, et al.
Published: (2025)
WarmPrior: Straightening Flow-Matching Policies with Temporal Priors
by: Kang, Sinjae, et al.
Published: (2026)
by: Kang, Sinjae, et al.
Published: (2026)
Benchmarking Mobile Device Control Agents across Diverse Configurations
by: Lee, Juyong, et al.
Published: (2024)
by: Lee, Juyong, et al.
Published: (2024)
Understanding Impact of Human Feedback via Influence Functions
by: Min, Taywon, et al.
Published: (2025)
by: Min, Taywon, et al.
Published: (2025)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
by: Kim, Jaekyeom, et al.
Published: (2024)
by: Kim, Jaekyeom, et al.
Published: (2024)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
by: Hahm, Dongyoon, et al.
Published: (2026)
by: Hahm, Dongyoon, et al.
Published: (2026)
State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital Living
by: Choi, Juheon, et al.
Published: (2025)
by: Choi, Juheon, et al.
Published: (2025)
Learning to Generate Unit Test via Adversarial Reinforcement Learning
by: Lee, Dongjun, et al.
Published: (2025)
by: Lee, Dongjun, et al.
Published: (2025)
Self-Training Elicits Concise Reasoning in Large Language Models
by: Munkhbat, Tergel, et al.
Published: (2025)
by: Munkhbat, Tergel, et al.
Published: (2025)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
by: Koh, Woosung, et al.
Published: (2024)
by: Koh, Woosung, et al.
Published: (2024)
Automated Kernel Discovery Towards Understanding High-dimensional Bayesian Optimization
by: Yun, Taeyoung, et al.
Published: (2026)
by: Yun, Taeyoung, et al.
Published: (2026)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
by: Kim, Changyeon, et al.
Published: (2025)
by: Kim, Changyeon, et al.
Published: (2025)
IgPose: A Generative Data-Augmented Pipeline for Robust Immunoglobulin-Antigen Binding Prediction
by: Bui, Tien-Cuong, et al.
Published: (2026)
by: Bui, Tien-Cuong, et al.
Published: (2026)
Enhancing LLM Agent Safety via Causal Influence Prompting
by: Hahm, Dongyoon, et al.
Published: (2025)
by: Hahm, Dongyoon, et al.
Published: (2025)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
by: Lee, Jung Hyun, et al.
Published: (2024)
by: Lee, Jung Hyun, et al.
Published: (2024)
Leveraging Human Feedback for Semantically-Relevant Skill Discovery
by: Hussonnois, Maxence, et al.
Published: (2026)
by: Hussonnois, Maxence, et al.
Published: (2026)
What is the Alignment Objective of GRPO?
by: Vojnovic, Milan, et al.
Published: (2025)
by: Vojnovic, Milan, et al.
Published: (2025)
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
by: Hahm, Dongyoon, et al.
Published: (2025)
by: Hahm, Dongyoon, et al.
Published: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
Enhanced Conditional Generation of Double Perovskite by Knowledge-Guided Language Model Feedback
by: Lee, Inhyo, et al.
Published: (2025)
by: Lee, Inhyo, et al.
Published: (2025)
Unsupervised Skill Discovery as Exploration for Learning Agile Locomotion
by: Rho, Seungeun, et al.
Published: (2025)
by: Rho, Seungeun, et al.
Published: (2025)
Adversarial Reinforcement Learning Framework for ESP Cheater Simulation
by: Park, Inkyu, et al.
Published: (2025)
by: Park, Inkyu, et al.
Published: (2025)
Quality over Quantity: Demonstration Curation via Influence Functions for Data-Centric Robot Learning
by: Lee, Haeone, et al.
Published: (2026)
by: Lee, Haeone, et al.
Published: (2026)
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy
by: Doo, JaeHyeok, et al.
Published: (2026)
by: Doo, JaeHyeok, et al.
Published: (2026)
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
by: Choi, Moonseok, et al.
Published: (2023)
by: Choi, Moonseok, et al.
Published: (2023)
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
by: Lee, Gihun, et al.
Published: (2024)
by: Lee, Gihun, et al.
Published: (2024)
FlickerFusion: Intra-trajectory Domain Generalizing Multi-Agent RL
by: Koh, Woosung, et al.
Published: (2024)
by: Koh, Woosung, et al.
Published: (2024)
Introducing Spectral Attention for Long-Range Dependency in Time Series Forecasting
by: Kang, Bong Gyun, et al.
Published: (2024)
by: Kang, Bong Gyun, et al.
Published: (2024)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023)
by: Lee, Jung Hyun, et al.
Published: (2023)
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
by: Lee, Sanghyun, et al.
Published: (2025)
by: Lee, Sanghyun, et al.
Published: (2025)
AMPED: Adaptive Multi-objective Projection for balancing Exploration and skill Diversification
by: Cho, Geonwoo, et al.
Published: (2025)
by: Cho, Geonwoo, et al.
Published: (2025)
LLM-based Agents for Automated Confounder Discovery and Subgroup Analysis in Causal Inference
by: Lee, Po-Han, et al.
Published: (2025)
by: Lee, Po-Han, et al.
Published: (2025)
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
by: Yoon, Hyungjun, et al.
Published: (2024)
by: Yoon, Hyungjun, et al.
Published: (2024)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
by: Lee, Joonhyung, et al.
Published: (2024)
by: Lee, Joonhyung, et al.
Published: (2024)
RAG-Enhanced Collaborative LLM Agents for Drug Discovery
by: Lee, Namkyeong, et al.
Published: (2025)
by: Lee, Namkyeong, et al.
Published: (2025)
Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models
by: Kim, Kyuyoung, et al.
Published: (2024)
by: Kim, Kyuyoung, et al.
Published: (2024)
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
by: Lee, Juyong, et al.
Published: (2024)
by: Lee, Juyong, et al.
Published: (2024)
FedSOL: Stabilized Orthogonal Learning with Proximal Restrictions in Federated Learning
by: Lee, Gihun, et al.
Published: (2023)
by: Lee, Gihun, et al.
Published: (2023)
Generative Visual Code Mobile World Models
by: Koh, Woosung, et al.
Published: (2026)
by: Koh, Woosung, et al.
Published: (2026)
Similar Items
-
Automated Filtering of Human Feedback Data for Aligning Text-to-Image Diffusion Models
by: Yang, Yongjin, et al.
Published: (2024) -
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
by: Yang, Yongjin, et al.
Published: (2025) -
WarmPrior: Straightening Flow-Matching Policies with Temporal Priors
by: Kang, Sinjae, et al.
Published: (2026) -
Benchmarking Mobile Device Control Agents across Diverse Configurations
by: Lee, Juyong, et al.
Published: (2024) -
Understanding Impact of Human Feedback via Influence Functions
by: Min, Taywon, et al.
Published: (2025)