CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Yusong, Wang, Haiyang, Wu, Shuzhe, Fan, Lue, Pan, Feiyang, Zhao, Sanyuan, Tu, Dandan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World
by: Lin, Yusong, et al.
Published: (2026)
by: Lin, Yusong, et al.
Published: (2026)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
by: Zhou, Qixing, et al.
Published: (2026)
by: Zhou, Qixing, et al.
Published: (2026)
Learning CLI Agents with Structured Action Credit under Selective Observation
by: Su, Haoyang, et al.
Published: (2026)
by: Su, Haoyang, et al.
Published: (2026)
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
by: Hu, Ruida, et al.
Published: (2026)
by: Hu, Ruida, et al.
Published: (2026)
CLI-RAG: A Retrieval-Augmented Framework for Clinically Structured and Context Aware Text Generation with LLMs
by: Keerthana, Garapati, et al.
Published: (2025)
by: Keerthana, Garapati, et al.
Published: (2025)
Render CLI チュートリアル
by: Kawai, Katsuhiko
Published: (2026)
by: Kawai, Katsuhiko
Published: (2026)
EconGym: A Scalable AI Testbed with Diverse Economic Tasks
by: Mi, Qirui, et al.
Published: (2025)
by: Mi, Qirui, et al.
Published: (2025)
MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
by: Wang, Bowen, et al.
Published: (2026)
by: Wang, Bowen, et al.
Published: (2026)
GEM: A Gym for Agentic LLMs
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
by: Agarwal, Anmol, et al.
Published: (2026)
by: Agarwal, Anmol, et al.
Published: (2026)
PeersimGym: An Environment for Solving the Task Offloading Problem with Reinforcement Learning
by: Metelo, Frederico, et al.
Published: (2024)
by: Metelo, Frederico, et al.
Published: (2024)
ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents
by: Savadikar, Chinmay, et al.
Published: (2026)
by: Savadikar, Chinmay, et al.
Published: (2026)
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
pyRDDLGym: From RDDL to Gym Environments
by: Taitler, Ayal, et al.
Published: (2022)
by: Taitler, Ayal, et al.
Published: (2022)
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
by: Feng, Yukang, et al.
Published: (2026)
by: Feng, Yukang, et al.
Published: (2026)
Online Dynamic Goal Recognition in Gym Environments
by: Matan, Shamir, et al.
Published: (2025)
by: Matan, Shamir, et al.
Published: (2025)
Interactive OT Gym: A Reinforcement Learning-Based Interactive Optical tweezer (OT)-Driven Microrobotics Simulation Platform
by: Tan, Zongcai, et al.
Published: (2025)
by: Tan, Zongcai, et al.
Published: (2025)
WorldGym: World Model as An Environment for Policy Evaluation
by: Quevedo, Julian, et al.
Published: (2025)
by: Quevedo, Julian, et al.
Published: (2025)
LOGIGEN: Logic-Driven Generation of Verifiable Agentic Tasks
by: Zeng, Yucheng, et al.
Published: (2026)
by: Zeng, Yucheng, et al.
Published: (2026)
RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
by: Li, Ruiying, et al.
Published: (2026)
by: Li, Ruiying, et al.
Published: (2026)
Gym-Anything: Turn any Software into an Agent Environment
by: Aggarwal, Pranjal, et al.
Published: (2026)
by: Aggarwal, Pranjal, et al.
Published: (2026)
EO-Gym: A Multimodal, Interactive Environment for Earth Observation Agents
by: Ma, Sai, et al.
Published: (2026)
by: Ma, Sai, et al.
Published: (2026)
Toward Scalable Terminal Task Synthesis via Skill Graphs
by: Fan, Zhiyuan, et al.
Published: (2026)
by: Fan, Zhiyuan, et al.
Published: (2026)
HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
by: Chen, Hongzheng, et al.
Published: (2025)
by: Chen, Hongzheng, et al.
Published: (2025)
EduGym: An Environment and Notebook Suite for Reinforcement Learning Education
by: Moerland, Thomas M., et al.
Published: (2023)
by: Moerland, Thomas M., et al.
Published: (2023)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
by: Khan, Zaid, et al.
Published: (2024)
by: Khan, Zaid, et al.
Published: (2024)
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
by: Zhang, Xichen, et al.
Published: (2026)
by: Zhang, Xichen, et al.
Published: (2026)
ClawGym: A Scalable Framework for Building Effective Claw Agents
by: Bai, Fei, et al.
Published: (2026)
by: Bai, Fei, et al.
Published: (2026)
STAMP: Scalable Task And Model-agnostic Collaborative Perception
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
Towards Agentic Self-Learning LLMs in Search Environment
by: Sun, Wangtao, et al.
Published: (2025)
by: Sun, Wangtao, et al.
Published: (2025)
DeepPresenter: Environment-Grounded Reflection for Agentic Presentation Generation
by: Zheng, Hao, et al.
Published: (2026)
by: Zheng, Hao, et al.
Published: (2026)
UserBench: An Interactive Gym Environment for User-Centric Agents
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
RoboGene: Boosting VLA Pre-training via Diversity-Driven Agentic Framework for Real-World Task Generation
by: Zhang, Yixue, et al.
Published: (2026)
by: Zhang, Yixue, et al.
Published: (2026)
AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
by: Xi, Zhiheng, et al.
Published: (2024)
by: Xi, Zhiheng, et al.
Published: (2024)
CubeRobot: Grounding Language in Rubik's Cube Manipulation via Vision-Language Model
by: Wang, Feiyang, et al.
Published: (2025)
by: Wang, Feiyang, et al.
Published: (2025)
NS-Gym: Open-Source Simulation Environments and Benchmarks for Non-Stationary Markov Decision Processes
by: Keplinger, Nathaniel S., et al.
Published: (2025)
by: Keplinger, Nathaniel S., et al.
Published: (2025)
6GAgentGym: Tool Use, Data Synthesis, and Agentic Learning for Network Management
by: Chen, Jiao, et al.
Published: (2026)
by: Chen, Jiao, et al.
Published: (2026)
AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning
by: Sun, Jingbo, et al.
Published: (2026)
by: Sun, Jingbo, et al.
Published: (2026)
Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning
by: Chen, Chishui, et al.
Published: (2026)
by: Chen, Chishui, et al.
Published: (2026)
Similar Items
-
Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World
by: Lin, Yusong, et al.
Published: (2026) -
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
by: Zhou, Qixing, et al.
Published: (2026) -
Learning CLI Agents with Structured Action Credit under Selective Observation
by: Su, Haoyang, et al.
Published: (2026) -
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
by: Hu, Ruida, et al.
Published: (2026) -
CLI-RAG: A Retrieval-Augmented Framework for Clinically Structured and Context Aware Text Generation with LLMs
by: Keerthana, Garapati, et al.
Published: (2025)