Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Siwei, Li, Yizhi, Song, Yuyang, Zhang, Wei, Wang, Yang, Batista-Navarro, Riza, Yang, Xian, Tang, Mingjie, Dai, Bryan, Yang, Jian, Lin, Chenghua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
by: Ren, Jincheng, et al.
Published: (2026)
by: Ren, Jincheng, et al.
Published: (2026)
AGRO-SQL: Agentic Group-Relative Optimization with High-Fidelity Data Synthesis
by: Yang, Cehua, et al.
Published: (2025)
by: Yang, Cehua, et al.
Published: (2025)
LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm
by: Wu, Siwei, et al.
Published: (2025)
by: Wu, Siwei, et al.
Published: (2025)
DockSmith: Scaling Reliable Coding Environments via an Agentic Docker Builder
by: Zhang, Jiaran, et al.
Published: (2026)
by: Zhang, Jiaran, et al.
Published: (2026)
Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
by: Qu, Xingwei, et al.
Published: (2024)
by: Qu, Xingwei, et al.
Published: (2024)
Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study
by: Sun, Yizheng, et al.
Published: (2025)
by: Sun, Yizheng, et al.
Published: (2025)
LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models
by: Sun, Yizheng, et al.
Published: (2025)
by: Sun, Yizheng, et al.
Published: (2025)
SWE-World: Building Software Engineering Agents in Docker-Free Environments
by: Sun, Shuang, et al.
Published: (2026)
by: Sun, Shuang, et al.
Published: (2026)
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
by: Wu, Yulong, et al.
Published: (2025)
by: Wu, Yulong, et al.
Published: (2025)
Aspect-based Sentiment Evaluation of Chess Moves (ASSESS): an NLP-based Method for Evaluating Chess Strategies from Textbooks
by: Alrdahi, Haifa, et al.
Published: (2024)
by: Alrdahi, Haifa, et al.
Published: (2024)
DocMMIR: A Framework for Document Multi-modal Information Retrieval
by: Li, Zirui, et al.
Published: (2025)
by: Li, Zirui, et al.
Published: (2025)
Multi-Docker-Eval: A `Shovel of the Gold Rush' Benchmark on Automatic Environment Building for Software Engineering
by: Fu, Kelin, et al.
Published: (2025)
by: Fu, Kelin, et al.
Published: (2025)
Arg-LLaDA: Argument Summarization via Large Language Diffusion Models and Sufficiency-Aware Refinement
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure
by: Li, Zirui, et al.
Published: (2026)
by: Li, Zirui, et al.
Published: (2026)
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
DRE: An Effective Dual-Refined Method for Integrating Small and Large Language Models in Open-Domain Dialogue Evaluation
by: Zhao, Kun, et al.
Published: (2025)
by: Zhao, Kun, et al.
Published: (2025)
Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
by: Cheng, Zihao, et al.
Published: (2026)
by: Cheng, Zihao, et al.
Published: (2026)
Investigating a Benchmark for Training-set free Evaluation of Linguistic Capabilities in Machine Reading Comprehension
by: Schlegel, Viktor, et al.
Published: (2024)
by: Schlegel, Viktor, et al.
Published: (2024)
Pay Attention to Real World Perturbations! Natural Robustness Evaluation in Machine Reading Comprehension
by: Wu, Yulong, et al.
Published: (2025)
by: Wu, Yulong, et al.
Published: (2025)
SLIDE: A Framework Integrating Small and Large Language Models for Open-Domain Dialogues Evaluation
by: Zhao, Kun, et al.
Published: (2024)
by: Zhao, Kun, et al.
Published: (2024)
Large Language Models in Argument Mining: A Survey
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Audio Contrastive-based Fine-tuning: Decoupling Representation Learning and Classification
by: Wang, Yang, et al.
Published: (2023)
by: Wang, Yang, et al.
Published: (2023)
Endless Terminals: Scaling RL Environments for Terminal Agents
by: Gandhi, Kanishk, et al.
Published: (2026)
by: Gandhi, Kanishk, et al.
Published: (2026)
TermiGen: High-Fidelity Environment and Robust Trajectory Synthesis for Terminal Agents
by: Zhu, Kaijie, et al.
Published: (2026)
by: Zhu, Kaijie, et al.
Published: (2026)
Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
by: Madusanka, Tharindu, et al.
Published: (2025)
by: Madusanka, Tharindu, et al.
Published: (2025)
$A^2Flow:$ Automating Agentic Workflow Generation via Self-Adaptive Abstraction Operators
by: Zhao, Mingming, et al.
Published: (2025)
by: Zhao, Mingming, et al.
Published: (2025)
Context as a Tool: Context Management for Long-Horizon SWE-Agents
by: Liu, Shukai, et al.
Published: (2025)
by: Liu, Shukai, et al.
Published: (2025)
CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
Learning to Generate and Evaluate Fact-checking Explanations with Transformers
by: Feher, Darius, et al.
Published: (2024)
by: Feher, Darius, et al.
Published: (2024)
Observing Micromotives and Macrobehavior of Large Language Models
by: Cheng, Yuyang, et al.
Published: (2024)
by: Cheng, Yuyang, et al.
Published: (2024)
Fast Swimming Robot Fish Under Countercurrent, Complex Trajectory, and Heavy Load Environments
by: Zhengcheng Yang, et al.
Published: (2025)
by: Zhengcheng Yang, et al.
Published: (2025)
Olbedo: An Albedo and Shading Aerial Dataset for Large-Scale Outdoor Environments
by: Song, Shuang, et al.
Published: (2026)
by: Song, Shuang, et al.
Published: (2026)
Revisiting Content-Based Music Recommendation: Efficient Feature Aggregation from Large-Scale Music Models
by: Zhou, Yizhi, et al.
Published: (2026)
by: Zhou, Yizhi, et al.
Published: (2026)
Adversarial Defence without Adversarial Defence: Enhancing Language Model Robustness via Instance-level Principal Component Removal
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
Rescaling Confidence: What Scale Design Reveals About LLM Metacognition
by: Dai, Yuyang
Published: (2026)
by: Dai, Yuyang
Published: (2026)
Follow Your Motion: A Generic Temporal Consistency Portrait Editing Framework with Trajectory Guidance
by: Yang, Haijie, et al.
Published: (2025)
by: Yang, Haijie, et al.
Published: (2025)
Contenerización en Docker
by: Ortiz Medina, Brandon
Published: (2026)
by: Ortiz Medina, Brandon
Published: (2026)
LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents
by: Peng, Xiaoxuan, et al.
Published: (2026)
by: Peng, Xiaoxuan, et al.
Published: (2026)
Distinct Critical Scaling of Quantum Fisher Information in a Quantum Rabi Triangle System
by: Tang, Yuyang, et al.
Published: (2025)
by: Tang, Yuyang, et al.
Published: (2025)
Train & Constrain: Phonologically Informed Tongue-Twister Generation from Topics and Paraphrases
by: Loakman, Tyler, et al.
Published: (2024)
by: Loakman, Tyler, et al.
Published: (2024)
Similar Items
-
A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
by: Ren, Jincheng, et al.
Published: (2026) -
AGRO-SQL: Agentic Group-Relative Optimization with High-Fidelity Data Synthesis
by: Yang, Cehua, et al.
Published: (2025) -
LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm
by: Wu, Siwei, et al.
Published: (2025) -
DockSmith: Scaling Reliable Coding Environments via an Agentic Docker Builder
by: Zhang, Jiaran, et al.
Published: (2026) -
Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
by: Qu, Xingwei, et al.
Published: (2024)