On Problems of Implicit Context Compression for Software Engineering Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gelvan, Kirill, Slinko, Igor, Steinbauer, Felix, Bogomolov, Egor, Kofler, Florian, Zharov, Yaroslav |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Step Rejection Fine-Tuning: A Practical Distillation Recipe
par: Slinko, Igor, et autres
Publié: (2026)
par: Slinko, Igor, et autres
Publié: (2026)
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
par: Lindenbauer, Tobias, et autres
Publié: (2025)
par: Lindenbauer, Tobias, et autres
Publié: (2025)
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
par: Lindenbauer, Tobias, et autres
Publié: (2025)
par: Lindenbauer, Tobias, et autres
Publié: (2025)
PIPer: On-Device Environment Setup via Online Reinforcement Learning
par: Kovrigin, Alexander, et autres
Publié: (2025)
par: Kovrigin, Alexander, et autres
Publié: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
par: Galimzyanov, Timur, et autres
Publié: (2024)
par: Galimzyanov, Timur, et autres
Publié: (2024)
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
par: Kovrigin, Alexander, et autres
Publié: (2024)
par: Kovrigin, Alexander, et autres
Publié: (2024)
Challenge on Optimization of Context Collection for Code Completion
par: Ustalov, Dmitry, et autres
Publié: (2025)
par: Ustalov, Dmitry, et autres
Publié: (2025)
EnvBench: A Benchmark for Automated Environment Setup
par: Eliseeva, Aleksandra, et autres
Publié: (2025)
par: Eliseeva, Aleksandra, et autres
Publié: (2025)
Dynamic Retrieval-Augmented Generation
par: Shapkin, Anton, et autres
Publié: (2023)
par: Shapkin, Anton, et autres
Publié: (2023)
Tool-Augmented LLMs as a Universal Interface for IDEs
par: Zharov, Yaroslav, et autres
Publié: (2024)
par: Zharov, Yaroslav, et autres
Publié: (2024)
Agentless: Demystifying LLM-based Software Engineering Agents
par: Xia, Chunqiu Steven, et autres
Publié: (2024)
par: Xia, Chunqiu Steven, et autres
Publié: (2024)
Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents
par: Zhang, Kexun, et autres
Publié: (2024)
par: Zhang, Kexun, et autres
Publié: (2024)
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
par: Xia, Chunqiu Steven, et autres
Publié: (2025)
par: Xia, Chunqiu Steven, et autres
Publié: (2025)
Agents in Software Engineering: Survey, Landscape, and Vision
par: Wang, Yanlin, et autres
Publié: (2024)
par: Wang, Yanlin, et autres
Publié: (2024)
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
par: Bogomolov, Egor, et autres
Publié: (2024)
par: Bogomolov, Egor, et autres
Publié: (2024)
SWE-Protégé: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents
par: Kon, Patrick Tser Jern, et autres
Publié: (2026)
par: Kon, Patrick Tser Jern, et autres
Publié: (2026)
SWE-smith: Scaling Data for Software Engineering Agents
par: Yang, John, et autres
Publié: (2025)
par: Yang, John, et autres
Publié: (2025)
Stack Trace Deduplication: Faster, More Accurately, and in More Realistic Scenarios
par: Shibaev, Egor, et autres
Publié: (2024)
par: Shibaev, Egor, et autres
Publié: (2024)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
par: Sonwane, Atharv, et autres
Publié: (2026)
par: Sonwane, Atharv, et autres
Publié: (2026)
Experiential Co-Learning of Software-Developing Agents
par: Qian, Chen, et autres
Publié: (2023)
par: Qian, Chen, et autres
Publié: (2023)
An Approach for Auto Generation of Labeling Functions for Software Engineering Chatbots
par: Alor, Ebube, et autres
Publié: (2024)
par: Alor, Ebube, et autres
Publié: (2024)
Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
par: Kuang, Jiayi, et autres
Publié: (2025)
par: Kuang, Jiayi, et autres
Publié: (2025)
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
par: Guo, Xuehang, et autres
Publié: (2025)
par: Guo, Xuehang, et autres
Publié: (2025)
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
par: Shetty, Manish, et autres
Publié: (2025)
par: Shetty, Manish, et autres
Publié: (2025)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
par: Wei, Yuxiang, et autres
Publié: (2025)
par: Wei, Yuxiang, et autres
Publié: (2025)
From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future
par: Jin, Haolin, et autres
Publié: (2024)
par: Jin, Haolin, et autres
Publié: (2024)
Kotlin ML Pack: Technical Report
par: Titov, Sergey, et autres
Publié: (2024)
par: Titov, Sergey, et autres
Publié: (2024)
Applying Large Language Models API to Issue Classification Problem
par: Aracena, Gabriel, et autres
Publié: (2024)
par: Aracena, Gabriel, et autres
Publié: (2024)
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
par: Wang, Lilin, et autres
Publié: (2025)
par: Wang, Lilin, et autres
Publié: (2025)
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
par: Yang, John, et autres
Publié: (2024)
par: Yang, John, et autres
Publié: (2024)
Introduction to Analytical Software Engineering Design Paradigm
par: Houichime, Tarik, et autres
Publié: (2025)
par: Houichime, Tarik, et autres
Publié: (2025)
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
par: Glukhov, Evgeniy, et autres
Publié: (2025)
par: Glukhov, Evgeniy, et autres
Publié: (2025)
Unified Software Engineering Agent as AI Software Engineer
par: Applis, Leonhard, et autres
Publié: (2025)
par: Applis, Leonhard, et autres
Publié: (2025)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
par: Ding, Yifeng, et autres
Publié: (2026)
par: Ding, Yifeng, et autres
Publié: (2026)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
par: Qiu, Jielin, et autres
Publié: (2025)
par: Qiu, Jielin, et autres
Publié: (2025)
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
par: Arora, Avi, et autres
Publié: (2025)
par: Arora, Avi, et autres
Publié: (2025)
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
par: Yuan, Danlong, et autres
Publié: (2026)
par: Yuan, Danlong, et autres
Publié: (2026)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
par: Cipollone, Daniele, et autres
Publié: (2025)
par: Cipollone, Daniele, et autres
Publié: (2025)
Multi-Agent Coordinated Rename Refactoring
par: Bellur, Abhiram, et autres
Publié: (2026)
par: Bellur, Abhiram, et autres
Publié: (2026)
From Code Generation to Software Testing: AI Copilot with Context-Based RAG
par: Wang, Yuchen, et autres
Publié: (2025)
par: Wang, Yuchen, et autres
Publié: (2025)
Documents similaires
-
Step Rejection Fine-Tuning: A Practical Distillation Recipe
par: Slinko, Igor, et autres
Publié: (2026) -
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
par: Lindenbauer, Tobias, et autres
Publié: (2025) -
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
par: Lindenbauer, Tobias, et autres
Publié: (2025) -
PIPer: On-Device Environment Setup via Online Reinforcement Learning
par: Kovrigin, Alexander, et autres
Publié: (2025) -
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
par: Galimzyanov, Timur, et autres
Publié: (2024)