Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?
Fuente:
arXiv
Saved in:
| Main Authors: | Garg, Spandan, Nitin, Vikram, Huang, Yufan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs?
by: Garg, Spandan, et al.
Published: (2026)
by: Garg, Spandan, et al.
Published: (2026)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
by: Garg, Spandan, et al.
Published: (2025)
by: Garg, Spandan, et al.
Published: (2025)
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
by: Liang, Shanchao, et al.
Published: (2025)
by: Liang, Shanchao, et al.
Published: (2025)
Can LLMs Replace Humans During Code Chunking?
by: Glasz, Christopher, et al.
Published: (2025)
by: Glasz, Christopher, et al.
Published: (2025)
PerfBench: Can Agents Resolve Real-World Performance Bugs?
by: Garg, Spandan, et al.
Published: (2025)
by: Garg, Spandan, et al.
Published: (2025)
RAPGen: An Approach for Fixing Code Inefficiencies in Zero-Shot
by: Garg, Spandan, et al.
Published: (2023)
by: Garg, Spandan, et al.
Published: (2023)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Agentic Scientific Simulation: Execution-Grounded Model Construction and Reconstruction
by: Lie, Knut-Andreas, et al.
Published: (2026)
by: Lie, Knut-Andreas, et al.
Published: (2026)
Project Prometheus: Bridging the Intent Gap in Agentic Program Repair via Reverse-Engineered Executable Specifications
by: Wang, Yongchao, et al.
Published: (2026)
by: Wang, Yongchao, et al.
Published: (2026)
The Semi-Executable Stack: Agentic Software Engineering and the Expanding Scope of SE
by: Feldt, Robert, et al.
Published: (2026)
by: Feldt, Robert, et al.
Published: (2026)
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
by: Fang, Sen, et al.
Published: (2025)
by: Fang, Sen, et al.
Published: (2025)
Teaching LLMs to Learn Tool Trialing and Execution through Environment Interaction
by: Gao, Xingjie, et al.
Published: (2026)
by: Gao, Xingjie, et al.
Published: (2026)
Agentic Frameworks for Reasoning Tasks: An Empirical Study
by: Rasheed, Zeeshan, et al.
Published: (2026)
by: Rasheed, Zeeshan, et al.
Published: (2026)
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
by: Roig, JV
Published: (2025)
by: Roig, JV
Published: (2025)
Runtime-Structured Task Decomposition for Agentic Coding Systems
by: Asthana, Shubhi, et al.
Published: (2026)
by: Asthana, Shubhi, et al.
Published: (2026)
Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs
by: Sakharova, Marina, et al.
Published: (2025)
by: Sakharova, Marina, et al.
Published: (2025)
Securing the AI Frontier: Urgent Ethical and Regulatory Imperatives for AI-Driven Cybersecurity
by: Kulothungan, Vikram
Published: (2025)
by: Kulothungan, Vikram
Published: (2025)
Evaluating LLMs for Visualization Tasks
by: Khan, Saadiq Rauf, et al.
Published: (2025)
by: Khan, Saadiq Rauf, et al.
Published: (2025)
Tree-of-Code: A Hybrid Approach for Robust Complex Task Planning and Execution
by: Ni, Ziyi, et al.
Published: (2024)
by: Ni, Ziyi, et al.
Published: (2024)
Generating Structured Plan Representation of Procedures with LLMs
by: Garg, Deepeka, et al.
Published: (2025)
by: Garg, Deepeka, et al.
Published: (2025)
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
by: Weidener, Lukas, et al.
Published: (2026)
by: Weidener, Lukas, et al.
Published: (2026)
DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
by: Chen, Aili, et al.
Published: (2026)
by: Chen, Aili, et al.
Published: (2026)
SWE-Hub: A Unified Production System for Scalable, Executable Software Engineering Tasks
by: Zeng, Yucheng, et al.
Published: (2026)
by: Zeng, Yucheng, et al.
Published: (2026)
Can LLMs Generate User Stories and Assess Their Quality?
by: Quattrocchi, Giovanni, et al.
Published: (2025)
by: Quattrocchi, Giovanni, et al.
Published: (2025)
Process-Supervised Reinforcement Learning for Code Generation
by: Ye, Yufan, et al.
Published: (2025)
by: Ye, Yufan, et al.
Published: (2025)
Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations
by: Honarvar, Shahin, et al.
Published: (2026)
by: Honarvar, Shahin, et al.
Published: (2026)
Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems
by: Ouyang, Yipeng, et al.
Published: (2026)
by: Ouyang, Yipeng, et al.
Published: (2026)
RA-Gen: A Controllable Code Generation Framework Using ReAct for Multi-Agent Task Execution
by: Liu, Aofan, et al.
Published: (2025)
by: Liu, Aofan, et al.
Published: (2025)
Ambiguity Detection and Elimination in Automated Executable Process Modeling
by: Matei, Ion, et al.
Published: (2026)
by: Matei, Ion, et al.
Published: (2026)
Tree-of-Code: A Tree-Structured Exploring Framework for End-to-End Code Generation and Execution in Complex Task Handling
by: Ni, Ziyi, et al.
Published: (2024)
by: Ni, Ziyi, et al.
Published: (2024)
SHERPA: A Model-Driven Framework for Large Language Model Execution
by: Chen, Boqi, et al.
Published: (2025)
by: Chen, Boqi, et al.
Published: (2025)
DeepCode: Open Agentic Coding
by: Li, Zongwei, et al.
Published: (2025)
by: Li, Zongwei, et al.
Published: (2025)
EduBot -- Can LLMs Solve Personalized Learning and Programming Assignments?
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
Constraint-Guided Multi-Agent Decompilation for Executable Binary Recovery
by: Zhang, Yifan, et al.
Published: (2026)
by: Zhang, Yifan, et al.
Published: (2026)
Pragmos: A Process Agentic Modeling System
by: Hernández-Ávalos, Pedro-Aarón, et al.
Published: (2026)
by: Hernández-Ávalos, Pedro-Aarón, et al.
Published: (2026)
DynamicsLLM: a Dynamic Analysis-based Tool for Generating Intelligent Execution Traces Using LLMs to Detect Android Behavioural Code Smells
by: Cherief, Houcine Abdelkader, et al.
Published: (2026)
by: Cherief, Houcine Abdelkader, et al.
Published: (2026)
How Robustly do LLMs Understand Execution Semantics?
by: Spiess, Claudio, et al.
Published: (2026)
by: Spiess, Claudio, et al.
Published: (2026)
Treefix: Enabling Execution with a Tree of Prefixes
by: Souza, Beatriz, et al.
Published: (2025)
by: Souza, Beatriz, et al.
Published: (2025)
PEFA-AI: Advancing Open-source LLMs for RTL generation using Progressive Error Feedback Agentic-AI
by: Narayanan, Athma, et al.
Published: (2025)
by: Narayanan, Athma, et al.
Published: (2025)
Insights from Benchmarking Frontier Language Models on Web App Code Generation
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
Similar Items
-
Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs?
by: Garg, Spandan, et al.
Published: (2026) -
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
by: Garg, Spandan, et al.
Published: (2025) -
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
by: Liang, Shanchao, et al.
Published: (2025) -
Can LLMs Replace Humans During Code Chunking?
by: Glasz, Christopher, et al.
Published: (2025) -
PerfBench: Can Agents Resolve Real-World Performance Bugs?
by: Garg, Spandan, et al.
Published: (2025)