Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.27905 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915917267468288 |
|---|---|
| author | Cruz, Christopher |
| author_facet | Cruz, Christopher |
| contents | We present ATLAS-RTC, a runtime control system for autoregressive language models that enforces structured output during decoding. ATLAS-RTC monitors generation at each step, detects drift from output contracts using lightweight signals, and applies targeted interventions such as biasing, masking, and rollback. Unlike post-hoc validation or static constrained decoding, it operates in a closed loop, enabling correction before errors materialize. Across structured generation and tool-calling tasks, ATLAS-RTC improves first-attempt success rates by 20 to 37.8 percentage points, with up to 88% latency reduction in failure-dominated settings. Results show that many failures arise from decoding artifacts rather than task misunderstanding, motivating runtime control as a distinct layer in LLM systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_27905 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ATLAS-RTC: Closing the Loop on LLM Agent Output with Token-Level Runtime Control Cruz, Christopher Machine Learning I.2.8 We present ATLAS-RTC, a runtime control system for autoregressive language models that enforces structured output during decoding. ATLAS-RTC monitors generation at each step, detects drift from output contracts using lightweight signals, and applies targeted interventions such as biasing, masking, and rollback. Unlike post-hoc validation or static constrained decoding, it operates in a closed loop, enabling correction before errors materialize. Across structured generation and tool-calling tasks, ATLAS-RTC improves first-attempt success rates by 20 to 37.8 percentage points, with up to 88% latency reduction in failure-dominated settings. Results show that many failures arise from decoding artifacts rather than task misunderstanding, motivating runtime control as a distinct layer in LLM systems. |
| title | ATLAS-RTC: Closing the Loop on LLM Agent Output with Token-Level Runtime Control |
| topic | Machine Learning I.2.8 |
| url | https://arxiv.org/abs/2603.27905 |