Tracking Capabilities for Safer Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Odersky, Martin, Zhao, Yaoyu, Xu, Yichen, Bračevac, Oliver, Pham, Cao Nguyen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LACUNA: Safe Agents as Recursive Program Holes
by: Zhao, Yaoyu, et al.
Published: (2026)
by: Zhao, Yaoyu, et al.
Published: (2026)
What's in the Box: Ergonomic and Expressive Capture Tracking over Generic Data Structures (Extended Version)
by: Xu, Yichen, et al.
Published: (2025)
by: Xu, Yichen, et al.
Published: (2025)
Agentic Proof Automation: A Case Study
by: Xu, Yichen, et al.
Published: (2026)
by: Xu, Yichen, et al.
Published: (2026)
Raw Pointer Rewriting with LLMs for Translating C to Safer Rust
by: Gao, Yifei, et al.
Published: (2025)
by: Gao, Yifei, et al.
Published: (2025)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
First-Class Refinement Types for Scala
by: Bovel, Matt, et al.
Published: (2026)
by: Bovel, Matt, et al.
Published: (2026)
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
by: Dai, Hankun, et al.
Published: (2025)
by: Dai, Hankun, et al.
Published: (2025)
FPMoE: A Sparse Mixture-of-Experts Approach to Functional Code Generation
by: Pham, Loc, et al.
Published: (2026)
by: Pham, Loc, et al.
Published: (2026)
Modeling Reachability Types with Logical Relations
by: Bao, Yuyan, et al.
Published: (2023)
by: Bao, Yuyan, et al.
Published: (2023)
Provable Coordination for LLM Agents via Message Sequence Charts
by: Bollig, Benedikt, et al.
Published: (2026)
by: Bollig, Benedikt, et al.
Published: (2026)
Language-Integrated Recursive Queries (Full Version)
by: Herlihy, Anna, et al.
Published: (2025)
by: Herlihy, Anna, et al.
Published: (2025)
Assessing GPT-4-Vision's Capabilities in UML-Based Code Generation
by: Antal, Gábor, et al.
Published: (2024)
by: Antal, Gábor, et al.
Published: (2024)
AutoPDL: Automatic Prompt Optimization for LLM Agents
by: Spiess, Claudio, et al.
Published: (2025)
by: Spiess, Claudio, et al.
Published: (2025)
Testing and Understanding Erroneous Planning in LLM Agents through Synthesized User Inputs
by: Ji, Zhenlan, et al.
Published: (2024)
by: Ji, Zhenlan, et al.
Published: (2024)
DriftScript: A Domain-Specific Language for Programming Non-Axiomatic Reasoning Agents
by: Brady, Seamus
Published: (2026)
by: Brady, Seamus
Published: (2026)
Safer-Instruct: Aligning Language Models with Automated Preference Data
by: Shi, Taiwei, et al.
Published: (2023)
by: Shi, Taiwei, et al.
Published: (2023)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
by: Duston, Titouan, et al.
Published: (2025)
by: Duston, Titouan, et al.
Published: (2025)
Controllable and Reliable Knowledge-Intensive Task-Oriented Conversational Agents with Declarative Genie Worksheets
by: Joshi, Harshit, et al.
Published: (2024)
by: Joshi, Harshit, et al.
Published: (2024)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
by: Zhu, Hengchuan, et al.
Published: (2025)
by: Zhu, Hengchuan, et al.
Published: (2025)
Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning
by: Pham, Anh, et al.
Published: (2025)
by: Pham, Anh, et al.
Published: (2025)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
by: Wang, Peiding, et al.
Published: (2025)
by: Wang, Peiding, et al.
Published: (2025)
AIOS Compiler: LLM as Interpreter for Natural Language Programming and Flow Programming of AI Agents
by: Xu, Shuyuan, et al.
Published: (2024)
by: Xu, Shuyuan, et al.
Published: (2024)
Adaptive Recursive Query Optimization
by: Herlihy, Anna, et al.
Published: (2023)
by: Herlihy, Anna, et al.
Published: (2023)
XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
by: Dong, Yixin, et al.
Published: (2024)
by: Dong, Yixin, et al.
Published: (2024)
Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
by: Xu, Xu, et al.
Published: (2025)
by: Xu, Xu, et al.
Published: (2025)
Data Petri Nets meet Probabilistic Programming (Extended version)
by: Kuhn, Martin, et al.
Published: (2024)
by: Kuhn, Martin, et al.
Published: (2024)
PDL: A Declarative Prompt Programming Language
by: Vaziri, Mandana, et al.
Published: (2024)
by: Vaziri, Mandana, et al.
Published: (2024)
VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable Code
by: Zeng, Lingfei, et al.
Published: (2025)
by: Zeng, Lingfei, et al.
Published: (2025)
Language-Based Agent Control
by: Zhou, Timothy, et al.
Published: (2026)
by: Zhou, Timothy, et al.
Published: (2026)
CACA Agent: Capability Collaboration based AI Agent
by: Xu, Peng, et al.
Published: (2024)
by: Xu, Peng, et al.
Published: (2024)
CodeV: Empowering LLMs with HDL Generation through Multi-Level Summarization
by: Zhao, Yang, et al.
Published: (2024)
by: Zhao, Yang, et al.
Published: (2024)
SGLang: Efficient Execution of Structured Language Model Programs
by: Zheng, Lianmin, et al.
Published: (2023)
by: Zheng, Lianmin, et al.
Published: (2023)
AC4A: Access Control for Agents
by: Sharma, Reshabh K, et al.
Published: (2026)
by: Sharma, Reshabh K, et al.
Published: (2026)
PAuth - Precise Task-Scoped Authorization For Agents
by: Sharma, Reshabh K, et al.
Published: (2026)
by: Sharma, Reshabh K, et al.
Published: (2026)
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
by: Cao, Jialun, et al.
Published: (2025)
by: Cao, Jialun, et al.
Published: (2025)
Pel, A Programming Language for Orchestrating AI Agents
by: Mohammadi, Behnam
Published: (2025)
by: Mohammadi, Behnam
Published: (2025)
Models That Know How Evaluations Are Designed Score Safer
by: Deckenbach, Katharina, et al.
Published: (2026)
by: Deckenbach, Katharina, et al.
Published: (2026)
Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
by: Chen, Hongyu, et al.
Published: (2025)
by: Chen, Hongyu, et al.
Published: (2025)
InvThink: Premortem Reasoning for Safer Language Models
by: Kim, Yubin, et al.
Published: (2025)
by: Kim, Yubin, et al.
Published: (2025)
Relational Programming with Foundation Models
by: Li, Ziyang, et al.
Published: (2024)
by: Li, Ziyang, et al.
Published: (2024)
Similar Items
-
LACUNA: Safe Agents as Recursive Program Holes
by: Zhao, Yaoyu, et al.
Published: (2026) -
What's in the Box: Ergonomic and Expressive Capture Tracking over Generic Data Structures (Extended Version)
by: Xu, Yichen, et al.
Published: (2025) -
Agentic Proof Automation: A Case Study
by: Xu, Yichen, et al.
Published: (2026) -
Raw Pointer Rewriting with LLMs for Translating C to Safer Rust
by: Gao, Yifei, et al.
Published: (2025) -
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)