It's LIT! Reliability-Optimized LLMs with Inspectable Tools
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Ruixin, Donnelly, Jon, Guo, Zhicheng, Khalighinejad, Ghazal, Huang, Haiyang, Barnett, Alina Jade, Rudin, Cynthia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rashomon Sets for Prototypical-Part Networks: Editing Interpretable Models in Real-Time
von: Donnelly, Jon, et al.
Veröffentlicht: (2025)
von: Donnelly, Jon, et al.
Veröffentlicht: (2025)
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
von: LeVine, Will, et al.
Veröffentlicht: (2026)
von: LeVine, Will, et al.
Veröffentlicht: (2026)
K-ASTRO: Structure-Aware Adaptation of LLMs for Code Vulnerability Detection
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
Towards Counterfactual Explanation and Assertion Inference for CPS Debugging
von: Ghazal, Zaid, et al.
Veröffentlicht: (2026)
von: Ghazal, Zaid, et al.
Veröffentlicht: (2026)
Monitoring Machine Learning Systems: A Multivocal Literature Review
von: Naveed, Hira, et al.
Veröffentlicht: (2025)
von: Naveed, Hira, et al.
Veröffentlicht: (2025)
Good Tools are Half the Work: Tool Usage in Deep Learning Projects
von: Panourgia, Evangelia, et al.
Veröffentlicht: (2023)
von: Panourgia, Evangelia, et al.
Veröffentlicht: (2023)
QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
von: Ke, Changxin, et al.
Veröffentlicht: (2026)
von: Ke, Changxin, et al.
Veröffentlicht: (2026)
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
von: Zhou, Zenghui, et al.
Veröffentlicht: (2026)
von: Zhou, Zenghui, et al.
Veröffentlicht: (2026)
Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs
von: Iskander, Shadi, et al.
Veröffentlicht: (2024)
von: Iskander, Shadi, et al.
Veröffentlicht: (2024)
PromSec: Prompt Optimization for Secure Generation of Functional Source Code with Large Language Models (LLMs)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
von: Jin, Yiyang, et al.
Veröffentlicht: (2025)
von: Jin, Yiyang, et al.
Veröffentlicht: (2025)
Disproving Program Equivalence with LLMs
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2025)
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2025)
Towards Continuous Assurance Case Creation for ADS with the Evidential Tool Bus
von: Sorokin, Lev, et al.
Veröffentlicht: (2024)
von: Sorokin, Lev, et al.
Veröffentlicht: (2024)
Teaching Code Refactoring Using LLMs
von: Khairnar, Anshul, et al.
Veröffentlicht: (2025)
von: Khairnar, Anshul, et al.
Veröffentlicht: (2025)
Towards Verified Code Reasoning by LLMs
von: Sistla, Meghana, et al.
Veröffentlicht: (2025)
von: Sistla, Meghana, et al.
Veröffentlicht: (2025)
A Machine Learning-Based Error Mitigation Approach For Reliable Software Development On IBM'S Quantum Computers
von: Muqeet, Asmar, et al.
Veröffentlicht: (2024)
von: Muqeet, Asmar, et al.
Veröffentlicht: (2024)
Ensuring Reliability in Programming Knowledge Tracing: A Re-evaluation of Attention-augmented Models and Experimental Protocols
von: Kim, Jaewook, et al.
Veröffentlicht: (2026)
von: Kim, Jaewook, et al.
Veröffentlicht: (2026)
A Semantic-based Optimization Approach for Repairing LLMs: Case Study on Code Generation
von: Gu, Jian, et al.
Veröffentlicht: (2025)
von: Gu, Jian, et al.
Veröffentlicht: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
Understanding Robustness of Model Editing in Code LLMs
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
Towards MLOps: A DevOps Tools Recommender System for Machine Learning System
von: Shah, Pir Sami Ullah, et al.
Veröffentlicht: (2024)
von: Shah, Pir Sami Ullah, et al.
Veröffentlicht: (2024)
DesCartes Builder: A Tool to Develop Machine-Learning Based Digital Twins
von: de Conto, Eduardo, et al.
Veröffentlicht: (2025)
von: de Conto, Eduardo, et al.
Veröffentlicht: (2025)
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
von: Lu, Yifei, et al.
Veröffentlicht: (2025)
von: Lu, Yifei, et al.
Veröffentlicht: (2025)
Breaking the Silence: the Threats of Using LLMs in Software Engineering
von: Sallou, June, et al.
Veröffentlicht: (2023)
von: Sallou, June, et al.
Veröffentlicht: (2023)
CigaR: Cost-efficient Program Repair with LLMs
von: Hidvégi, Dávid, et al.
Veröffentlicht: (2024)
von: Hidvégi, Dávid, et al.
Veröffentlicht: (2024)
Automating API Documentation with LLMs: A BERTopic Approach
von: Naghshzan, AmirHossein
Veröffentlicht: (2025)
von: Naghshzan, AmirHossein
Veröffentlicht: (2025)
ThrowBench: Benchmarking LLMs by Predicting Runtime Exceptions
von: Prenner, Julian Aron, et al.
Veröffentlicht: (2025)
von: Prenner, Julian Aron, et al.
Veröffentlicht: (2025)
Unsupervised Evaluation of Code LLMs with Round-Trip Correctness
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2024)
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2024)
iServe: An Intent-based Serving System for LLMs
von: Liakopoulos, Dimitrios, et al.
Veröffentlicht: (2025)
von: Liakopoulos, Dimitrios, et al.
Veröffentlicht: (2025)
Renaissance of Literate Programming in the Era of LLMs: Enhancing LLM-Based Code Generation in Large-Scale Projects
von: Zhang, Wuyang, et al.
Veröffentlicht: (2024)
von: Zhang, Wuyang, et al.
Veröffentlicht: (2024)
Finding Missed Code Size Optimizations in Compilers using LLMs
von: Italiano, Davide, et al.
Veröffentlicht: (2024)
von: Italiano, Davide, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
SkillScope: A Tool to Predict Fine-Grained Skills Needed to Solve Issues on GitHub
von: Carter, Benjamin C., et al.
Veröffentlicht: (2025)
von: Carter, Benjamin C., et al.
Veröffentlicht: (2025)
ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs
von: Ding, Peng, et al.
Veröffentlicht: (2025)
von: Ding, Peng, et al.
Veröffentlicht: (2025)
Deformable ProtoPNet: An Interpretable Image Classifier Using Deformable Prototypes
von: Donnelly, Jon, et al.
Veröffentlicht: (2021)
von: Donnelly, Jon, et al.
Veröffentlicht: (2021)
TritonRL: Training LLMs to Think and Code Triton Without Cheating
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
von: Haque, Mirazul, et al.
Veröffentlicht: (2025)
von: Haque, Mirazul, et al.
Veröffentlicht: (2025)
SkVM: Revisiting Language VM for Skills across Heterogenous LLMs and Harnesses
von: Chen, Le, et al.
Veröffentlicht: (2026)
von: Chen, Le, et al.
Veröffentlicht: (2026)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
von: Bhargava, Vaishnavi, et al.
Veröffentlicht: (2024)
von: Bhargava, Vaishnavi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Rashomon Sets for Prototypical-Part Networks: Editing Interpretable Models in Real-Time
von: Donnelly, Jon, et al.
Veröffentlicht: (2025) -
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
von: LeVine, Will, et al.
Veröffentlicht: (2026) -
K-ASTRO: Structure-Aware Adaptation of LLMs for Code Vulnerability Detection
von: Zhang, Yifan, et al.
Veröffentlicht: (2022) -
Towards Counterfactual Explanation and Assertion Inference for CPS Debugging
von: Ghazal, Zaid, et al.
Veröffentlicht: (2026) -
Monitoring Machine Learning Systems: A Multivocal Literature Review
von: Naveed, Hira, et al.
Veröffentlicht: (2025)