Lachesis: Predicting LLM Inference Accuracy using Structural Properties of Reasoning Paths
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Naryeong, Kang, Sungmin, An, Gabin, Yoo, Shin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault Localization
von: Kang, Sungmin, et al.
Veröffentlicht: (2023)
von: Kang, Sungmin, et al.
Veröffentlicht: (2023)
Identifying Bug Inducing Commits by Combining Fault Localisation and Code Change Histories
von: An, Gabin, et al.
Veröffentlicht: (2025)
von: An, Gabin, et al.
Veröffentlicht: (2025)
Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap
von: Kim, Naryeong, et al.
Veröffentlicht: (2026)
von: Kim, Naryeong, et al.
Veröffentlicht: (2026)
COSMosFL: Ensemble of Small Language Models for Fault Localisation
von: Cho, Hyunjoon, et al.
Veröffentlicht: (2025)
von: Cho, Hyunjoon, et al.
Veröffentlicht: (2025)
Finding the Needle in the Crash Stack: Industrial-Scale Crash Root Cause Localization with AutoCrashFL
von: Kang, Sungmin, et al.
Veröffentlicht: (2025)
von: Kang, Sungmin, et al.
Veröffentlicht: (2025)
METAMON: Finding Inconsistencies between Program Documentation and Behavior using Metamorphic LLM Queries
von: Lee, Hyeonseok, et al.
Veröffentlicht: (2025)
von: Lee, Hyeonseok, et al.
Veröffentlicht: (2025)
Identifying Inaccurate Descriptions in LLM-generated Code Comments via Test Execution
von: Kang, Sungmin, et al.
Veröffentlicht: (2024)
von: Kang, Sungmin, et al.
Veröffentlicht: (2024)
Beyond pip install: Evaluating LLM Agents for the Automated Installation of Python Projects
von: Milliken, Louis, et al.
Veröffentlicht: (2024)
von: Milliken, Louis, et al.
Veröffentlicht: (2024)
Capturing Semantic Flow of ML-based Systems
von: Yoo, Shin, et al.
Veröffentlicht: (2025)
von: Yoo, Shin, et al.
Veröffentlicht: (2025)
Predictive Prompt Analysis
von: Lee, Jae Yong, et al.
Veröffentlicht: (2025)
von: Lee, Jae Yong, et al.
Veröffentlicht: (2025)
Rigorous Assessment of Model Inference Accuracy using Language Cardinality
von: Clun, Donato, et al.
Veröffentlicht: (2022)
von: Clun, Donato, et al.
Veröffentlicht: (2022)
CSA-Trans: Code Structure Aware Transformer for AST
von: Oh, Saeyoon, et al.
Veröffentlicht: (2024)
von: Oh, Saeyoon, et al.
Veröffentlicht: (2024)
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
von: Yoon, Juyeon, et al.
Veröffentlicht: (2025)
von: Yoon, Juyeon, et al.
Veröffentlicht: (2025)
DANDI: Diffusion as Normative Distribution for Deep Neural Network Input
von: Kim, Somin, et al.
Veröffentlicht: (2025)
von: Kim, Somin, et al.
Veröffentlicht: (2025)
Adaptive Testing for LLM-Based Applications: A Diversity-based Approach
von: Yoon, Juyeon, et al.
Veröffentlicht: (2025)
von: Yoon, Juyeon, et al.
Veröffentlicht: (2025)
MuFF: Stable and Sensitive Post-training Mutation Testing for Deep Learning
von: Kim, Jinhan, et al.
Veröffentlicht: (2025)
von: Kim, Jinhan, et al.
Veröffentlicht: (2025)
New Formulation of DNN Statistical Mutation Killing for Ensuring Monotonicity: A Technical Report
von: Kim, Jinhan, et al.
Veröffentlicht: (2025)
von: Kim, Jinhan, et al.
Veröffentlicht: (2025)
Revisiting "Revisiting Neuron Coverage for DNN Testing: A Layer-Wise and Distribution-Aware Criterion": A Critical Review and Implications on DNN Coverage Testing
von: Kim, Jinhan, et al.
Veröffentlicht: (2026)
von: Kim, Jinhan, et al.
Veröffentlicht: (2026)
Fault Localisation and Repair for DL Systems: An Empirical Study with LLMs
von: Kim, Jinhan, et al.
Veröffentlicht: (2025)
von: Kim, Jinhan, et al.
Veröffentlicht: (2025)
Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
von: Tan, Honghao, et al.
Veröffentlicht: (2026)
von: Tan, Honghao, et al.
Veröffentlicht: (2026)
AutoCodeSherpa: Symbolic Explanations in AI Coding Agents
von: Kang, Sungmin, et al.
Veröffentlicht: (2025)
von: Kang, Sungmin, et al.
Veröffentlicht: (2025)
Structural Anchors and Reasoning Fragility:Understanding CoT Robustness in LLM4Code
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
Think Like Human Developers: Harnessing Community Knowledge for Structured Code Reasoning
von: Yang, Chengran, et al.
Veröffentlicht: (2025)
von: Yang, Chengran, et al.
Veröffentlicht: (2025)
Beyond Translation Accuracy: Addressing False Failures in LLM-Based Code Translation
von: Rabbi, Fazle, et al.
Veröffentlicht: (2026)
von: Rabbi, Fazle, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning
von: Su, Qisheng, et al.
Veröffentlicht: (2026)
von: Su, Qisheng, et al.
Veröffentlicht: (2026)
Less Is More: Measuring How LLM Involvement affects Chatbot Accuracy in Static Analysis
von: Narasimhan, Krishna
Veröffentlicht: (2026)
von: Narasimhan, Krishna
Veröffentlicht: (2026)
A First Look at Bugs in LLM Inference Engines
von: Liu, Mugeng, et al.
Veröffentlicht: (2025)
von: Liu, Mugeng, et al.
Veröffentlicht: (2025)
Improving Dynamic Specification Inference with LLM-Generated Counterexamples
von: Balestra, Agustín, et al.
Veröffentlicht: (2026)
von: Balestra, Agustín, et al.
Veröffentlicht: (2026)
Testing Refactoring Engine via Historical Bug Report driven LLM
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
On Reasoning-Centric LLM-based Automated Theorem Proving
von: Sun, Yican, et al.
Veröffentlicht: (2026)
von: Sun, Yican, et al.
Veröffentlicht: (2026)
Minimizing False Positives in Static Bug Detection via LLM-Enhanced Path Feasibility Analysis
von: Du, Xueying, et al.
Veröffentlicht: (2025)
von: Du, Xueying, et al.
Veröffentlicht: (2025)
Using Causal Inference to Test Systems with Hidden and Interacting Variables: An Evaluative Case Study
von: Foster, Michael, et al.
Veröffentlicht: (2025)
von: Foster, Michael, et al.
Veröffentlicht: (2025)
PALM: Path-aware LLM-based Test Generation with Comprehension
von: Wu, Yaoxuan, et al.
Veröffentlicht: (2025)
von: Wu, Yaoxuan, et al.
Veröffentlicht: (2025)
LLM-Guided Issue Generation from Uncovered Code Segments
von: Pressato, Diany, et al.
Veröffentlicht: (2026)
von: Pressato, Diany, et al.
Veröffentlicht: (2026)
LLM-based Unit Test Generation via Property Retrieval
von: Zhang, Zhe, et al.
Veröffentlicht: (2024)
von: Zhang, Zhe, et al.
Veröffentlicht: (2024)
Identifying Root Causes of Null Pointer Exceptions with Logical Inferences
von: Kim, Jindae, et al.
Veröffentlicht: (2024)
von: Kim, Jindae, et al.
Veröffentlicht: (2024)
ProfInfer: An eBPF-based Fine-Grained LLM Inference Profiler
von: Zou, Bohua, et al.
Veröffentlicht: (2026)
von: Zou, Bohua, et al.
Veröffentlicht: (2026)
Neurosymbolic Architectural Reasoning: Towards Formal Analysis through Neural Software Architecture Inference
von: Herbold, Steffen, et al.
Veröffentlicht: (2025)
von: Herbold, Steffen, et al.
Veröffentlicht: (2025)
Effective LLM Code Refinement via Property-Oriented and Structurally Minimal Feedback
von: He, Lehan, et al.
Veröffentlicht: (2025)
von: He, Lehan, et al.
Veröffentlicht: (2025)
Beyond Rules: LLM-Powered Linting for Quantum Programs
von: Cassieri, Pietro, et al.
Veröffentlicht: (2026)
von: Cassieri, Pietro, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault Localization
von: Kang, Sungmin, et al.
Veröffentlicht: (2023) -
Identifying Bug Inducing Commits by Combining Fault Localisation and Code Change Histories
von: An, Gabin, et al.
Veröffentlicht: (2025) -
Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap
von: Kim, Naryeong, et al.
Veröffentlicht: (2026) -
COSMosFL: Ensemble of Small Language Models for Fault Localisation
von: Cho, Hyunjoon, et al.
Veröffentlicht: (2025) -
Finding the Needle in the Crash Stack: Industrial-Scale Crash Root Cause Localization with AutoCrashFL
von: Kang, Sungmin, et al.
Veröffentlicht: (2025)