Ensuring Reliability in Programming Knowledge Tracing: A Re-evaluation of Attention-augmented Models and Experimental Protocols
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kim, Jaewook, Kim, Hyeoncheol |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Ensuring Functional Correctness of Large Code Models with Selective Generation
par: Jeong, Jaewoo, et autres
Publié: (2025)
par: Jeong, Jaewoo, et autres
Publié: (2025)
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
par: Jin, Yiyang, et autres
Publié: (2025)
par: Jin, Yiyang, et autres
Publié: (2025)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
par: Haque, Mirazul, et autres
Publié: (2025)
par: Haque, Mirazul, et autres
Publié: (2025)
Context-Augmented Code Generation Using Programming Knowledge Graphs
par: Seddik, Shahd, et autres
Publié: (2026)
par: Seddik, Shahd, et autres
Publié: (2026)
High-quality data augmentation for code comment classification
par: Borsani, Thomas, et autres
Publié: (2026)
par: Borsani, Thomas, et autres
Publié: (2026)
Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap
par: Kim, Naryeong, et autres
Publié: (2026)
par: Kim, Naryeong, et autres
Publié: (2026)
From Prediction to Application: Language Model-based Code Knowledge Tracing with Domain Adaptive Pre-Training and Automatic Feedback System with Pedagogical Prompting for Comprehensive Programming Education
par: Lee, Unggi, et autres
Publié: (2024)
par: Lee, Unggi, et autres
Publié: (2024)
Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
par: Zhou, Yongxi, et autres
Publié: (2026)
par: Zhou, Yongxi, et autres
Publié: (2026)
Tracing Stereotypes in Pre-trained Transformers: From Biased Neurons to Fairer Models
par: Voria, Gianmario, et autres
Publié: (2026)
par: Voria, Gianmario, et autres
Publié: (2026)
Teaching LLMs Program Semantics via Symbolic Execution Traces
par: Bayer, Jonas, et autres
Publié: (2026)
par: Bayer, Jonas, et autres
Publié: (2026)
Capturing Semantic Flow of ML-based Systems
par: Yoo, Shin, et autres
Publié: (2025)
par: Yoo, Shin, et autres
Publié: (2025)
RePlay: a Recommendation Framework for Experimentation and Production Use
par: Vasilev, Alexey, et autres
Publié: (2024)
par: Vasilev, Alexey, et autres
Publié: (2024)
Fault Localization via Fine-tuning Large Language Models with Mutation Generated Stack Traces
par: Jambigi, Neetha, et autres
Publié: (2025)
par: Jambigi, Neetha, et autres
Publié: (2025)
DANDI: Diffusion as Normative Distribution for Deep Neural Network Input
par: Kim, Somin, et autres
Publié: (2025)
par: Kim, Somin, et autres
Publié: (2025)
It's LIT! Reliability-Optimized LLMs with Inspectable Tools
par: Zhang, Ruixin, et autres
Publié: (2025)
par: Zhang, Ruixin, et autres
Publié: (2025)
RepairBench: Leaderboard of Frontier Models for Program Repair
par: Silva, André, et autres
Publié: (2024)
par: Silva, André, et autres
Publié: (2024)
ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization
par: Lian, Junbo Jacob, et autres
Publié: (2026)
par: Lian, Junbo Jacob, et autres
Publié: (2026)
Stack Trace-Based Crash Deduplication with Transformer Adaptation
par: Mamun, Md Afif Al, et autres
Publié: (2025)
par: Mamun, Md Afif Al, et autres
Publié: (2025)
On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code
par: Weyssow, Martin, et autres
Publié: (2023)
par: Weyssow, Martin, et autres
Publié: (2023)
Can Hessian-Based Insights Support Fault Diagnosis in Attention-based Models?
par: Jahan, Sigma, et autres
Publié: (2025)
par: Jahan, Sigma, et autres
Publié: (2025)
Bayesian Program Learning by Decompiling Amortized Knowledge
par: Palmarini, Alessandro B., et autres
Publié: (2023)
par: Palmarini, Alessandro B., et autres
Publié: (2023)
Identifying Flaky Tests in Quantum Code: A Machine Learning Approach
par: Kaur, Khushdeep, et autres
Publié: (2025)
par: Kaur, Khushdeep, et autres
Publié: (2025)
A Machine Learning-Based Error Mitigation Approach For Reliable Software Development On IBM'S Quantum Computers
par: Muqeet, Asmar, et autres
Publié: (2024)
par: Muqeet, Asmar, et autres
Publié: (2024)
A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
par: Awal, Md. Abdul, et autres
Publié: (2025)
par: Awal, Md. Abdul, et autres
Publié: (2025)
MLXP: A Framework for Conducting Replicable Experiments in Python
par: Arbel, Michael, et autres
Publié: (2024)
par: Arbel, Michael, et autres
Publié: (2024)
Program Structure-aware Language Models: Targeted Software Testing beyond Textual Semantics
par: Tran, Khang, et autres
Publié: (2026)
par: Tran, Khang, et autres
Publié: (2026)
Knowledge Tracing in Programming Education Integrating Students' Questions
par: Kim, Doyoun, et autres
Publié: (2025)
par: Kim, Doyoun, et autres
Publié: (2025)
FairQuant: Certifying and Quantifying Fairness of Deep Neural Networks
par: Kim, Brian Hyeongseok, et autres
Publié: (2024)
par: Kim, Brian Hyeongseok, et autres
Publié: (2024)
ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning
par: Jiang, Juyong, et autres
Publié: (2026)
par: Jiang, Juyong, et autres
Publié: (2026)
Disproving Program Equivalence with LLMs
par: Allamanis, Miltiadis, et autres
Publié: (2025)
par: Allamanis, Miltiadis, et autres
Publié: (2025)
ReCodeAgent: A Multi-Agent Workflow for Language-agnostic Translation and Validation of Large-scale Repositories
par: Ibrahimzada, Ali Reza, et autres
Publié: (2026)
par: Ibrahimzada, Ali Reza, et autres
Publié: (2026)
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
par: LeVine, Will, et autres
Publié: (2026)
par: LeVine, Will, et autres
Publié: (2026)
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
par: Yoon, Juyeon, et autres
Publié: (2025)
par: Yoon, Juyeon, et autres
Publié: (2025)
Where's the Bug? Attention Probing for Scalable Fault Localization
par: Stein, Adam, et autres
Publié: (2025)
par: Stein, Adam, et autres
Publié: (2025)
Bootstrapping Coding Agents: The Specification Is the Program
par: Monperrus, Martin
Publié: (2026)
par: Monperrus, Martin
Publié: (2026)
A Progressive Transformer for Unifying Binary Code Embedding and Knowledge Transfer
par: Lu, Hanxiao, et autres
Publié: (2024)
par: Lu, Hanxiao, et autres
Publié: (2024)
JetTrain: IDE-Native Machine Learning Experiments
par: Trofimov, Artem, et autres
Publié: (2024)
par: Trofimov, Artem, et autres
Publié: (2024)
CigaR: Cost-efficient Program Repair with LLMs
par: Hidvégi, Dávid, et autres
Publié: (2024)
par: Hidvégi, Dávid, et autres
Publié: (2024)
Software Engineering Principles for Fairer Systems: Experiments with GroupCART
par: Peng, Kewen, et autres
Publié: (2025)
par: Peng, Kewen, et autres
Publié: (2025)
A Survey on LLM-based Code Generation for Low-Resource and Domain-Specific Programming Languages
par: Joel, Sathvik, et autres
Publié: (2024)
par: Joel, Sathvik, et autres
Publié: (2024)
Documents similaires
-
Ensuring Functional Correctness of Large Code Models with Selective Generation
par: Jeong, Jaewoo, et autres
Publié: (2025) -
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
par: Jin, Yiyang, et autres
Publié: (2025) -
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
par: Haque, Mirazul, et autres
Publié: (2025) -
Context-Augmented Code Generation Using Programming Knowledge Graphs
par: Seddik, Shahd, et autres
Publié: (2026) -
High-quality data augmentation for code comment classification
par: Borsani, Thomas, et autres
Publié: (2026)