Does In-IDE Calibration of Large Language Models work at Scale?
Fuente:
arXiv
Saved in:
| Main Authors: | Koohestani, Roham, Sergeyuk, Agnia, Gros, David, Spiess, Claudio, Titov, Sergey, Devanbu, Prem, Izadi, Maliheh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
In-IDE Human-AI Experience in the Era of Large Language Models; A Literature Review
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Rethinking IDE Customization for Enhanced HAX: A Hyperdimensional Perspective
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Prompt-with-Me: in-IDE Structured Prompt Management for LLM-Driven Software Engineering
by: Li, Ziyou, et al.
Published: (2025)
by: Li, Ziyou, et al.
Published: (2025)
Leveraging Large Language Models for Enhancing the Understandability of Generated Unit Tests
by: Deljouyi, Amirhossein, et al.
Published: (2024)
by: Deljouyi, Amirhossein, et al.
Published: (2024)
Localized Calibrated Uncertainty in Code Language Models
by: Gros, David, et al.
Published: (2025)
by: Gros, David, et al.
Published: (2025)
HyperSeq: A Hyper-Adaptive Representation for Predictive Sequencing of States
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Human-AI Experience in Integrated Development Environments: A Systematic Literature Review
by: Sergeyuk, Agnia, et al.
Published: (2025)
by: Sergeyuk, Agnia, et al.
Published: (2025)
AST-PAC: AST-guided Membership Inference for Code
by: Koohestani, Roham, et al.
Published: (2026)
by: Koohestani, Roham, et al.
Published: (2026)
TriCEGAR: A Trace-Driven Abstraction Mechanism for Agentic AI
by: Koohestani, Roham, et al.
Published: (2026)
by: Koohestani, Roham, et al.
Published: (2026)
Developer Interaction Patterns with Proactive AI: A Five-Day Field Study
by: Kuo, Nadine, et al.
Published: (2026)
by: Kuo, Nadine, et al.
Published: (2026)
AgentGuard: Runtime Verification of AI Agents
by: Koohestani, Roham
Published: (2025)
by: Koohestani, Roham
Published: (2025)
Code4MeV2: a Research-oriented Code-completion Platform
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time
by: Popescu, Razvan Mihai, et al.
Published: (2026)
by: Popescu, Razvan Mihai, et al.
Published: (2026)
A Multi-agent Onboarding Assistant based on Large Language Models, Retrieval Augmented Generation, and Chain-of-Thought
by: Ionescu, Andrei Cristian, et al.
Published: (2025)
by: Ionescu, Andrei Cristian, et al.
Published: (2025)
How Robustly do LLMs Understand Execution Semantics?
by: Spiess, Claudio, et al.
Published: (2026)
by: Spiess, Claudio, et al.
Published: (2026)
Developer Needs and Feasible Features for AI Assistants in IDEs
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
On LLMs' Internal Representation of Code Correctness
by: Ribeiro, Francisco, et al.
Published: (2025)
by: Ribeiro, Francisco, et al.
Published: (2025)
Calibration and Correctness of Language Models for Code
by: Spiess, Claudio, et al.
Published: (2024)
by: Spiess, Claudio, et al.
Published: (2024)
Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
by: Mundhra, Yash, et al.
Published: (2025)
by: Mundhra, Yash, et al.
Published: (2025)
Traces of Memorisation in Large Language Models for Code
by: Al-Kaswan, Ali, et al.
Published: (2023)
by: Al-Kaswan, Ali, et al.
Published: (2023)
Using AI-Based Coding Assistants in Practice: State of Affairs, Perceptions, and Ways Forward
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Assessing Consensus of Developers' Views on Code Readability
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Reassessing Java Code Readability Models with a Human-Centered Approach
by: Sergeyuk, Agnia, et al.
Published: (2024)
by: Sergeyuk, Agnia, et al.
Published: (2024)
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
by: Al-Kaswan, Ali, et al.
Published: (2025)
by: Al-Kaswan, Ali, et al.
Published: (2025)
AI in Software Engineering: Perceived Roles and Their Impact on Adoption
by: Zakharov, Ilya, et al.
Published: (2025)
by: Zakharov, Ilya, et al.
Published: (2025)
From Teacher to Colleague: How Coding Experience Shapes Developer Perceptions of AI Tools
by: Zakharov, Ilya, et al.
Published: (2025)
by: Zakharov, Ilya, et al.
Published: (2025)
Calibration of Large Language Models on Code Summarization
by: Virk, Yuvraj, et al.
Published: (2024)
by: Virk, Yuvraj, et al.
Published: (2024)
Understanding Prompt Programming Tasks and Questions
by: Liang, Jenny T., et al.
Published: (2025)
by: Liang, Jenny T., et al.
Published: (2025)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
by: Cipollone, Daniele, et al.
Published: (2025)
by: Cipollone, Daniele, et al.
Published: (2025)
Ecosystem of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
Beyond Functional Correctness: Design Issues in AI IDE-Generated Large-Scale Projects
by: Kashif, Syed Mohammad, et al.
Published: (2026)
by: Kashif, Syed Mohammad, et al.
Published: (2026)
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
by: Katzy, Jonathan, et al.
Published: (2024)
by: Katzy, Jonathan, et al.
Published: (2024)
Are Agents Probabilistic Automata? A Trace-Based, Memory-Constrained Theory of Agentic AI
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
by: Bouzenia, Islem, et al.
Published: (2024)
by: Bouzenia, Islem, et al.
Published: (2024)
Themisto: Jupyter-Based Runtime Benchmark
by: Grotov, Konstantin, et al.
Published: (2025)
by: Grotov, Konstantin, et al.
Published: (2025)
A Transformer-Based Approach for Smart Invocation of Automatic Code Completion
by: de Moor, Aral, et al.
Published: (2024)
by: de Moor, Aral, et al.
Published: (2024)
Robustness, Security, Privacy, Explainability, Efficiency, and Usability of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
In-IDE Toolkit for Developers of AI-Based Features
by: Sokolov, Yaroslav, et al.
Published: (2026)
by: Sokolov, Yaroslav, et al.
Published: (2026)
Similar Items
-
In-IDE Human-AI Experience in the Era of Large Language Models; A Literature Review
by: Sergeyuk, Agnia, et al.
Published: (2024) -
Rethinking IDE Customization for Enhanced HAX: A Hyperdimensional Perspective
by: Koohestani, Roham, et al.
Published: (2025) -
Prompt-with-Me: in-IDE Structured Prompt Management for LLM-Driven Software Engineering
by: Li, Ziyou, et al.
Published: (2025) -
Leveraging Large Language Models for Enhancing the Understandability of Generated Unit Tests
by: Deljouyi, Amirhossein, et al.
Published: (2024) -
Localized Calibrated Uncertainty in Code Language Models
by: Gros, David, et al.
Published: (2025)