Not All Tokens Matter: Data-Centric Optimization for Efficient Code Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Afrin, Saima, Cheng, Zaiyu, Sharma, Tushar, Serebrenik, Alexander, Di Penta, Massimiliano, Mastropaolo, Antonio |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation
by: Giagnorio, Alessandro, et al.
Published: (2025)
by: Giagnorio, Alessandro, et al.
Published: (2025)
Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough?
by: Vitale, Antonio, et al.
Published: (2025)
by: Vitale, Antonio, et al.
Published: (2025)
Resource-Efficient & Effective Code Summarization
by: Afrin, Saima, et al.
Published: (2025)
by: Afrin, Saima, et al.
Published: (2025)
Prompt-Driven Code Summarization: A Systematic Literature Review
by: Farjana, Afia, et al.
Published: (2026)
by: Farjana, Afia, et al.
Published: (2026)
An Empirical Study on the Effects of System Prompts in Instruction-Tuned Models for Code Generation
by: Cheng, Zaiyu, et al.
Published: (2026)
by: Cheng, Zaiyu, et al.
Published: (2026)
Parameter-Efficient Multi-Task Fine-Tuning in Code-Related Tasks
by: Haque, Md Zahidul, et al.
Published: (2026)
by: Haque, Md Zahidul, et al.
Published: (2026)
A Systematic Literature Review of Parameter-Efficient Fine-Tuning for Large Code Models
by: Afrin, Saima, et al.
Published: (2025)
by: Afrin, Saima, et al.
Published: (2025)
Is Quantization a Deal-breaker? Empirical Insights from Large Code Models
by: Afrin, Saima, et al.
Published: (2025)
by: Afrin, Saima, et al.
Published: (2025)
How the Training Procedure Impacts the Performance of Deep Learning-based Vulnerability Patching
by: Mastropaolo, Antonio, et al.
Published: (2024)
by: Mastropaolo, Antonio, et al.
Published: (2024)
A Taxonomy of Self-Admitted Technical Debt in Deep Learning Systems
by: Pepe, Federica, et al.
Published: (2024)
by: Pepe, Federica, et al.
Published: (2024)
Automated Refactoring of Non-Idiomatic Python Code: A Differentiated Replication with LLMs
by: Midolo, Alessandro, et al.
Published: (2025)
by: Midolo, Alessandro, et al.
Published: (2025)
Unveiling ChatGPT's Usage in Open Source Projects: A Mining-based Study
by: Tufano, Rosalia, et al.
Published: (2024)
by: Tufano, Rosalia, et al.
Published: (2024)
Developers and Generative AI: A Study of Self-Admitted Usage in Open Source Projects
by: Tufano, Rosalia, et al.
Published: (2026)
by: Tufano, Rosalia, et al.
Published: (2026)
Smart but Costly? Benchmarking LLMs on Functional Accuracy and Energy Efficiency
by: Mehditabar, Mohammadjavad, et al.
Published: (2025)
by: Mehditabar, Mohammadjavad, et al.
Published: (2025)
On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization
by: Crupi, Giuseppe, et al.
Published: (2025)
by: Crupi, Giuseppe, et al.
Published: (2025)
Towards Summarizing Code Snippets Using Pre-Trained Transformers
by: Mastropaolo, Antonio, et al.
Published: (2024)
by: Mastropaolo, Antonio, et al.
Published: (2024)
From Human to Machine Refactoring: Assessing GPT-4's Impact on Python Class Quality and Readability
by: Midolo, Alessandro, et al.
Published: (2026)
by: Midolo, Alessandro, et al.
Published: (2026)
Do LLMs Provide Links to Code Similar to what they Generate? A Study with Gemini and Bing CoPilot
by: Bifolco, Daniele, et al.
Published: (2025)
by: Bifolco, Daniele, et al.
Published: (2025)
CodeGenLink: A Tool to Find the Likely Origin and License of Automatically Generated Code
by: Bifolco, Daniele, et al.
Published: (2025)
by: Bifolco, Daniele, et al.
Published: (2025)
On the Effect of Token Merging on Pre-trained Models for Code
by: Saad, Mootez, et al.
Published: (2025)
by: Saad, Mootez, et al.
Published: (2025)
Guidelines to Prompt Large Language Models for Code Generation: An Empirical Characterization
by: Midolo, Alessandro, et al.
Published: (2026)
by: Midolo, Alessandro, et al.
Published: (2026)
Ethics of Care for Software Engineering
by: Serebrenik, Alexander, et al.
Published: (2026)
by: Serebrenik, Alexander, et al.
Published: (2026)
Teaching Empirical Methods at Eindhoven University of Technology
by: Serebrenik, Alexander, et al.
Published: (2024)
by: Serebrenik, Alexander, et al.
Published: (2024)
Machine Learning in the Wild: Early Evidence of Non-Compliant ML-Automation in Open-Source Software
by: Arshid, Zohaib, et al.
Published: (2026)
by: Arshid, Zohaib, et al.
Published: (2026)
Augmenting Software Bills of Materials with Software Vulnerability Description: A Preliminary Study on GitHub
by: Fucci, Davide, et al.
Published: (2025)
by: Fucci, Davide, et al.
Published: (2025)
"Let it be Chaos in the Plumbing!" Usage and Efficacy of Chaos Engineering in DevOps Pipelines
by: Fossati, Stefano, et al.
Published: (2025)
by: Fossati, Stefano, et al.
Published: (2025)
A Path Less Traveled: Reimagining Software Engineering Automation via a Neurosymbolic Paradigm
by: Mastropaolo, Antonio, et al.
Published: (2025)
by: Mastropaolo, Antonio, et al.
Published: (2025)
[Replication Package] Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation
by: Giagnorio, Alessandro, et al.
Published: (2025)
by: Giagnorio, Alessandro, et al.
Published: (2025)
CodeGreen: Towards Improving Precision and Portability in Software Energy Measurement
by: Rajput, Saurabhsingh, et al.
Published: (2026)
by: Rajput, Saurabhsingh, et al.
Published: (2026)
Code Review Automation: Strengths and Weaknesses of the State of the Art
by: Tufano, Rosalia, et al.
Published: (2024)
by: Tufano, Rosalia, et al.
Published: (2024)
On Inter-dataset Code Duplication and Data Leakage in Large Language Models
by: López, José Antonio Hernández, et al.
Published: (2024)
by: López, José Antonio Hernández, et al.
Published: (2024)
How are MLOps Frameworks Used in Open Source Projects? An Empirical Characterization
by: Zampetti, Fiorella, et al.
Published: (2026)
by: Zampetti, Fiorella, et al.
Published: (2026)
Quality Gatekeepers: Investigating the Effects ofCode Review Bots on Pull Request Activities
by: Wessel, Mairieli, et al.
Published: (2021)
by: Wessel, Mairieli, et al.
Published: (2021)
Automatic Categorization of GitHub Actions with Transformers and Few-shot Learning
by: Nguyen, Phuong T., et al.
Published: (2024)
by: Nguyen, Phuong T., et al.
Published: (2024)
HackRep: A Large-Scale Dataset of GitHub Hackathon Projects
by: Halmans, Sjoerd, et al.
Published: (2026)
by: Halmans, Sjoerd, et al.
Published: (2026)
Fine-grained Multi-Document Extraction and Generation of Code Change Rationale
by: Sun, Mehedi, et al.
Published: (2026)
by: Sun, Mehedi, et al.
Published: (2026)
BOMs Away! Inside the Minds of Stakeholders: A Comprehensive Study of Bills of Materials for Software Systems
by: Stalnaker, Trevor, et al.
Published: (2023)
by: Stalnaker, Trevor, et al.
Published: (2023)
CONCORD: Towards a DSL for Configurable Graph Code Representation
by: Saad, Mootez, et al.
Published: (2024)
by: Saad, Mootez, et al.
Published: (2024)
Negativity in Self-Admitted Technical Debt: How Sentiment Influences Prioritization
by: Cassee, Nathan, et al.
Published: (2025)
by: Cassee, Nathan, et al.
Published: (2025)
Exploring the Effect of Multiple Natural Languages on Code Suggestion Using GitHub Copilot
by: Koyanagi, Kei, et al.
Published: (2024)
by: Koyanagi, Kei, et al.
Published: (2024)
Similar Items
-
Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation
by: Giagnorio, Alessandro, et al.
Published: (2025) -
Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough?
by: Vitale, Antonio, et al.
Published: (2025) -
Resource-Efficient & Effective Code Summarization
by: Afrin, Saima, et al.
Published: (2025) -
Prompt-Driven Code Summarization: A Systematic Literature Review
by: Farjana, Afia, et al.
Published: (2026) -
An Empirical Study on the Effects of System Prompts in Instruction-Tuned Models for Code Generation
by: Cheng, Zaiyu, et al.
Published: (2026)