Simple and Effective Baselines for Code Summarisation Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Robinson, Jade, Kummerfeld, Jonathan K. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MEC$^3$O: Multi-Expert Consensus for Code Time Complexity Prediction
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
TCProF: Time-Complexity Prediction SSL Framework
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
MicroRemed: Benchmarking LLMs in Microservices Remediation
by: Zhang, Lingzhe, et al.
Published: (2025)
by: Zhang, Lingzhe, et al.
Published: (2025)
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)
by: Lee, Hokyung, et al.
Published: (2024)
Is It Time To Treat Prompts As Code? A Multi-Use Case Study For Prompt Optimization Using DSPy
by: Lemos, Francisca, et al.
Published: (2025)
by: Lemos, Francisca, et al.
Published: (2025)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
by: Xia, Bowei, et al.
Published: (2026)
by: Xia, Bowei, et al.
Published: (2026)
GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries
by: Fachada, Nuno, et al.
Published: (2025)
by: Fachada, Nuno, et al.
Published: (2025)
QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges
by: Basit, Abdul, et al.
Published: (2025)
by: Basit, Abdul, et al.
Published: (2025)
XPath Agent: An Efficient XPath Programming Agent Based on LLM for Web Crawler
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
GraphSkill: Documentation-Guided Hierarchical Retrieval-Augmented Coding for Complex Graph Reasoning
by: Wang, Fali, et al.
Published: (2026)
by: Wang, Fali, et al.
Published: (2026)
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
by: Lim, Soohan, et al.
Published: (2025)
by: Lim, Soohan, et al.
Published: (2025)
CELI: Controller-Embedded Language Model Interactions
by: Wagner, Jan-Samuel, et al.
Published: (2024)
by: Wagner, Jan-Samuel, et al.
Published: (2024)
When LLM meets Fuzzy-TOPSIS for Personnel Selection through Automated Profile Analysis
by: Hoque, Shahria, et al.
Published: (2026)
by: Hoque, Shahria, et al.
Published: (2026)
CATER: Leveraging LLM to Pioneer a Multidimensional, Reference-Independent Paradigm in Translation Quality Evaluation
by: IIDA, Kurando, et al.
Published: (2024)
by: IIDA, Kurando, et al.
Published: (2024)
Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
by: Görge, Rebekka, et al.
Published: (2025)
by: Görge, Rebekka, et al.
Published: (2025)
GEML: A Grammar-based Evolutionary Machine Learning Approach for Design-Pattern Detection
by: Barbudo, Rafael, et al.
Published: (2024)
by: Barbudo, Rafael, et al.
Published: (2024)
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
by: Fang, Xi, et al.
Published: (2025)
by: Fang, Xi, et al.
Published: (2025)
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
by: Tran, Hung, et al.
Published: (2026)
by: Tran, Hung, et al.
Published: (2026)
ExpliCa: Evaluating Explicit Causal Reasoning in Large Language Models
by: Miliani, Martina, et al.
Published: (2025)
by: Miliani, Martina, et al.
Published: (2025)
Enabling Low-Resource Language Retrieval: Establishing Baselines for Urdu MS MARCO
by: Butt, Umer, et al.
Published: (2024)
by: Butt, Umer, et al.
Published: (2024)
Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context
by: An, Tao
Published: (2025)
by: An, Tao
Published: (2025)
One Agent to Serve All: a Lite-Adaptive Stylized AI Assistant for Millions of Multi-Style Official Accounts
by: Fan, Xingyu, et al.
Published: (2025)
by: Fan, Xingyu, et al.
Published: (2025)
Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
by: Xu, Beining, et al.
Published: (2025)
by: Xu, Beining, et al.
Published: (2025)
A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents
by: Soman, Sumit, et al.
Published: (2025)
by: Soman, Sumit, et al.
Published: (2025)
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
by: Alam, Firoj, et al.
Published: (2025)
by: Alam, Firoj, et al.
Published: (2025)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
by: Schmidt, Craig W., et al.
Published: (2025)
by: Schmidt, Craig W., et al.
Published: (2025)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
by: Lee, Yejin, et al.
Published: (2025)
by: Lee, Yejin, et al.
Published: (2025)
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
by: Lee, Yejin, et al.
Published: (2025)
by: Lee, Yejin, et al.
Published: (2025)
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
by: Tu, Songjun, et al.
Published: (2025)
by: Tu, Songjun, et al.
Published: (2025)
Efficient Adaptive Rejection Sampling for Accelerating Speculative Decoding in Large Language Models
by: Sun, Chendong, et al.
Published: (2025)
by: Sun, Chendong, et al.
Published: (2025)
Reducing Hallucinations in Summarization via Reinforcement Learning with Entity Hallucination Index
by: Katwe, Praveenkumar, et al.
Published: (2025)
by: Katwe, Praveenkumar, et al.
Published: (2025)
Bielik 11B v2 Technical Report
by: Ociepa, Krzysztof, et al.
Published: (2025)
by: Ociepa, Krzysztof, et al.
Published: (2025)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
by: Marcuzzo, Matteo, et al.
Published: (2025)
by: Marcuzzo, Matteo, et al.
Published: (2025)
Pun Unintended: LLMs and the Illusion of Humor Understanding
by: Zangari, Alessandro, et al.
Published: (2025)
by: Zangari, Alessandro, et al.
Published: (2025)
Multi-chain Graph Refinement and Selection for Reliable Reasoning in Large Language Models
by: Yang, Yujiao, et al.
Published: (2025)
by: Yang, Yujiao, et al.
Published: (2025)
The Impact of Role Design in In-Context Learning for Large Language Models
by: Rouzegar, Hamidreza, et al.
Published: (2025)
by: Rouzegar, Hamidreza, et al.
Published: (2025)
Semantic Synergy: Unlocking Policy Insights and Learning Pathways Through Advanced Skill Mapping
by: Koundouri, Phoebe, et al.
Published: (2025)
by: Koundouri, Phoebe, et al.
Published: (2025)
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
by: Marjanović, Sara Vera, et al.
Published: (2024)
by: Marjanović, Sara Vera, et al.
Published: (2024)
TREX: Tokenizer Regression for Optimal Data Mixture
by: Won, Inho, et al.
Published: (2026)
by: Won, Inho, et al.
Published: (2026)
Similar Items
-
MEC$^3$O: Multi-Expert Consensus for Code Time Complexity Prediction
by: Hahn, Joonghyuk, et al.
Published: (2025) -
TCProF: Time-Complexity Prediction SSL Framework
by: Hahn, Joonghyuk, et al.
Published: (2025) -
MicroRemed: Benchmarking LLMs in Microservices Remediation
by: Zhang, Lingzhe, et al.
Published: (2025) -
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025) -
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)