Beyond BLEU: A Semantic Evaluation Method for Code Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Näumann, Julius, Keidel, Sven, Sharifloo, Amir Molzam, Mezini, Mira |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Graph-Language Fusion for Structure-Aware Code Generation
by: Tiftikci, Mert, et al.
Published: (2026)
by: Tiftikci, Mert, et al.
Published: (2026)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
by: Sharifloo, Amir Molzam, et al.
Published: (2025)
by: Sharifloo, Amir Molzam, et al.
Published: (2025)
A Direct-Style Effect Notation for Sequential and Parallel Programs
by: Richter, David, et al.
Published: (2023)
by: Richter, David, et al.
Published: (2023)
DeCo: A Core Calculus for Incremental Functional Programming with Generic Data Types
by: Böhler, Timon, et al.
Published: (2026)
by: Böhler, Timon, et al.
Published: (2026)
PRDTs: Composable Knowledge-Based Consensus Protocols with Replicated Data Types
by: Haas, Julian, et al.
Published: (2025)
by: Haas, Julian, et al.
Published: (2025)
Distributed Locking as a Data Type
by: Haas, Julian, et al.
Published: (2024)
by: Haas, Julian, et al.
Published: (2024)
Abstracting Denotational Interpreters
by: Graf, Sebastian, et al.
Published: (2024)
by: Graf, Sebastian, et al.
Published: (2024)
Compiling with Arrays
by: Richter, David, et al.
Published: (2024)
by: Richter, David, et al.
Published: (2024)
LoRe: A Programming Model for Verifiably Safe Local-First Software
by: Haas, Julian, et al.
Published: (2023)
by: Haas, Julian, et al.
Published: (2023)
Evaluating Extremely Low-Resource Machine Translation: A Comparative Study of ChrF++ and BLEU Metrics
by: Kumar, Sanjeev, et al.
Published: (2026)
by: Kumar, Sanjeev, et al.
Published: (2026)
Evaluating and Mitigating Errors in LLM-Generated Web API Integrations
by: Maninger, Daniel, et al.
Published: (2025)
by: Maninger, Daniel, et al.
Published: (2025)
TensorBLEU: Vectorized GPU-based BLEU Score Implementation for Per-Sentence In-Training Evaluation
by: Filipek, Adam
Published: (2025)
by: Filipek, Adam
Published: (2025)
A Critical Study of What Code-LLMs (Do Not) Learn
by: Anand, Abhinav, et al.
Published: (2024)
by: Anand, Abhinav, et al.
Published: (2024)
SignBLEU: Automatic Evaluation of Multi-channel Sign Language Translation
by: Kim, Jung-Ho, et al.
Published: (2024)
by: Kim, Jung-Ho, et al.
Published: (2024)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
by: Noah, Amit Finkman, et al.
Published: (2024)
by: Noah, Amit Finkman, et al.
Published: (2024)
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
by: Bae, Suyoung, et al.
Published: (2026)
by: Bae, Suyoung, et al.
Published: (2026)
How Programming Concepts and Neurons Are Shared in Code Language Models
by: Kargaran, Amir Hossein, et al.
Published: (2025)
by: Kargaran, Amir Hossein, et al.
Published: (2025)
Semantic Source Code Segmentation using Small and Large Language Models
by: Dahou, Abdelhalim, et al.
Published: (2025)
by: Dahou, Abdelhalim, et al.
Published: (2025)
A Multi-Perspective Architecture for Semantic Code Search
by: Haldar, Rajarshi, et al.
Published: (2020)
by: Haldar, Rajarshi, et al.
Published: (2020)
SymCode: A Neurosymbolic Approach to Mathematical Reasoning via Verifiable Code Generation
by: Nezhad, Sina Bagheri, et al.
Published: (2025)
by: Nezhad, Sina Bagheri, et al.
Published: (2025)
Efficient Algorithms for Partial Constraint Satisfaction Problems over Control-flow Graphs
by: Cai, Xuran, et al.
Published: (2026)
by: Cai, Xuran, et al.
Published: (2026)
A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond
by: Sun, Qiushi, et al.
Published: (2024)
by: Sun, Qiushi, et al.
Published: (2024)
Evaluating Program Semantics Reasoning with Type Inference in System F
by: He, Yifeng, et al.
Published: (2025)
by: He, Yifeng, et al.
Published: (2025)
On Code-Induced Reasoning in LLMs
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
Revisiting Code Similarity Evaluation with Abstract Syntax Tree Edit Distance
by: Song, Yewei, et al.
Published: (2024)
by: Song, Yewei, et al.
Published: (2024)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
by: Chang, Yapei, et al.
Published: (2025)
by: Chang, Yapei, et al.
Published: (2025)
Constrained Code Generation with Discrete Diffusion
by: Shao, Lize, et al.
Published: (2026)
by: Shao, Lize, et al.
Published: (2026)
CodeMind: Evaluating Large Language Models for Code Reasoning
by: Liu, Changshu, et al.
Published: (2024)
by: Liu, Changshu, et al.
Published: (2024)
Hear Your Code Fail, Voice-Assisted Debugging for Python
by: Amiri, Sayed Mahbub Hasan, et al.
Published: (2025)
by: Amiri, Sayed Mahbub Hasan, et al.
Published: (2025)
\texttt{ReMind}: Understanding Deductive Code Reasoning in LLMs
by: Gao, Jun, et al.
Published: (2025)
by: Gao, Jun, et al.
Published: (2025)
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
by: Lin, Honglin, et al.
Published: (2025)
by: Lin, Honglin, et al.
Published: (2025)
Compiling by Proving: Language-Agnostic Automatic Optimization from Formal Semantics
by: Zhao, Jianhong, et al.
Published: (2025)
by: Zhao, Jianhong, et al.
Published: (2025)
LLM4Decompile: Decompiling Binary Code with Large Language Models
by: Tan, Hanzhuo, et al.
Published: (2024)
by: Tan, Hanzhuo, et al.
Published: (2024)
OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
by: Huang, Siming, et al.
Published: (2024)
by: Huang, Siming, et al.
Published: (2024)
VeriAgent: A Tool-Integrated Multi-Agent System with Evolving Memory for PPA-Aware RTL Code Generation
by: Wang, Yaoxiang, et al.
Published: (2026)
by: Wang, Yaoxiang, et al.
Published: (2026)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
SaraCoder: Orchestrating Semantic and Structural Cues for Resource-Optimized Repository-Level Code Completion
by: Chen, Xiaohan, et al.
Published: (2025)
by: Chen, Xiaohan, et al.
Published: (2025)
Uncertainty Quantification for Evaluating Machine Translation Bias
by: Staliūnaitė, Ieva Raminta, et al.
Published: (2025)
by: Staliūnaitė, Ieva Raminta, et al.
Published: (2025)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
by: Jing, Huihao, et al.
Published: (2026)
by: Jing, Huihao, et al.
Published: (2026)
Self-Infilling Code Generation
by: Zheng, Lin, et al.
Published: (2023)
by: Zheng, Lin, et al.
Published: (2023)
Similar Items
-
Deep Graph-Language Fusion for Structure-Aware Code Generation
by: Tiftikci, Mert, et al.
Published: (2026) -
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
by: Sharifloo, Amir Molzam, et al.
Published: (2025) -
A Direct-Style Effect Notation for Sequential and Parallel Programs
by: Richter, David, et al.
Published: (2023) -
DeCo: A Core Calculus for Incremental Functional Programming with Generic Data Types
by: Böhler, Timon, et al.
Published: (2026) -
PRDTs: Composable Knowledge-Based Consensus Protocols with Replicated Data Types
by: Haas, Julian, et al.
Published: (2025)