Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhuohao, Chen, Wenqing, Yu, Jianxing, Lu, Zhichao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
Generating Equivalent Representations of Code By A Self-Reflection Approach
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
von: Peng, Qiwei, et al.
Veröffentlicht: (2024)
von: Peng, Qiwei, et al.
Veröffentlicht: (2024)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
von: Zhang, William, et al.
Veröffentlicht: (2024)
von: Zhang, William, et al.
Veröffentlicht: (2024)
SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
von: Petrukha, Ivan, et al.
Veröffentlicht: (2025)
von: Petrukha, Ivan, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
SEVerA: Verified Synthesis of Self-Evolving Agents
von: Banerjee, Debangshu, et al.
Veröffentlicht: (2026)
von: Banerjee, Debangshu, et al.
Veröffentlicht: (2026)
EigenData: A Self-Evolving Multi-Agent Platform for Function-Calling Data Synthesis, Auditing, and Repair
von: Chen, Jiaao, et al.
Veröffentlicht: (2026)
von: Chen, Jiaao, et al.
Veröffentlicht: (2026)
PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models
von: Chen, Simin, et al.
Veröffentlicht: (2024)
von: Chen, Simin, et al.
Veröffentlicht: (2024)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
von: Peng, Yun, et al.
Veröffentlicht: (2024)
von: Peng, Yun, et al.
Veröffentlicht: (2024)
Neural Models for Source Code Synthesis and Completion
von: Niyogi, Mitodru
Veröffentlicht: (2024)
von: Niyogi, Mitodru
Veröffentlicht: (2024)
CodeFuse-Query: A Data-Centric Static Code Analysis System for Large-Scale Organizations
von: Xie, Xiaoheng, et al.
Veröffentlicht: (2024)
von: Xie, Xiaoheng, et al.
Veröffentlicht: (2024)
Python Symbolic Execution with LLM-powered Code Generation
von: Wang, Wenhan, et al.
Veröffentlicht: (2024)
von: Wang, Wenhan, et al.
Veröffentlicht: (2024)
MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings
von: Gurioli, Andrea, et al.
Veröffentlicht: (2025)
von: Gurioli, Andrea, et al.
Veröffentlicht: (2025)
Is Self-Repair a Silver Bullet for Code Generation?
von: Olausson, Theo X., et al.
Veröffentlicht: (2023)
von: Olausson, Theo X., et al.
Veröffentlicht: (2023)
How Programming Concepts and Neurons Are Shared in Code Language Models
von: Kargaran, Amir Hossein, et al.
Veröffentlicht: (2025)
von: Kargaran, Amir Hossein, et al.
Veröffentlicht: (2025)
Semantic Source Code Segmentation using Small and Large Language Models
von: Dahou, Abdelhalim, et al.
Veröffentlicht: (2025)
von: Dahou, Abdelhalim, et al.
Veröffentlicht: (2025)
Revisiting Code Similarity Evaluation with Abstract Syntax Tree Edit Distance
von: Song, Yewei, et al.
Veröffentlicht: (2024)
von: Song, Yewei, et al.
Veröffentlicht: (2024)
SWE-QA: Can Language Models Answer Repository-level Code Questions?
von: Peng, Weihan, et al.
Veröffentlicht: (2025)
von: Peng, Weihan, et al.
Veröffentlicht: (2025)
CodeMind: Evaluating Large Language Models for Code Reasoning
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
Contract Based Verification of Non-functional Requirements for Embedded Automotive C Code
von: Amilon, Jesper, et al.
Veröffentlicht: (2026)
von: Amilon, Jesper, et al.
Veröffentlicht: (2026)
Hidden in Plain Sight: Where Developers Confess Self-Admitted Technical Debt
von: Sridharan, Murali, et al.
Veröffentlicht: (2025)
von: Sridharan, Murali, et al.
Veröffentlicht: (2025)
Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation
von: Chen, Le, et al.
Veröffentlicht: (2025)
von: Chen, Le, et al.
Veröffentlicht: (2025)
SaraCoder: Orchestrating Semantic and Structural Cues for Resource-Optimized Repository-Level Code Completion
von: Chen, Xiaohan, et al.
Veröffentlicht: (2025)
von: Chen, Xiaohan, et al.
Veröffentlicht: (2025)
LLM-Guided Compositional Program Synthesis
von: Khan, Ruhma, et al.
Veröffentlicht: (2025)
von: Khan, Ruhma, et al.
Veröffentlicht: (2025)
The CodeInverter Suite: Control-Flow and Data-Mapping Augmented Binary Decompilation with LLMs
von: Liu, Peipei, et al.
Veröffentlicht: (2025)
von: Liu, Peipei, et al.
Veröffentlicht: (2025)
A Unified Framework for Automated Code Transformation and Pragma Insertion
von: Pouget, Stéphane, et al.
Veröffentlicht: (2024)
von: Pouget, Stéphane, et al.
Veröffentlicht: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Building A Proof-Oriented Programmer That Is 64% Better Than GPT-4o Under Data Scarcity
von: Zhang, Dylan, et al.
Veröffentlicht: (2025)
von: Zhang, Dylan, et al.
Veröffentlicht: (2025)
A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
Code Broker: A Multi-Agent System for Automated Code Quality Assessment
von: Attrah, Samer
Veröffentlicht: (2026)
von: Attrah, Samer
Veröffentlicht: (2026)
Defusing Logic Bombs in Symbolic Execution with LLM-Generated Ghost Code
von: Bouras, Dimitrios Stamatios, et al.
Veröffentlicht: (2026)
von: Bouras, Dimitrios Stamatios, et al.
Veröffentlicht: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
COBOLAssist: Analyzing and Fixing Compilation Errors for LLM-Powered COBOL Code Generation
von: Dau, Anh T. V., et al.
Veröffentlicht: (2026)
von: Dau, Anh T. V., et al.
Veröffentlicht: (2026)
TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar
von: Li, Yinxi, et al.
Veröffentlicht: (2025)
von: Li, Yinxi, et al.
Veröffentlicht: (2025)
AutoCode: LLMs as Problem Setters for Competitive Programming
von: Zhou, Shang, et al.
Veröffentlicht: (2025)
von: Zhou, Shang, et al.
Veröffentlicht: (2025)
ReSyn: A Generalized Recursive Regular Expression Synthesis Framework
von: Kim, Seongmin, et al.
Veröffentlicht: (2026)
von: Kim, Seongmin, et al.
Veröffentlicht: (2026)
Novel Preprocessing Technique for Data Embedding in Engineering Code Generation Using Large Language Model
von: Lin, Yu-Chen, et al.
Veröffentlicht: (2023)
von: Lin, Yu-Chen, et al.
Veröffentlicht: (2023)
Enhancing Large Language Models in Coding Through Multi-Perspective Self-Consistency
von: Huang, Baizhou, et al.
Veröffentlicht: (2023)
von: Huang, Baizhou, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024) -
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026) -
Generating Equivalent Representations of Code By A Self-Reflection Approach
von: Li, Jia, et al.
Veröffentlicht: (2024) -
HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
von: Peng, Qiwei, et al.
Veröffentlicht: (2024) -
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
von: Zhang, William, et al.
Veröffentlicht: (2024)