MHRC-Bench: A Multilingual Hardware Repository-Level Code Completion benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zou, Qingyun, Cui, Jiahao, Chen, Nuo, He, Bingsheng, Wong, Weng-Fai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
JudgeLRM: Large Reasoning Models as a Judge
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
HLStrans: Dataset for C-to-HLS Hardware Code Synthesis
von: Zou, Qingyun, et al.
Veröffentlicht: (2025)
von: Zou, Qingyun, et al.
Veröffentlicht: (2025)
Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation
von: Chen, Nuo, et al.
Veröffentlicht: (2026)
von: Chen, Nuo, et al.
Veröffentlicht: (2026)
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
von: Dong, Honghua, et al.
Veröffentlicht: (2025)
von: Dong, Honghua, et al.
Veröffentlicht: (2025)
Towards Repository-Level Program Verification with Large Language Models
von: Zhong, Si Cheng, et al.
Veröffentlicht: (2025)
von: Zhong, Si Cheng, et al.
Veröffentlicht: (2025)
Beyond Brainstorming: What Drives High-Quality Scientific Ideas? Lessons from Multi-Agent Collaboration
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable Code
von: Zeng, Lingfei, et al.
Veröffentlicht: (2025)
von: Zeng, Lingfei, et al.
Veröffentlicht: (2025)
Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation
von: Chen, Le, et al.
Veröffentlicht: (2025)
von: Chen, Le, et al.
Veröffentlicht: (2025)
Analysis of AdvFusion: Adapter-based Multilingual Learning for Code Large Language Models
von: Esmaeili, Amirreza, et al.
Veröffentlicht: (2025)
von: Esmaeili, Amirreza, et al.
Veröffentlicht: (2025)
Bench4HLS: End-to-End Evaluation of LLMs in High-Level Synthesis Code Generation
von: Khan, M Zafir Sadik, et al.
Veröffentlicht: (2026)
von: Khan, M Zafir Sadik, et al.
Veröffentlicht: (2026)
CodeV: Empowering LLMs with HDL Generation through Multi-Level Summarization
von: Zhao, Yang, et al.
Veröffentlicht: (2024)
von: Zhao, Yang, et al.
Veröffentlicht: (2024)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks
von: Dandamudi, Rohit, et al.
Veröffentlicht: (2024)
von: Dandamudi, Rohit, et al.
Veröffentlicht: (2024)
Towards Formal Verification of LLM-Generated Code from Natural Language Prompts
von: Councilman, Aaron, et al.
Veröffentlicht: (2025)
von: Councilman, Aaron, et al.
Veröffentlicht: (2025)
Insights from the Usage of the Ansible Lightspeed Code Completion Service
von: Sahoo, Priyam, et al.
Veröffentlicht: (2024)
von: Sahoo, Priyam, et al.
Veröffentlicht: (2024)
OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
von: Ding, Deming, et al.
Veröffentlicht: (2026)
von: Ding, Deming, et al.
Veröffentlicht: (2026)
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
von: Ravi, Nikil, et al.
Veröffentlicht: (2026)
von: Ravi, Nikil, et al.
Veröffentlicht: (2026)
Fine-Tuning Multilingual Language Models for Code Review: An Empirical Study on Industrial C# Projects
von: Begolli, Igli, et al.
Veröffentlicht: (2025)
von: Begolli, Igli, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
How Do Humans Write Code? Large Models Do It the Same Way Too
von: Li, Long, et al.
Veröffentlicht: (2024)
von: Li, Long, et al.
Veröffentlicht: (2024)
Chain of Execution Supervision Promotes General Reasoning in Large Language Models
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules
von: Le, Hung, et al.
Veröffentlicht: (2023)
von: Le, Hung, et al.
Veröffentlicht: (2023)
QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code
von: Fang, Hainan, et al.
Veröffentlicht: (2025)
von: Fang, Hainan, et al.
Veröffentlicht: (2025)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
Grammar-Based Code Representation: Is It a Worthy Pursuit for LLMs?
von: Liang, Qingyuan, et al.
Veröffentlicht: (2025)
von: Liang, Qingyuan, et al.
Veröffentlicht: (2025)
CatCode: A Comprehensive Evaluation Framework for LLMs On the Mixture of Code and Text
von: Lin, Zhenru, et al.
Veröffentlicht: (2024)
von: Lin, Zhenru, et al.
Veröffentlicht: (2024)
VisCoder2: Building Multi-Language Visualization Coding Agents
von: Ni, Yuansheng, et al.
Veröffentlicht: (2025)
von: Ni, Yuansheng, et al.
Veröffentlicht: (2025)
SaraCoder: Orchestrating Semantic and Structural Cues for Resource-Optimized Repository-Level Code Completion
von: Chen, Xiaohan, et al.
Veröffentlicht: (2025)
von: Chen, Xiaohan, et al.
Veröffentlicht: (2025)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
LILO: Learning Interpretable Libraries by Compressing and Documenting Code
von: Grand, Gabriel, et al.
Veröffentlicht: (2023)
von: Grand, Gabriel, et al.
Veröffentlicht: (2023)
CodeMind: Evaluating Large Language Models for Code Reasoning
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
CodeS: Natural Language to Code Repository via Multi-Layer Sketch
von: Zan, Daoguang, et al.
Veröffentlicht: (2024)
von: Zan, Daoguang, et al.
Veröffentlicht: (2024)
Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
CSSG: Measuring Code Similarity with Semantic Graphs
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench
von: Zou, Qingyun, et al.
Veröffentlicht: (2026) -
HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning
von: Zou, Qingyun, et al.
Veröffentlicht: (2026) -
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025) -
JudgeLRM: Large Reasoning Models as a Judge
von: Chen, Nuo, et al.
Veröffentlicht: (2025) -
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
von: Wang, Yanli, et al.
Veröffentlicht: (2024)