CatCode: A Comprehensive Evaluation Framework for LLMs On the Mixture of Code and Text
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Zhenru, Yao, Yiqun, Yuan, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Grammar-Based Code Representation: Is It a Worthy Pursuit for LLMs?
von: Liang, Qingyuan, et al.
Veröffentlicht: (2025)
von: Liang, Qingyuan, et al.
Veröffentlicht: (2025)
FPMoE: A Sparse Mixture-of-Experts Approach to Functional Code Generation
von: Pham, Loc, et al.
Veröffentlicht: (2026)
von: Pham, Loc, et al.
Veröffentlicht: (2026)
CodeMind: Evaluating Large Language Models for Code Reasoning
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
Assessing Code Understanding in LLMs
von: Laneve, Cosimo, et al.
Veröffentlicht: (2025)
von: Laneve, Cosimo, et al.
Veröffentlicht: (2025)
CodeV: Empowering LLMs with HDL Generation through Multi-Level Summarization
von: Zhao, Yang, et al.
Veröffentlicht: (2024)
von: Zhao, Yang, et al.
Veröffentlicht: (2024)
ECO: Enhanced Code Optimization via Performance-Aware Prompting for Code-LLMs
von: Kim, Su-Hyeon, et al.
Veröffentlicht: (2025)
von: Kim, Su-Hyeon, et al.
Veröffentlicht: (2025)
Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility
von: Maveli, Nickil, et al.
Veröffentlicht: (2026)
von: Maveli, Nickil, et al.
Veröffentlicht: (2026)
AutoCode: LLMs as Problem Setters for Competitive Programming
von: Zhou, Shang, et al.
Veröffentlicht: (2025)
von: Zhou, Shang, et al.
Veröffentlicht: (2025)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable Code
von: Zeng, Lingfei, et al.
Veröffentlicht: (2025)
von: Zeng, Lingfei, et al.
Veröffentlicht: (2025)
Code Repair with LLMs gives an Exploration-Exploitation Tradeoff
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes
von: Chon, Heejae, et al.
Veröffentlicht: (2024)
von: Chon, Heejae, et al.
Veröffentlicht: (2024)
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
von: Dai, Hankun, et al.
Veröffentlicht: (2025)
von: Dai, Hankun, et al.
Veröffentlicht: (2025)
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
von: Fang, Sen, et al.
Veröffentlicht: (2025)
von: Fang, Sen, et al.
Veröffentlicht: (2025)
From Prompts to Performance: Evaluating LLMs for Task-based Parallel Code Generation
von: Bantel, Linus, et al.
Veröffentlicht: (2026)
von: Bantel, Linus, et al.
Veröffentlicht: (2026)
AutoMCQ -- Automatically Generate Code Comprehension Questions using GenAI
von: Goodfellow, Martin, et al.
Veröffentlicht: (2025)
von: Goodfellow, Martin, et al.
Veröffentlicht: (2025)
CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules
von: Le, Hung, et al.
Veröffentlicht: (2023)
von: Le, Hung, et al.
Veröffentlicht: (2023)
Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation
von: Chen, Le, et al.
Veröffentlicht: (2025)
von: Chen, Le, et al.
Veröffentlicht: (2025)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2025)
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2025)
CSSG: Measuring Code Similarity with Semantic Graphs
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
Code Broker: A Multi-Agent System for Automated Code Quality Assessment
von: Attrah, Samer
Veröffentlicht: (2026)
von: Attrah, Samer
Veröffentlicht: (2026)
MCTS-SQL: Light-Weight LLMs can Master the Text-to-SQL through Monte Carlo Tree Search
von: Yuan, Shuozhi, et al.
Veröffentlicht: (2025)
von: Yuan, Shuozhi, et al.
Veröffentlicht: (2025)
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
von: Bae, Suyoung, et al.
Veröffentlicht: (2026)
von: Bae, Suyoung, et al.
Veröffentlicht: (2026)
Bench4HLS: End-to-End Evaluation of LLMs in High-Level Synthesis Code Generation
von: Khan, M Zafir Sadik, et al.
Veröffentlicht: (2026)
von: Khan, M Zafir Sadik, et al.
Veröffentlicht: (2026)
A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
MHRC-Bench: A Multilingual Hardware Repository-Level Code Completion benchmark
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
von: Zou, Qingyun, et al.
Veröffentlicht: (2026)
Executing as You Generate: Hiding Execution Latency in LLM Code Generation
von: Sun, Zhensu, et al.
Veröffentlicht: (2026)
von: Sun, Zhensu, et al.
Veröffentlicht: (2026)
Code Simulation Challenges for Large Language Models
von: La Malfa, Emanuele, et al.
Veröffentlicht: (2024)
von: La Malfa, Emanuele, et al.
Veröffentlicht: (2024)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
von: Peng, Yun, et al.
Veröffentlicht: (2024)
von: Peng, Yun, et al.
Veröffentlicht: (2024)
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
von: Shi, Yuling, et al.
Veröffentlicht: (2024)
von: Shi, Yuling, et al.
Veröffentlicht: (2024)
DOCE: Finding the Sweet Spot for Execution-Based Code Generation
von: Li, Haau-Sing, et al.
Veröffentlicht: (2024)
von: Li, Haau-Sing, et al.
Veröffentlicht: (2024)
Agentic Code Reasoning
von: Ugare, Shubham, et al.
Veröffentlicht: (2026)
von: Ugare, Shubham, et al.
Veröffentlicht: (2026)
Conditioning LLMs to Generate Code-Switched Text
von: Heredia, Maite, et al.
Veröffentlicht: (2025)
von: Heredia, Maite, et al.
Veröffentlicht: (2025)
REINFOREST: Reinforcing Semantic Code Similarity for Cross-Lingual Code Search Models
von: Saieva, Anthony, et al.
Veröffentlicht: (2023)
von: Saieva, Anthony, et al.
Veröffentlicht: (2023)
A Case Study on the Effectiveness of LLMs in Verification with Proof Assistants
von: Bayazıt, Barış, et al.
Veröffentlicht: (2025)
von: Bayazıt, Barış, et al.
Veröffentlicht: (2025)
LongCodeBench: Evaluating Coding LLMs at 1M Context Windows
von: Rando, Stefano, et al.
Veröffentlicht: (2025)
von: Rando, Stefano, et al.
Veröffentlicht: (2025)
A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks
von: Dandamudi, Rohit, et al.
Veröffentlicht: (2024)
von: Dandamudi, Rohit, et al.
Veröffentlicht: (2024)
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
von: Lang, Nguyet-Anh H., et al.
Veröffentlicht: (2026)
von: Lang, Nguyet-Anh H., et al.
Veröffentlicht: (2026)
Towards Formal Verification of LLM-Generated Code from Natural Language Prompts
von: Councilman, Aaron, et al.
Veröffentlicht: (2025)
von: Councilman, Aaron, et al.
Veröffentlicht: (2025)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Grammar-Based Code Representation: Is It a Worthy Pursuit for LLMs?
von: Liang, Qingyuan, et al.
Veröffentlicht: (2025) -
FPMoE: A Sparse Mixture-of-Experts Approach to Functional Code Generation
von: Pham, Loc, et al.
Veröffentlicht: (2026) -
CodeMind: Evaluating Large Language Models for Code Reasoning
von: Liu, Changshu, et al.
Veröffentlicht: (2024) -
Assessing Code Understanding in LLMs
von: Laneve, Cosimo, et al.
Veröffentlicht: (2025) -
CodeV: Empowering LLMs with HDL Generation through Multi-Level Summarization
von: Zhao, Yang, et al.
Veröffentlicht: (2024)