Evaluating Code Generation of LLMs in Advanced Computer Science Problems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Catir, Emir, Claesson, Robin, Tsoupidi, Rodothea Myrsini |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Protecting Cryptographic Libraries against Side-Channel and Code-Reuse Attacks
von: Tsoupidi, Rodothea Myrsini, et al.
Veröffentlicht: (2024)
von: Tsoupidi, Rodothea Myrsini, et al.
Veröffentlicht: (2024)
Adaptive Generation of Bias-Eliciting Questions for LLMs
von: Staab, Robin, et al.
Veröffentlicht: (2025)
von: Staab, Robin, et al.
Veröffentlicht: (2025)
AI Mentors for Student Projects: Spotting Early Issues in Computer Science Proposals
von: Aher, Gati, et al.
Veröffentlicht: (2025)
von: Aher, Gati, et al.
Veröffentlicht: (2025)
Homoglyph-based Adversarial Perturbation of Introductory Computer Science Theory Problems
von: Alexander, Aidan, et al.
Veröffentlicht: (2026)
von: Alexander, Aidan, et al.
Veröffentlicht: (2026)
ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle
von: Miroyan, Mihran, et al.
Veröffentlicht: (2025)
von: Miroyan, Mihran, et al.
Veröffentlicht: (2025)
Advancing Problem-Based Learning in Biomedical Engineering in the Era of Generative AI
von: Nnamdi, Micky C., et al.
Veröffentlicht: (2025)
von: Nnamdi, Micky C., et al.
Veröffentlicht: (2025)
A Comparative Study of Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics
von: Liu, Suqing, et al.
Veröffentlicht: (2025)
von: Liu, Suqing, et al.
Veröffentlicht: (2025)
The Provenance Problem: LLMs and the Breakdown of Citation Norms
von: Earp, Brian D., et al.
Veröffentlicht: (2025)
von: Earp, Brian D., et al.
Veröffentlicht: (2025)
Computational Thinking with Computer Vision: Developing AI Competency in an Introductory Computer Science Course
von: Chowdhury, Tahiya
Veröffentlicht: (2025)
von: Chowdhury, Tahiya
Veröffentlicht: (2025)
A Review of Generative AI in Computer Science Education: Challenges and Opportunities in Accuracy, Authenticity, and Assessment
von: Reihanian, Iman, et al.
Veröffentlicht: (2025)
von: Reihanian, Iman, et al.
Veröffentlicht: (2025)
Prompt Programming: A Platform for Dialogue-based Computational Problem Solving with Generative AI Models
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
Ensuring Computer Science Learning in the AI Era: Open Generative AI Policies and Assignment-Driven Written Quizzes
von: Chung, Chan-Jin
Veröffentlicht: (2026)
von: Chung, Chan-Jin
Veröffentlicht: (2026)
Automated Assessment of Students' Code Comprehension using LLMs
von: Oli, Priti, et al.
Veröffentlicht: (2023)
von: Oli, Priti, et al.
Veröffentlicht: (2023)
PaperRepro: Automated Computational Reproducibility Assessment for Social Science Papers
von: Zhang, Linhao, et al.
Veröffentlicht: (2026)
von: Zhang, Linhao, et al.
Veröffentlicht: (2026)
Prompt Design Matters for Computational Social Science Tasks but in Unpredictable Ways
von: Atreja, Shubham, et al.
Veröffentlicht: (2024)
von: Atreja, Shubham, et al.
Veröffentlicht: (2024)
Arti-"fickle" Intelligence: Using LLMs as a Tool for Inference in the Political and Social Sciences
von: Argyle, Lisa P., et al.
Veröffentlicht: (2025)
von: Argyle, Lisa P., et al.
Veröffentlicht: (2025)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
von: Rios-Sialer, Ian
Veröffentlicht: (2026)
von: Rios-Sialer, Ian
Veröffentlicht: (2026)
Navigating Pitfalls: Evaluating LLMs in Machine Learning Programming Education
von: Kumar, Smitha, et al.
Veröffentlicht: (2025)
von: Kumar, Smitha, et al.
Veröffentlicht: (2025)
The Potential of Answer Classes in Large-scale Written Computer-Science Exams -- Vol. 2
von: Lohr, Dominic, et al.
Veröffentlicht: (2024)
von: Lohr, Dominic, et al.
Veröffentlicht: (2024)
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
Theory Trace Card: Theory-Driven Socio-Cognitive Evaluation of LLMs
von: Karimi-Malekabadi, Farzan, et al.
Veröffentlicht: (2026)
von: Karimi-Malekabadi, Farzan, et al.
Veröffentlicht: (2026)
WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making
von: Li, Zongjie, et al.
Veröffentlicht: (2026)
von: Li, Zongjie, et al.
Veröffentlicht: (2026)
Reinforcement Learning and Life Cycle Assessment for a Circular Economy -- Towards Progressive Computer Science
von: Buchner, Johannes
Veröffentlicht: (2025)
von: Buchner, Johannes
Veröffentlicht: (2025)
Intelligent Computing Social Modeling and Methodological Innovations in Political Science in the Era of Large Language Models
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
Generative AI for Education (GAIED): Advances, Opportunities, and Challenges
von: Denny, Paul, et al.
Veröffentlicht: (2024)
von: Denny, Paul, et al.
Veröffentlicht: (2024)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
von: Ghosh, Rajarshi, et al.
Veröffentlicht: (2025)
von: Ghosh, Rajarshi, et al.
Veröffentlicht: (2025)
On the Credibility of Evaluating LLMs using Survey Questions
von: Libovický, Jindřich
Veröffentlicht: (2026)
von: Libovický, Jindřich
Veröffentlicht: (2026)
Toward an Engineering of Science: Rebalancing Generation and Verification in the Age of AI
von: Ma, Jiaqi W.
Veröffentlicht: (2026)
von: Ma, Jiaqi W.
Veröffentlicht: (2026)
Computational Hermeneutics: Evaluating generative AI as a cultural technology
von: Kommers, Cody, et al.
Veröffentlicht: (2026)
von: Kommers, Cody, et al.
Veröffentlicht: (2026)
Measuring and Mitigating Bias in Code Generated by Large Language Models
von: Chen, Yuxi, et al.
Veröffentlicht: (2026)
von: Chen, Yuxi, et al.
Veröffentlicht: (2026)
Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users
von: Kempermann, Manon, et al.
Veröffentlicht: (2025)
von: Kempermann, Manon, et al.
Veröffentlicht: (2025)
What Is Actually Being Annotated? Inter-Prompt Reliability as a Measurement Problem in LLM-Based Social Science Labeling
von: Liu, Jingyuan
Veröffentlicht: (2026)
von: Liu, Jingyuan
Veröffentlicht: (2026)
Responsible Data Stewardship: Generative AI and the Digital Waste Problem
von: Utz, Vanessa
Veröffentlicht: (2025)
von: Utz, Vanessa
Veröffentlicht: (2025)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
von: Dawson, Fiifi, et al.
Veröffentlicht: (2024)
von: Dawson, Fiifi, et al.
Veröffentlicht: (2024)
MoPHES:Leveraging on-device LLMs as Agent for Mobile Psychological Health Evaluation and Support
von: Wei, Xun, et al.
Veröffentlicht: (2025)
von: Wei, Xun, et al.
Veröffentlicht: (2025)
Adoption and Impact of ChatGPT in Computer Science Education: A Case Study on a Database Administration Course
von: López-Fernández, Daniel, et al.
Veröffentlicht: (2024)
von: López-Fernández, Daniel, et al.
Veröffentlicht: (2024)
Simulating Validity: Modal Decoupling in MLLM Generated Feedback on Science Drawings
von: Bewersdorff, Arne, et al.
Veröffentlicht: (2026)
von: Bewersdorff, Arne, et al.
Veröffentlicht: (2026)
Evaluating Generative AI for CS1 Code Grading: Direct vs Reverse Methods
von: Memon, Ahmad, et al.
Veröffentlicht: (2025)
von: Memon, Ahmad, et al.
Veröffentlicht: (2025)
Developing a Multi-Agent System to Generate Next Generation Science Assessments with Evidence-Centered Design
von: Yang, Yaxuan, et al.
Veröffentlicht: (2026)
von: Yang, Yaxuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Protecting Cryptographic Libraries against Side-Channel and Code-Reuse Attacks
von: Tsoupidi, Rodothea Myrsini, et al.
Veröffentlicht: (2024) -
Adaptive Generation of Bias-Eliciting Questions for LLMs
von: Staab, Robin, et al.
Veröffentlicht: (2025) -
AI Mentors for Student Projects: Spotting Early Issues in Computer Science Proposals
von: Aher, Gati, et al.
Veröffentlicht: (2025) -
Homoglyph-based Adversarial Perturbation of Introductory Computer Science Theory Problems
von: Alexander, Aidan, et al.
Veröffentlicht: (2026) -
ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle
von: Miroyan, Mihran, et al.
Veröffentlicht: (2025)