Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
Fuente:
arXiv
Salvato in:
| Autori principali: | Galimzyanov, Timur, Titov, Sergey, Golubev, Yaroslav, Bogomolov, Egor |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
di: Bogomolov, Egor, et al.
Pubblicazione: (2024)
di: Bogomolov, Egor, et al.
Pubblicazione: (2024)
Dynamic Retrieval-Augmented Generation
di: Shapkin, Anton, et al.
Pubblicazione: (2023)
di: Shapkin, Anton, et al.
Pubblicazione: (2023)
Challenge on Optimization of Context Collection for Code Completion
di: Ustalov, Dmitry, et al.
Pubblicazione: (2025)
di: Ustalov, Dmitry, et al.
Pubblicazione: (2025)
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
di: Lindenbauer, Tobias, et al.
Pubblicazione: (2025)
di: Lindenbauer, Tobias, et al.
Pubblicazione: (2025)
PIPer: On-Device Environment Setup via Online Reinforcement Learning
di: Kovrigin, Alexander, et al.
Pubblicazione: (2025)
di: Kovrigin, Alexander, et al.
Pubblicazione: (2025)
Step Rejection Fine-Tuning: A Practical Distillation Recipe
di: Slinko, Igor, et al.
Pubblicazione: (2026)
di: Slinko, Igor, et al.
Pubblicazione: (2026)
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
di: Glukhov, Evgeniy, et al.
Pubblicazione: (2025)
di: Glukhov, Evgeniy, et al.
Pubblicazione: (2025)
Themisto: Jupyter-Based Runtime Benchmark
di: Grotov, Konstantin, et al.
Pubblicazione: (2025)
di: Grotov, Konstantin, et al.
Pubblicazione: (2025)
Stack Trace Deduplication: Faster, More Accurately, and in More Realistic Scenarios
di: Shibaev, Egor, et al.
Pubblicazione: (2024)
di: Shibaev, Egor, et al.
Pubblicazione: (2024)
On Problems of Implicit Context Compression for Software Engineering Agents
di: Gelvan, Kirill, et al.
Pubblicazione: (2026)
di: Gelvan, Kirill, et al.
Pubblicazione: (2026)
EnvBench: A Benchmark for Automated Environment Setup
di: Eliseeva, Aleksandra, et al.
Pubblicazione: (2025)
di: Eliseeva, Aleksandra, et al.
Pubblicazione: (2025)
Practical Code RAG at Scale: Task-Aware Retrieval Design Choices under Compute Budgets
di: Galimzyanov, Timur, et al.
Pubblicazione: (2025)
di: Galimzyanov, Timur, et al.
Pubblicazione: (2025)
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
di: Lindenbauer, Tobias, et al.
Pubblicazione: (2025)
di: Lindenbauer, Tobias, et al.
Pubblicazione: (2025)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
di: Patel, Harsh, et al.
Pubblicazione: (2024)
di: Patel, Harsh, et al.
Pubblicazione: (2024)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
di: Majdinasab, Vahid, et al.
Pubblicazione: (2025)
di: Majdinasab, Vahid, et al.
Pubblicazione: (2025)
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
di: Kovrigin, Alexander, et al.
Pubblicazione: (2024)
di: Kovrigin, Alexander, et al.
Pubblicazione: (2024)
Operational Robustness of LLMs on Code Generation
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2026)
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2026)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
di: Thillen, Alex, et al.
Pubblicazione: (2026)
di: Thillen, Alex, et al.
Pubblicazione: (2026)
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
di: Samsonau, Sergey V.
Pubblicazione: (2026)
di: Samsonau, Sergey V.
Pubblicazione: (2026)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
di: Bodla, Krishna Vamshi, et al.
Pubblicazione: (2025)
di: Bodla, Krishna Vamshi, et al.
Pubblicazione: (2025)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
di: Xie, Zichen, et al.
Pubblicazione: (2026)
di: Xie, Zichen, et al.
Pubblicazione: (2026)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
di: Daghighfarsoodeh, Alireza, et al.
Pubblicazione: (2025)
di: Daghighfarsoodeh, Alireza, et al.
Pubblicazione: (2025)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
di: Zheng, Qinkai, et al.
Pubblicazione: (2023)
di: Zheng, Qinkai, et al.
Pubblicazione: (2023)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2026)
di: Kumarappan, Adarsh, et al.
Pubblicazione: (2026)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
Untangling Knots: Leveraging LLM for Error Resolution in Computational Notebooks
di: Grotov, Konstantin, et al.
Pubblicazione: (2024)
di: Grotov, Konstantin, et al.
Pubblicazione: (2024)
Mellum: Production-Grade in-IDE Contextual Code Completion with Multi-File Project Understanding
di: Pavlichenko, Nikita, et al.
Pubblicazione: (2025)
di: Pavlichenko, Nikita, et al.
Pubblicazione: (2025)
On LLMs' Internal Representation of Code Correctness
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
di: Gu, Alex, et al.
Pubblicazione: (2024)
di: Gu, Alex, et al.
Pubblicazione: (2024)
miniCodeProps: a Minimal Benchmark for Proving Code Properties
di: Lohn, Evan, et al.
Pubblicazione: (2024)
di: Lohn, Evan, et al.
Pubblicazione: (2024)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
di: Rahman, Musfiqur, et al.
Pubblicazione: (2025)
di: Rahman, Musfiqur, et al.
Pubblicazione: (2025)
Evaluating the Use of LLMs for Documentation to Code Traceability
di: Alor, Ebube, et al.
Pubblicazione: (2025)
di: Alor, Ebube, et al.
Pubblicazione: (2025)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
di: Cipollone, Daniele, et al.
Pubblicazione: (2025)
di: Cipollone, Daniele, et al.
Pubblicazione: (2025)
Free and Customizable Code Documentation with LLMs: A Fine-Tuning Approach
di: Chakrabarty, Sayak, et al.
Pubblicazione: (2024)
di: Chakrabarty, Sayak, et al.
Pubblicazione: (2024)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
di: Rodriguez-Cardenas, Daniel, et al.
Pubblicazione: (2025)
di: Rodriguez-Cardenas, Daniel, et al.
Pubblicazione: (2025)
The Struggles of LLMs in Cross-lingual Code Clone Detection
di: Moumoula, Micheline Bénédicte, et al.
Pubblicazione: (2024)
di: Moumoula, Micheline Bénédicte, et al.
Pubblicazione: (2024)
LLMs in Coding and their Impact on the Commercial Software Engineering Landscape
di: Belozerov, Vladislav, et al.
Pubblicazione: (2025)
di: Belozerov, Vladislav, et al.
Pubblicazione: (2025)
Rethinking Repetition Problems of LLMs in Code Generation
di: Dong, Yihong, et al.
Pubblicazione: (2025)
di: Dong, Yihong, et al.
Pubblicazione: (2025)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
di: Vulićević, Jelena Ilić
Pubblicazione: (2026)
di: Vulićević, Jelena Ilić
Pubblicazione: (2026)
Documenti analoghi
-
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
di: Bogomolov, Egor, et al.
Pubblicazione: (2024) -
Dynamic Retrieval-Augmented Generation
di: Shapkin, Anton, et al.
Pubblicazione: (2023) -
Challenge on Optimization of Context Collection for Code Completion
di: Ustalov, Dmitry, et al.
Pubblicazione: (2025) -
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
di: Lindenbauer, Tobias, et al.
Pubblicazione: (2025) -
PIPer: On-Device Environment Setup via Online Reinforcement Learning
di: Kovrigin, Alexander, et al.
Pubblicazione: (2025)