Themisto: Jupyter-Based Runtime Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Grotov, Konstantin, Titov, Sergey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Untangling Knots: Leveraging LLM for Error Resolution in Computational Notebooks
by: Grotov, Konstantin, et al.
Published: (2024)
by: Grotov, Konstantin, et al.
Published: (2024)
PIPer: On-Device Environment Setup via Online Reinforcement Learning
by: Kovrigin, Alexander, et al.
Published: (2025)
by: Kovrigin, Alexander, et al.
Published: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024)
by: Galimzyanov, Timur, et al.
Published: (2024)
Observing Fine-Grained Changes in Jupyter Notebooks During Development Time
by: Titov, Sergey, et al.
Published: (2025)
by: Titov, Sergey, et al.
Published: (2025)
Hidden Gems in the Rough: Computational Notebooks as an Uncharted Oasis for IDEs
by: Titov, Sergey, et al.
Published: (2024)
by: Titov, Sergey, et al.
Published: (2024)
AIPC: Agent-Based Automation for AI Model Deployment with Qualcomm AI Runtime
by: Su, Jianhao, et al.
Published: (2026)
by: Su, Jianhao, et al.
Published: (2026)
Debugging and Runtime Analysis of Neural Networks with VLMs (A Case Study)
by: Hu, Boyue Caroline, et al.
Published: (2025)
by: Hu, Boyue Caroline, et al.
Published: (2025)
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
by: Samsonau, Sergey V.
Published: (2026)
by: Samsonau, Sergey V.
Published: (2026)
Are Large Language Models Memorizing Bug Benchmarks?
by: Ramos, Daniel, et al.
Published: (2024)
by: Ramos, Daniel, et al.
Published: (2024)
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
by: Gu, Alex, et al.
Published: (2024)
by: Gu, Alex, et al.
Published: (2024)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
by: Lange, Robert Tjarko, et al.
Published: (2025)
by: Lange, Robert Tjarko, et al.
Published: (2025)
SoundnessBench: A Soundness Benchmark for Neural Network Verifiers
by: Zhou, Xingjian, et al.
Published: (2024)
by: Zhou, Xingjian, et al.
Published: (2024)
Real Faults in Deep Learning Fault Benchmarks: How Real Are They?
by: Jahangirova, Gunel, et al.
Published: (2024)
by: Jahangirova, Gunel, et al.
Published: (2024)
MTAD: Tools and Benchmarks for Multivariate Time Series Anomaly Detection
by: Liu, Jinyang, et al.
Published: (2024)
by: Liu, Jinyang, et al.
Published: (2024)
Benchmarking Large Language Models with Integer Sequence Generation Tasks
by: O'Malley, Daniel, et al.
Published: (2024)
by: O'Malley, Daniel, et al.
Published: (2024)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
by: Deshpande, Darshan, et al.
Published: (2026)
by: Deshpande, Darshan, et al.
Published: (2026)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
by: Xie, Zichen, et al.
Published: (2026)
by: Xie, Zichen, et al.
Published: (2026)
miniCodeProps: a Minimal Benchmark for Proving Code Properties
by: Lohn, Evan, et al.
Published: (2024)
by: Lohn, Evan, et al.
Published: (2024)
CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification
by: Xu, Jiacheng, et al.
Published: (2025)
by: Xu, Jiacheng, et al.
Published: (2025)
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
by: Shbita, Basel, et al.
Published: (2025)
by: Shbita, Basel, et al.
Published: (2025)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
by: Xiao, Yijia, et al.
Published: (2025)
by: Xiao, Yijia, et al.
Published: (2025)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization
by: Wang, Juntong, et al.
Published: (2026)
by: Wang, Juntong, et al.
Published: (2026)
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
by: Feng, Yunfei, et al.
Published: (2026)
by: Feng, Yunfei, et al.
Published: (2026)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
by: Qiu, Ruizhong, et al.
Published: (2024)
by: Qiu, Ruizhong, et al.
Published: (2024)
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
by: Jha, Smriti, et al.
Published: (2026)
by: Jha, Smriti, et al.
Published: (2026)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
by: Majdinasab, Vahid, et al.
Published: (2025)
by: Majdinasab, Vahid, et al.
Published: (2025)
WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks
by: Wornow, Michael, et al.
Published: (2024)
by: Wornow, Michael, et al.
Published: (2024)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
by: Li, Yuangang, et al.
Published: (2026)
by: Li, Yuangang, et al.
Published: (2026)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
Lookup multivariate Kolmogorov-Arnold Networks
by: Pozdnyakov, Sergey, et al.
Published: (2025)
by: Pozdnyakov, Sergey, et al.
Published: (2025)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
by: Diehl, Patrick, et al.
Published: (2025)
by: Diehl, Patrick, et al.
Published: (2025)
Evaluating Robustness of Large Language Models in Enterprise Applications: Benchmarks for Perturbation Consistency Across Formats and Languages
by: Bogavelli, Tara, et al.
Published: (2026)
by: Bogavelli, Tara, et al.
Published: (2026)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
by: Zheng, Qinkai, et al.
Published: (2023)
by: Zheng, Qinkai, et al.
Published: (2023)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
by: Xu, WeiZhe, et al.
Published: (2026)
by: Xu, WeiZhe, et al.
Published: (2026)
Natural Language Requirements Testability Measurement Based on Requirement Smells
by: Zakeri-Nasrabadi, Morteza, et al.
Published: (2024)
by: Zakeri-Nasrabadi, Morteza, et al.
Published: (2024)
Analysing the Behaviour of Tree-Based Neural Networks in Regression Tasks
by: Samoaa, Peter, et al.
Published: (2024)
by: Samoaa, Peter, et al.
Published: (2024)
Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family Classification
by: Wu, Zehao, et al.
Published: (2025)
by: Wu, Zehao, et al.
Published: (2025)
Similar Items
-
Untangling Knots: Leveraging LLM for Error Resolution in Computational Notebooks
by: Grotov, Konstantin, et al.
Published: (2024) -
PIPer: On-Device Environment Setup via Online Reinforcement Learning
by: Kovrigin, Alexander, et al.
Published: (2025) -
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024) -
Observing Fine-Grained Changes in Jupyter Notebooks During Development Time
by: Titov, Sergey, et al.
Published: (2025) -
Hidden Gems in the Rough: Computational Notebooks as an Uncharted Oasis for IDEs
by: Titov, Sergey, et al.
Published: (2024)