PyEvalAI: AI-assisted evaluation of Jupyter Notebooks for immediate personalized feedback
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wandel, Nils, Stotko, David, Schier, Alexander, Klein, Reinhard |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Physics-guided Shape-from-Template: Monocular Video Perception through Neural Surrogate Models
par: Stotko, David, et autres
Publié: (2023)
par: Stotko, David, et autres
Publié: (2023)
Metamizer: a versatile neural optimizer for fast and accurate physics simulations
par: Wandel, Nils, et autres
Publié: (2024)
par: Wandel, Nils, et autres
Publié: (2024)
JELAI: Integrating AI and Learning Analytics in Jupyter Notebooks
par: Torre, Manuel Valle, et autres
Publié: (2025)
par: Torre, Manuel Valle, et autres
Publié: (2025)
SAFT: Shape and Appearance of Fabrics from Template via Differentiable Physical Simulations from Monocular Video
par: Stotko, David, et autres
Publié: (2025)
par: Stotko, David, et autres
Publié: (2025)
Can AI expose tax loopholes? Towards a new generation of legal policy assistants
par: Fratrič, Peter, et autres
Publié: (2025)
par: Fratrič, Peter, et autres
Publié: (2025)
Debug Smarter, Not Harder: AI Agents for Error Resolution in Computational Notebooks
par: Grotov, Konstantin, et autres
Publié: (2024)
par: Grotov, Konstantin, et autres
Publié: (2024)
Generative AI in clinical practice: novel qualitative evidence of risk and responsible use of Google's NotebookLM
par: Reuter, Max, et autres
Publié: (2025)
par: Reuter, Max, et autres
Publié: (2025)
Towards AI-assisted Academic Writing
par: Liebling, Daniel J., et autres
Publié: (2025)
par: Liebling, Daniel J., et autres
Publié: (2025)
AI-assisted Code Authoring at Scale: Fine-tuning, deploying, and mixed methods evaluation
par: Murali, Vijayaraghavan, et autres
Publié: (2023)
par: Murali, Vijayaraghavan, et autres
Publié: (2023)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
par: Garg, Madhav Krishan, et autres
Publié: (2025)
par: Garg, Madhav Krishan, et autres
Publié: (2025)
Malware analysis assisted by AI with R2AI
par: Apvrille, Axelle, et autres
Publié: (2025)
par: Apvrille, Axelle, et autres
Publié: (2025)
Themisto: Jupyter-Based Runtime Benchmark
par: Grotov, Konstantin, et autres
Publié: (2025)
par: Grotov, Konstantin, et autres
Publié: (2025)
Academic journals' AI policies fail to curb the surge in AI-assisted academic writing
par: He, Yongyuan, et autres
Publié: (2025)
par: He, Yongyuan, et autres
Publié: (2025)
Evaluating the role of `Constitutions' for learning from AI feedback
par: Redgate, Saskia, et autres
Publié: (2024)
par: Redgate, Saskia, et autres
Publié: (2024)
Reducing research bureaucracy in UK higher education: Can generative AI assist with the internal evaluation of quality?
par: Fletcher, Gordon, et autres
Publié: (2025)
par: Fletcher, Gordon, et autres
Publié: (2025)
TextBO: Bayesian Optimization in Language Space for Eval-Efficient Self-Improving AI
par: Kang, Enoch Hyunwook, et autres
Publié: (2025)
par: Kang, Enoch Hyunwook, et autres
Publié: (2025)
AgentEval: Generative Agents as Reliable Proxies for Human Evaluation of AI-Generated Content
par: Vu, Thanh, et autres
Publié: (2025)
par: Vu, Thanh, et autres
Publié: (2025)
LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks
par: Gao, Hengjian, et autres
Publié: (2026)
par: Gao, Hengjian, et autres
Publié: (2026)
Quantifying truth and authenticity in AI-assisted candidate evaluation: A multi-domain pilot analysis
par: Lee, Eldred, et autres
Publié: (2025)
par: Lee, Eldred, et autres
Publié: (2025)
Conversational AI for Rapid Scientific Prototyping: A Case Study on ESA's ELOPE Competition
par: Einecke, Nils
Publié: (2026)
par: Einecke, Nils
Publié: (2026)
New care pathways for supporting transitional care from hospitals to home using AI and personalized digital assistance
par: Anghel, Ionut, et autres
Publié: (2025)
par: Anghel, Ionut, et autres
Publié: (2025)
PopPy: Opportunistically Exploiting Parallelism in Python Compound AI Applications
par: Mell, Stephen, et autres
Publié: (2026)
par: Mell, Stephen, et autres
Publié: (2026)
Towards evaluations-based safety cases for AI scheming
par: Balesni, Mikita, et autres
Publié: (2024)
par: Balesni, Mikita, et autres
Publié: (2024)
Human-Alignment Influences the Utility of AI-assisted Decision Making
par: Benz, Nina L. Corvelo, et autres
Publié: (2025)
par: Benz, Nina L. Corvelo, et autres
Publié: (2025)
Architectural Constraints Alignment in AI-assisted, Platform-based Service Development
par: Irion, Julius, et autres
Publié: (2026)
par: Irion, Julius, et autres
Publié: (2026)
PyGen: A Collaborative Human-AI Approach to Python Package Creation
par: Barua, Saikat, et autres
Publié: (2024)
par: Barua, Saikat, et autres
Publié: (2024)
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
par: Ristea, Dan, et autres
Publié: (2024)
par: Ristea, Dan, et autres
Publié: (2024)
Multi-line AI-assisted Code Authoring
par: Dunay, Omer, et autres
Publié: (2024)
par: Dunay, Omer, et autres
Publié: (2024)
In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
par: Durner, Nils
Publié: (2025)
par: Durner, Nils
Publié: (2025)
PsychEval: A Multi-Session and Multi-Therapy Benchmark for High-Realism AI Psychological Counselor
par: Pan, Qianjun, et autres
Publié: (2026)
par: Pan, Qianjun, et autres
Publié: (2026)
What happens when reviewers receive AI feedback in their reviews?
par: Chen, Shiping, et autres
Publié: (2026)
par: Chen, Shiping, et autres
Publié: (2026)
ROSA: Reconstructing Object Shape and Appearance Textures by Adaptive Detail Transfer
par: Kaltheuner, Julian, et autres
Publié: (2025)
par: Kaltheuner, Julian, et autres
Publié: (2025)
Log analysis is necessary for credible evaluation of AI agents
par: Kirgis, Peter, et autres
Publié: (2026)
par: Kirgis, Peter, et autres
Publié: (2026)
FPO++: Efficient Encoding and Rendering of Dynamic Neural Radiance Fields by Analyzing and Enhancing Fourier PlenOctrees
par: Rabich, Saskia, et autres
Publié: (2023)
par: Rabich, Saskia, et autres
Publié: (2023)
Incomplete Gamma Kernels: Generalizing Locally Optimal Projection Operators
par: Stotko, Patrick, et autres
Publié: (2022)
par: Stotko, Patrick, et autres
Publié: (2022)
ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation
par: Huang, Yizheng, et autres
Publié: (2026)
par: Huang, Yizheng, et autres
Publié: (2026)
ChildEval: When large language models meet children's personalities
par: Luo, Yanyan, et autres
Publié: (2026)
par: Luo, Yanyan, et autres
Publié: (2026)
AI in radiological imaging of soft-tissue and bone tumours: a systematic review evaluating against CLAIM and FUTURE-AI guidelines
par: Spaanderman, Douwe J., et autres
Publié: (2024)
par: Spaanderman, Douwe J., et autres
Publié: (2024)
Formative Study for AI-assisted Data Visualization
par: Saber, Rania, et autres
Publié: (2024)
par: Saber, Rania, et autres
Publié: (2024)
Exploring the Challenges and Opportunities of AI-assisted Codebase Generation
par: Eibl, Philipp, et autres
Publié: (2025)
par: Eibl, Philipp, et autres
Publié: (2025)
Documents similaires
-
Physics-guided Shape-from-Template: Monocular Video Perception through Neural Surrogate Models
par: Stotko, David, et autres
Publié: (2023) -
Metamizer: a versatile neural optimizer for fast and accurate physics simulations
par: Wandel, Nils, et autres
Publié: (2024) -
JELAI: Integrating AI and Learning Analytics in Jupyter Notebooks
par: Torre, Manuel Valle, et autres
Publié: (2025) -
SAFT: Shape and Appearance of Fabrics from Template via Differentiable Physical Simulations from Monocular Video
par: Stotko, David, et autres
Publié: (2025) -
Can AI expose tax loopholes? Towards a new generation of legal policy assistants
par: Fratrič, Peter, et autres
Publié: (2025)