Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Donglin, Li, Daming, Shi, Hanyuan, Zhang, Jialu |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch
par: Si, Yuan, et autres
Publié: (2025)
par: Si, Yuan, et autres
Publié: (2025)
ScratchEval : A Multimodal Evaluation Framework for LLMs in Block-Based Programming
par: Si, Yuan, et autres
Publié: (2026)
par: Si, Yuan, et autres
Publié: (2026)
EcoScratch: Cost-Effective Multimodal Repair for Scratch Using Execution Feedback
par: Si, Yuan, et autres
Publié: (2026)
par: Si, Yuan, et autres
Publié: (2026)
Stitch: Step-by-step LLM Guided Tutoring for Scratch
par: Si, Yuan, et autres
Publié: (2025)
par: Si, Yuan, et autres
Publié: (2025)
A Systematic Study of Time Limit Exceeded Errors in Online Programming Assignments
par: Zhang, Jialu, et autres
Publié: (2025)
par: Zhang, Jialu, et autres
Publié: (2025)
Empirical Evaluation of Large Language Models in Automated Program Repair
par: Sun, Jiajun, et autres
Publié: (2025)
par: Sun, Jiajun, et autres
Publié: (2025)
HerAgent: Rethinking the Automated Environment Deployment via Hierarchical Test Pyramid
par: Li, Xiang, et autres
Publié: (2026)
par: Li, Xiang, et autres
Publié: (2026)
Evaluating the Generalizability of LLMs in Automated Program Repair
par: Li, Fengjie, et autres
Publié: (2025)
par: Li, Fengjie, et autres
Publié: (2025)
Hybrid Automated Program Repair by Combining Large Language Models and Program Analysis
par: Li, Fengjie, et autres
Publié: (2024)
par: Li, Fengjie, et autres
Publié: (2024)
ProgramBench: Can Language Models Rebuild Programs From Scratch?
par: Yang, John, et autres
Publié: (2026)
par: Yang, John, et autres
Publié: (2026)
ComPass: Contrastive Learning for Automated Patch Correctness Assessment in Program Repair
par: Zhang, Quanjun, et autres
Publié: (2026)
par: Zhang, Quanjun, et autres
Publié: (2026)
A Survey on Feedback Types in Automated Programming Assessment Systems
par: Frankford, Eduard, et autres
Publié: (2025)
par: Frankford, Eduard, et autres
Publié: (2025)
An Online Integrated Development Environment for Automated Programming Assessment Systems
par: Frankford, Eduard, et autres
Publié: (2025)
par: Frankford, Eduard, et autres
Publié: (2025)
BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models
par: Li, Yuanhao, et autres
Publié: (2026)
par: Li, Yuanhao, et autres
Publié: (2026)
Enhancing NeuroEvolution-Based Game Testing: A Branch Coverage Approach for Scratch Programs
par: Sohail, Khizra, et autres
Publié: (2025)
par: Sohail, Khizra, et autres
Publié: (2025)
SpecGen: Automated Generation of Formal Program Specifications via Large Language Models
par: Ma, Lezhi, et autres
Publié: (2024)
par: Ma, Lezhi, et autres
Publié: (2024)
RepoZero: Can LLMs Generate a Code Repository from Scratch?
par: Zhang, Zhaoxi, et autres
Publié: (2026)
par: Zhang, Zhaoxi, et autres
Publié: (2026)
DebugRepair: Enhancing LLM-Based Automated Program Repair via Self-Directed Debugging
par: Wu, Linhao, et autres
Publié: (2026)
par: Wu, Linhao, et autres
Publié: (2026)
Voice-Controlled Scratch for Children with (Motor) Disabilities
par: Goller, Elias, et autres
Publié: (2026)
par: Goller, Elias, et autres
Publié: (2026)
Is Measurement Enough? Rethinking Output Validation in Quantum Program Testing
par: Ye, Jiaming, et autres
Publié: (2025)
par: Ye, Jiaming, et autres
Publié: (2025)
Chatbot-Based Assessment of Code Understanding in Automated Programming Assessment Systems
par: Frankford, Eduard, et autres
Publié: (2026)
par: Frankford, Eduard, et autres
Publié: (2026)
ThinkRepair: Self-Directed Automated Program Repair
par: Yin, Xin, et autres
Publié: (2024)
par: Yin, Xin, et autres
Publié: (2024)
Rethinking Cognitive Complexity for Unit Tests: Toward a Readability-Aware Metric Grounded in Developer Perception
par: Ouédraogo, Wendkûuni C., et autres
Publié: (2025)
par: Ouédraogo, Wendkûuni C., et autres
Publié: (2025)
The Impact of Program Reduction on Automated Program Repair
par: Vidziunas, Linas, et autres
Publié: (2024)
par: Vidziunas, Linas, et autres
Publié: (2024)
Rethinking Legal Compliance Automation: Opportunities with Large Language Models
par: Hassani, Shabnam, et autres
Publié: (2024)
par: Hassani, Shabnam, et autres
Publié: (2024)
Automated Assessment in Mobile Programming Courses: Leveraging GitHub Classroom and Flutter for Enhanced Student Outcomes
par: Alves, Pedro, et autres
Publié: (2025)
par: Alves, Pedro, et autres
Publié: (2025)
Logging Like Humans for LLMs: Rethinking Logging via Execution and Runtime Feedback
par: Wang, Xin, et autres
Publié: (2026)
par: Wang, Xin, et autres
Publié: (2026)
Boosting Redundancy-based Automated Program Repair by Fine-grained Pattern Mining
par: Jiang, Jiajun, et autres
Publié: (2023)
par: Jiang, Jiajun, et autres
Publié: (2023)
RepoScope: Leveraging Call Chain-Aware Multi-View Context for Repository-Level Code Generation
par: Liu, Yang, et autres
Publié: (2025)
par: Liu, Yang, et autres
Publié: (2025)
Less Training, More Repairing Please: Revisiting Automated Program Repair via Zero-shot Learning
par: Xia, Chunqiu Steven, et autres
Publié: (2022)
par: Xia, Chunqiu Steven, et autres
Publié: (2022)
Integrating Symbolic Execution with LLMs for Automated Generation of Program Specifications
par: Yang, Fanpeng, et autres
Publié: (2025)
par: Yang, Fanpeng, et autres
Publié: (2025)
SoK: Automated Vulnerability Repair: Methods, Tools, and Assessments
par: Hu, Yiwei, et autres
Publié: (2025)
par: Hu, Yiwei, et autres
Publié: (2025)
Rethinking the Capability of Fine-Tuned Language Models for Automated Vulnerability Repair
par: Han, Woorim, et autres
Publié: (2025)
par: Han, Woorim, et autres
Publié: (2025)
SpecEval: Evaluating Code Comprehension in Large Language Models via Program Specifications
par: Ma, Lezhi, et autres
Publié: (2024)
par: Ma, Lezhi, et autres
Publié: (2024)
Exploring and Lifting the Robustness of LLM-powered Automated Program Repair with Metamorphic Testing
par: Xue, Pengyu, et autres
Publié: (2024)
par: Xue, Pengyu, et autres
Publié: (2024)
Rethinking Kernel Program Repair: Benchmarking and Enhancing LLMs with RGym
par: Shehada, Kareem, et autres
Publié: (2025)
par: Shehada, Kareem, et autres
Publié: (2025)
Characterizing Real-World Bugs in Tile Programs for Automated Bug Detection
par: Rathnasuriya, Ravishka, et autres
Publié: (2026)
par: Rathnasuriya, Ravishka, et autres
Publié: (2026)
Automated Static Warning Identification via Path-based Semantic Representation
par: Zhang, Yuwei, et autres
Publié: (2023)
par: Zhang, Yuwei, et autres
Publié: (2023)
Enhancing Automated Program Repair with Solution Design
par: Zhao, Jiuang, et autres
Publié: (2024)
par: Zhao, Jiuang, et autres
Publié: (2024)
From Task to Tutorial: An Automated GUI Framework for Excel Tutorial Document and Video Creation
par: Xie, Yuhang, et autres
Publié: (2025)
par: Xie, Yuhang, et autres
Publié: (2025)
Documents similaires
-
ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch
par: Si, Yuan, et autres
Publié: (2025) -
ScratchEval : A Multimodal Evaluation Framework for LLMs in Block-Based Programming
par: Si, Yuan, et autres
Publié: (2026) -
EcoScratch: Cost-Effective Multimodal Repair for Scratch Using Execution Feedback
par: Si, Yuan, et autres
Publié: (2026) -
Stitch: Step-by-step LLM Guided Tutoring for Scratch
par: Si, Yuan, et autres
Publié: (2025) -
A Systematic Study of Time Limit Exceeded Errors in Online Programming Assignments
par: Zhang, Jialu, et autres
Publié: (2025)