RETUYT-INCO at BEA 2025 Shared Task: How Far Can Lightweight Models Go in AI-powered Tutor Evaluation?
Fuente:
arXiv
Saved in:
| Main Authors: | Góngora, Santiago, Sastre, Ignacio, Robaina, Santiago, Remersaro, Ignacio, Chiruzzo, Luis, Rosá, Aiala |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RETUYT-INCO at BEA 2026 Shared Task 2: Meta-prompting in Rubric-based Scoring for German
by: Sastre, Ignacio, et al.
Published: (2026)
by: Sastre, Ignacio, et al.
Published: (2026)
A Platform for Generating Educational Activities to Teach English as a Second Language
by: Rosá, Aiala, et al.
Published: (2025)
by: Rosá, Aiala, et al.
Published: (2025)
Memory Tokens: Large Language Models Can Generate Reversible Sentence Embeddings
by: Sastre, Ignacio, et al.
Published: (2025)
by: Sastre, Ignacio, et al.
Published: (2025)
Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions
by: Sastre, Ignacio, et al.
Published: (2026)
by: Sastre, Ignacio, et al.
Published: (2026)
Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
by: Kochmar, Ekaterina, et al.
Published: (2025)
by: Kochmar, Ekaterina, et al.
Published: (2025)
Derivation Prompting: A Logic-Based Method for Improving Retrieval-Augmented Generation
by: Sastre, Ignacio, et al.
Published: (2026)
by: Sastre, Ignacio, et al.
Published: (2026)
NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
BD at BEA 2025 Shared Task: MPNet Ensembles for Pedagogical Mistake Identification and Localization in AI Tutor Responses
by: Rohan, Shadman, et al.
Published: (2025)
by: Rohan, Shadman, et al.
Published: (2025)
PAYADOR: A Minimalist Approach to Grounding Language Models on Structured Data for Interactive Storytelling and Role-playing Games
by: Góngora, Santiago, et al.
Published: (2025)
by: Góngora, Santiago, et al.
Published: (2025)
MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors
by: Hikal, Baraa, et al.
Published: (2025)
by: Hikal, Baraa, et al.
Published: (2025)
Ensayos físico-químicos para el estudio de la degradación de bolsas de supermercado
by: J. Remersaro
Published: (2010)
by: J. Remersaro
Published: (2010)
The ADAIO System at the BEA-2023 Shared Task on Generating AI Teacher Responses in Educational Dialogues
by: Adigwe, Adaeze, et al.
Published: (2023)
by: Adigwe, Adaeze, et al.
Published: (2023)
How Far Can You Go with Your Studies at UNED?
by: Cristina Gutiérrez-Carranza
Published: (2020)
by: Cristina Gutiérrez-Carranza
Published: (2020)
Sakura at BEA 2026 Shared Task 1: What Makes Vocabulary Difficult?
by: Nohejl, Adam, et al.
Published: (2026)
by: Nohejl, Adam, et al.
Published: (2026)
EmocioCine: Evaluación de la Inteligencia Emocional en población infanto-juvenil a partir de escenas cinematográficas
by: Santiago Sastre
Published: (2023)
by: Santiago Sastre
Published: (2023)
Financial Named Entity Recognition: How Far Can LLM Go?
by: Lu, Yi-Te, et al.
Published: (2025)
by: Lu, Yi-Te, et al.
Published: (2025)
Breaking the Degeneracy of Sense Codons – How Far Can We Go?
by: Clark A. Jones, et al.
Published: (2024)
by: Clark A. Jones, et al.
Published: (2024)
Consideraciones actuales sobre el embarazo en la adolescencia
by: José Ignacio Robaina-Castillo
Published: (2019)
by: José Ignacio Robaina-Castillo
Published: (2019)
How Far Can In-Context Alignment Go? Exploring the State of In-Context Alignment
by: Huang, Heyan, et al.
Published: (2024)
by: Huang, Heyan, et al.
Published: (2024)
LLM Compression: How Far Can We Go in Balancing Size and Performance?
by: Sk, Sahil, et al.
Published: (2025)
by: Sk, Sahil, et al.
Published: (2025)
How Far Can We Go with Practical Function-Level Program Repair?
by: Xiang, Jiahong, et al.
Published: (2024)
by: Xiang, Jiahong, et al.
Published: (2024)
How Far do Lindbladians Go?
by: Cai, Jihong, et al.
Published: (2025)
by: Cai, Jihong, et al.
Published: (2025)
CON LA CULTURA A CUESTAS. LA CENTRALIDAD DE LA PERSONA EN EL DEBATE MULTICULTURAL
by: Santiago Sastre Ariza
Published: (2007)
by: Santiago Sastre Ariza
Published: (2007)
PARA VER CON MEJOR LUZ. UNAAPROXIMACIÓN AL TRABAJO DE LA DOGMÁTICA JURÍDICA
by: Santiago Sastre Ariza
Published: (2006)
by: Santiago Sastre Ariza
Published: (2006)
CON LA CULTURA A CUESTAS. LA CENTRALIDAD DE LA PERSONA EN EL DEBATE MULTICULTURAL
by: Santiago Sastre Ariza
Published: (2007)
by: Santiago Sastre Ariza
Published: (2007)
PARA VER CON MEJOR LUZ. UNA APROXIMACIÓN AL TRABAJO DE LA DOGMÁTICA JURÍDICA
by: Santiago Sastre Ariza
Published: (2006)
by: Santiago Sastre Ariza
Published: (2006)
Online Peer-Tutoring: A Renewed Impetus for Autonomous English Learning
by: Luis Ignacio Herrera Bohórquez
Published: (2019)
by: Luis Ignacio Herrera Bohórquez
Published: (2019)
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
by: Chen, Xuetian, et al.
Published: (2025)
by: Chen, Xuetian, et al.
Published: (2025)
BEA Discoveries 2010: BEA beyond the Buzz
by: Fox, Bette-Lee, et al.
Published: (2010)
by: Fox, Bette-Lee, et al.
Published: (2010)
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
by: Li, Bangzheng, et al.
Published: (2023)
by: Li, Bangzheng, et al.
Published: (2023)
Can AI Agents Generate Microservices? How Far are We?
by: Adnan, Bassam, et al.
Published: (2026)
by: Adnan, Bassam, et al.
Published: (2026)
Runtime Failure Hunting for Physics Engine Based Software Systems: How Far Can We Go?
by: Li, Shuqing, et al.
Published: (2025)
by: Li, Shuqing, et al.
Published: (2025)
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
by: Lu, Yi, et al.
Published: (2025)
by: Lu, Yi, et al.
Published: (2025)
Pedagogy-driven Evaluation of Generative AI-powered Intelligent Tutoring Systems
by: Maurya, Kaushal Kumar, et al.
Published: (2025)
by: Maurya, Kaushal Kumar, et al.
Published: (2025)
À l'ombre de l'INCO
by: Dennie, Donald
Published: (2015)
by: Dennie, Donald
Published: (2015)
Valorization of logistics infrastructures using the SWOT-Delphi-CAME methodology. The case of the Albacete railway logistics platform
by: J. Ignacio Parra-Santiago
Published: (2021)
by: J. Ignacio Parra-Santiago
Published: (2021)
AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?
by: Karmakar, Ranit, et al.
Published: (2026)
by: Karmakar, Ranit, et al.
Published: (2026)
How Far Can Pretrained LLMs Go in Symbolic Music? Controlled Comparisons of Supervised and Preference-based Adaptation
by: Kumar, Deepak, et al.
Published: (2026)
by: Kumar, Deepak, et al.
Published: (2026)
Trastornos afectivos estacionales, “winter blues”.
by: Miren Aiala Gatón Moreno
Published: (2015)
by: Miren Aiala Gatón Moreno
Published: (2015)
A Bivariate Model to Predict Library Circulation.
by: Barr, Aiala, et al.
Published: (1991)
by: Barr, Aiala, et al.
Published: (1991)
Similar Items
-
RETUYT-INCO at BEA 2026 Shared Task 2: Meta-prompting in Rubric-based Scoring for German
by: Sastre, Ignacio, et al.
Published: (2026) -
A Platform for Generating Educational Activities to Teach English as a Second Language
by: Rosá, Aiala, et al.
Published: (2025) -
Memory Tokens: Large Language Models Can Generate Reversible Sentence Embeddings
by: Sastre, Ignacio, et al.
Published: (2025) -
Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions
by: Sastre, Ignacio, et al.
Published: (2026) -
Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
by: Kochmar, Ekaterina, et al.
Published: (2025)