Training data generation for context-dependent rubric-based short answer grading
Fuente:
arXiv
Guardado en:
| Autores principales: | Šindelář, Pavel, Slivka, Dávid, Bouma, Christopher, Prášil, Filip, Bojar, Ondřej |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
por: Šindelář, Pavel, et al.
Publicado: (2025)
por: Šindelář, Pavel, et al.
Publicado: (2025)
Finetuning LLMs for EvaCun 2025 token prediction shared task
por: Jon, Josef, et al.
Publicado: (2025)
por: Jon, Josef, et al.
Publicado: (2025)
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
por: Barančíková, Petra, et al.
Publicado: (2025)
por: Barančíková, Petra, et al.
Publicado: (2025)
Understanding the role of FFNs in driving multilingual behaviour in LLMs
por: Bhattacharya, Sunit, et al.
Publicado: (2024)
por: Bhattacharya, Sunit, et al.
Publicado: (2024)
Quality and Quantity of Machine Translation References for Automatic Metrics
por: Zouhar, Vilém, et al.
Publicado: (2024)
por: Zouhar, Vilém, et al.
Publicado: (2024)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
por: Luu, Nam, et al.
Publicado: (2025)
por: Luu, Nam, et al.
Publicado: (2025)
Continuous Rating as Reliable Human Evaluation of Simultaneous Speech Translation
por: Javorský, Dávid, et al.
Publicado: (2022)
por: Javorský, Dávid, et al.
Publicado: (2022)
Prompting LLMs: Length Control for Isometric Machine Translation
por: Javorský, Dávid, et al.
Publicado: (2025)
por: Javorský, Dávid, et al.
Publicado: (2025)
MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines
por: Javorský, Dávid, et al.
Publicado: (2025)
por: Javorský, Dávid, et al.
Publicado: (2025)
ChatGPT for automated grading of short answer questions in mechanical ventilation
por: Jade, Tejas, et al.
Publicado: (2025)
por: Jade, Tejas, et al.
Publicado: (2025)
ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
por: Stankov, Vladislav, et al.
Publicado: (2025)
por: Stankov, Vladislav, et al.
Publicado: (2025)
Multimodal Shannon Game with Images
por: Zouhar, Vilém, et al.
Publicado: (2023)
por: Zouhar, Vilém, et al.
Publicado: (2023)
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
por: Polák, Peter, et al.
Publicado: (2023)
por: Polák, Peter, et al.
Publicado: (2023)
Evaluating Optimal Reference Translations
por: Zouhar, Vilém, et al.
Publicado: (2023)
por: Zouhar, Vilém, et al.
Publicado: (2023)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
por: Polák, Peter, et al.
Publicado: (2025)
por: Polák, Peter, et al.
Publicado: (2025)
Corpus of Cross-lingual Dialogues with Minutes and Detection of Misunderstandings
por: Čechovič, Marko, et al.
Publicado: (2025)
por: Čechovič, Marko, et al.
Publicado: (2025)
Findings of the Third Automatic Minuting (AutoMin) Challenge
por: Shinde, Kartik, et al.
Publicado: (2025)
por: Shinde, Kartik, et al.
Publicado: (2025)
Ratas framework: A comprehensive genai-based approach to rubric-based marking of real-world textual exams
por: Safilian, Masoud, et al.
Publicado: (2025)
por: Safilian, Masoud, et al.
Publicado: (2025)
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
por: Papi, Sara, et al.
Publicado: (2024)
por: Papi, Sara, et al.
Publicado: (2024)
FusionMind -- Improving question and answering with external context fusion
por: Verma, Shreyas, et al.
Publicado: (2023)
por: Verma, Shreyas, et al.
Publicado: (2023)
ConSens: Assessing context grounding in open-book question answering
por: Vankov, Ivan, et al.
Publicado: (2025)
por: Vankov, Ivan, et al.
Publicado: (2025)
Czech Dataset for Complex Aspect-Based Sentiment Analysis Tasks
por: Šmíd, Jakub, et al.
Publicado: (2025)
por: Šmíd, Jakub, et al.
Publicado: (2025)
CLIPPER: Compression enables long-context synthetic data generation
por: Pham, Chau Minh, et al.
Publicado: (2025)
por: Pham, Chau Minh, et al.
Publicado: (2025)
Extract, Match, and Score: An Evaluation Paradigm for Long Question-context-answer Triplets in Financial Analysis
por: Hu, Bo, et al.
Publicado: (2025)
por: Hu, Bo, et al.
Publicado: (2025)
Retrieval augmented text-to-SQL generation for epidemiological question answering using electronic health records
por: Ziletti, Angelo, et al.
Publicado: (2024)
por: Ziletti, Angelo, et al.
Publicado: (2024)
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
por: Sil, Pritam, et al.
Publicado: (2026)
por: Sil, Pritam, et al.
Publicado: (2026)
PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
por: Batorski, Pawel, et al.
Publicado: (2025)
por: Batorski, Pawel, et al.
Publicado: (2025)
LLM Compression: How Far Can We Go in Balancing Size and Performance?
por: Sk, Sahil, et al.
Publicado: (2025)
por: Sk, Sahil, et al.
Publicado: (2025)
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
por: Kamzela, Wiktor, et al.
Publicado: (2025)
por: Kamzela, Wiktor, et al.
Publicado: (2025)
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
por: Sperber, Matthias, et al.
Publicado: (2024)
por: Sperber, Matthias, et al.
Publicado: (2024)
From text to multimodal: a survey of adversarial example generation in question answering systems
por: Yigit, Gulsum, et al.
Publicado: (2023)
por: Yigit, Gulsum, et al.
Publicado: (2023)
Enhancing textual textbook question answering with large language models and retrieval augmented generation
por: Alawwad, Hessa Abdulrahman, et al.
Publicado: (2024)
por: Alawwad, Hessa Abdulrahman, et al.
Publicado: (2024)
CMRAG: Co-modality-based visual document retrieval and question answering
por: Chen, Wang, et al.
Publicado: (2025)
por: Chen, Wang, et al.
Publicado: (2025)
Exploring Multiple Strategies to Improve Multilingual Coreference Resolution in CorefUD
por: Pražák, Ondřej, et al.
Publicado: (2024)
por: Pražák, Ondřej, et al.
Publicado: (2024)
factgenie: A Framework for Span-based Evaluation of Generated Texts
por: Kasner, Zdeněk, et al.
Publicado: (2024)
por: Kasner, Zdeněk, et al.
Publicado: (2024)
MEEDAV: A Synchronous Web Viewer for EEG, Eye-Tracking and Speech Data
por: Pijálek, Jan, et al.
Publicado: (2026)
por: Pijálek, Jan, et al.
Publicado: (2026)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
por: Maar, Jim, et al.
Publicado: (2026)
por: Maar, Jim, et al.
Publicado: (2026)
LADM: Long-context Training Data Selection with Attention-based Dependency Measurement for LLMs
por: Chen, Jianghao, et al.
Publicado: (2025)
por: Chen, Jianghao, et al.
Publicado: (2025)
TANQ: An open domain dataset of table answered questions
por: Akhtar, Mubashara, et al.
Publicado: (2024)
por: Akhtar, Mubashara, et al.
Publicado: (2024)
A dependently-typed calculus of event telicity and culminativity
por: Kovalev, Pavel, et al.
Publicado: (2025)
por: Kovalev, Pavel, et al.
Publicado: (2025)
Ejemplares similares
-
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
por: Šindelář, Pavel, et al.
Publicado: (2025) -
Finetuning LLMs for EvaCun 2025 token prediction shared task
por: Jon, Josef, et al.
Publicado: (2025) -
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
por: Barančíková, Petra, et al.
Publicado: (2025) -
Understanding the role of FFNs in driving multilingual behaviour in LLMs
por: Bhattacharya, Sunit, et al.
Publicado: (2024) -
Quality and Quantity of Machine Translation References for Automatic Metrics
por: Zouhar, Vilém, et al.
Publicado: (2024)