Finnish SQuAD: A Simple Approach to Machine Translation of Span Annotations
Fuente:
arXiv
Saved in:
| Main Authors: | Nuutinen, Emil, Rastas, Iiro, Ginter, Filip |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EuSQuAD: Automatically Translated and Aligned SQuAD2.0 for Basque
by: García-Pablos, Aitor, et al.
Published: (2024)
by: García-Pablos, Aitor, et al.
Published: (2024)
Hybrid-SQuAD: Hybrid Scholarly Question Answering Dataset
by: Taffa, Tilahun Abedissa, et al.
Published: (2024)
by: Taffa, Tilahun Abedissa, et al.
Published: (2024)
The Death of Feature Engineering? BERT with Linguistic Features on SQuAD 2.0
by: Li, Jiawei, et al.
Published: (2024)
by: Li, Jiawei, et al.
Published: (2024)
emrQA-msquad: A Medical Dataset Structured with the SQuAD V2.0 Framework, Enriched with emrQA Medical Information
by: Eladio, Jimenez, et al.
Published: (2024)
by: Eladio, Jimenez, et al.
Published: (2024)
When is dataset cartography ineffective? Using training dynamics does not improve robustness against Adversarial SQuAD
by: Mandal, Paul K.
Published: (2025)
by: Mandal, Paul K.
Published: (2025)
Extracting Social Connections from Finnish Karelian Refugee Interviews Using LLMs
by: Laato, Joonatan, et al.
Published: (2025)
by: Laato, Joonatan, et al.
Published: (2025)
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
Semantic Search as Extractive Paraphrase Span Detection
by: Kanerva, Jenna, et al.
Published: (2021)
by: Kanerva, Jenna, et al.
Published: (2021)
AmaSQuAD: A Benchmark for Amharic Extractive Question Answering
by: Hailemariam, Nebiyou Daniel, et al.
Published: (2025)
by: Hailemariam, Nebiyou Daniel, et al.
Published: (2025)
Advancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
by: Pal, Aniket, et al.
Published: (2024)
by: Pal, Aniket, et al.
Published: (2024)
HeySQuAD: A Spoken Question Answering Dataset
by: Wu, Yijing, et al.
Published: (2023)
by: Wu, Yijing, et al.
Published: (2023)
FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering
by: Henriksson, Erik, et al.
Published: (2025)
by: Henriksson, Erik, et al.
Published: (2025)
MahaSQuAD: Bridging Linguistic Divides in Marathi Question-Answering
by: Ghatage, Ruturaj, et al.
Published: (2024)
by: Ghatage, Ruturaj, et al.
Published: (2024)
IndicSQuAD: A Comprehensive Multilingual Question Answering Dataset for Indic Languages
by: Endait, Sharvi, et al.
Published: (2025)
by: Endait, Sharvi, et al.
Published: (2025)
OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches
by: Kanerva, Jenna, et al.
Published: (2025)
by: Kanerva, Jenna, et al.
Published: (2025)
Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance
by: Myntti, Amanda, et al.
Published: (2026)
by: Myntti, Amanda, et al.
Published: (2026)
Span-Level Machine Translation Meta-Evaluation
by: Perrella, Stefano, et al.
Published: (2026)
by: Perrella, Stefano, et al.
Published: (2026)
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
by: Kasner, Zdeněk, et al.
Published: (2025)
by: Kasner, Zdeněk, et al.
Published: (2025)
ConGA: Guidelines for Contextual Gender Annotation. A Framework for Annotating Gender in Machine Translation
by: Rescigno, Argentina Anna, et al.
Published: (2026)
by: Rescigno, Argentina Anna, et al.
Published: (2026)
Spectroscopic Quasar Anomaly Detection (SQuAD) I: Rest-Frame UV Spectra from SDSS DR16
by: Tiwari, Arihant, et al.
Published: (2024)
by: Tiwari, Arihant, et al.
Published: (2024)
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025)
by: Riley, Parker, et al.
Published: (2025)
Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs
by: Laato, Joonatan, et al.
Published: (2026)
by: Laato, Joonatan, et al.
Published: (2026)
Large Language Models as Annotators for Machine Translation Quality Estimation
by: Wang, Sidi, et al.
Published: (2026)
by: Wang, Sidi, et al.
Published: (2026)
SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic Languages
by: Liu, Hannah, et al.
Published: (2025)
by: Liu, Hannah, et al.
Published: (2025)
To Aggregate or Not to Aggregate. That is the Question: A Case Study on Annotation Subjectivity in Span Prediction
by: Kurniawan, Kemal, et al.
Published: (2024)
by: Kurniawan, Kemal, et al.
Published: (2024)
CharSpan: Utilizing Lexical Similarity to Enable Zero-Shot Machine Translation for Extremely Low-resource Languages
by: Maurya, Kaushal Kumar, et al.
Published: (2023)
by: Maurya, Kaushal Kumar, et al.
Published: (2023)
Matching Meaning at Scale: Evaluating Semantic Search for 18th-Century Intellectual History through the Case of Locke
by: Wu, Yu, et al.
Published: (2026)
by: Wu, Yu, et al.
Published: (2026)
Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations
by: Tan, Shaomu, et al.
Published: (2025)
by: Tan, Shaomu, et al.
Published: (2025)
Enhanced Hallucination Detection in Neural Machine Translation through Simple Detector Aggregation
by: Himmi, Anas, et al.
Published: (2024)
by: Himmi, Anas, et al.
Published: (2024)
A Bayesian Optimization Approach to Machine Translation Reranking
by: Cheng, Julius, et al.
Published: (2024)
by: Cheng, Julius, et al.
Published: (2024)
Combining Qualitative and Computational Approaches for Literary Analysis of Finnish Novels
by: Ohman, Emily, et al.
Published: (2024)
by: Ohman, Emily, et al.
Published: (2024)
Can GPT-4 Identify Propaganda? Annotation and Detection of Propaganda Spans in News Articles
by: Hasanain, Maram, et al.
Published: (2024)
by: Hasanain, Maram, et al.
Published: (2024)
Compositional Translation: A Novel LLM-based Approach for Low-resource Machine Translation
by: Zebaze, Armel, et al.
Published: (2025)
by: Zebaze, Armel, et al.
Published: (2025)
PubMedCausal: A Span-Level Annotated Corpus for Causal Relation Extraction in Biomedical Text
by: Kunle-John, Ifeoluwa, et al.
Published: (2026)
by: Kunle-John, Ifeoluwa, et al.
Published: (2026)
SQuAI: Scientific Question-Answering with Multi-Agent Retrieval-Augmented Generation
by: Besrour, Ines, et al.
Published: (2025)
by: Besrour, Ines, et al.
Published: (2025)
Investigating Affect Mining Techniques for Annotation Sample Selection in the Creation of Finnish Affective Speech Corpus
by: Lahtinen, Kalle, et al.
Published: (2025)
by: Lahtinen, Kalle, et al.
Published: (2025)
Translation via Annotation: A Computational Study of Translating Classical Chinese into Japanese
by: Li, Zilong, et al.
Published: (2025)
by: Li, Zilong, et al.
Published: (2025)
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
by: Ki, Dayeon, et al.
Published: (2024)
by: Ki, Dayeon, et al.
Published: (2024)
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation
by: Lyu, Boxuan, et al.
Published: (2025)
by: Lyu, Boxuan, et al.
Published: (2025)
LATA: A Tool for LLM-Assisted Translation Annotation
by: Huang, Baorong, et al.
Published: (2026)
by: Huang, Baorong, et al.
Published: (2026)
Similar Items
-
EuSQuAD: Automatically Translated and Aligned SQuAD2.0 for Basque
by: García-Pablos, Aitor, et al.
Published: (2024) -
Hybrid-SQuAD: Hybrid Scholarly Question Answering Dataset
by: Taffa, Tilahun Abedissa, et al.
Published: (2024) -
The Death of Feature Engineering? BERT with Linguistic Features on SQuAD 2.0
by: Li, Jiawei, et al.
Published: (2024) -
emrQA-msquad: A Medical Dataset Structured with the SQuAD V2.0 Framework, Enriched with emrQA Medical Information
by: Eladio, Jimenez, et al.
Published: (2024) -
When is dataset cartography ineffective? Using training dynamics does not improve robustness against Adversarial SQuAD
by: Mandal, Paul K.
Published: (2025)