Curation of a Palaeohispanic Dataset for Machine Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Martínez-Fernández, Gonzalo, Quesada, Jose F, Riscos-Núñez, Agustín, Salguero-Lamillar, Francisco José |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AI Thinking as a Meaning-Centered Framework: Reimagining Language Technologies Through Community Agency
por: Quesada, Jose F
Publicado: (2025)
por: Quesada, Jose F
Publicado: (2025)
Machine-Assisted Script Curation
por: Ciosici, Manuel R., et al.
Publicado: (2021)
por: Ciosici, Manuel R., et al.
Publicado: (2021)
SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
por: Chopra, Muskaan, et al.
Publicado: (2025)
por: Chopra, Muskaan, et al.
Publicado: (2025)
LLM Unlearning Without an Expert Curated Dataset
por: Zhu, Xiaoyuan, et al.
Publicado: (2025)
por: Zhu, Xiaoyuan, et al.
Publicado: (2025)
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
por: Djuhera, Aladin, et al.
Publicado: (2025)
por: Djuhera, Aladin, et al.
Publicado: (2025)
Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
por: Rachamalla, Neel Prabhanjan, et al.
Publicado: (2025)
por: Rachamalla, Neel Prabhanjan, et al.
Publicado: (2025)
ORBIT: Cost-Effective Dataset Curation for Large Language Model Domain Adaptation with an Astronomy Case Study
por: Modesitt, Eric, et al.
Publicado: (2024)
por: Modesitt, Eric, et al.
Publicado: (2024)
SkillOS: Learning Skill Curation for Self-Evolving Agents
por: Ouyang, Siru, et al.
Publicado: (2026)
por: Ouyang, Siru, et al.
Publicado: (2026)
Machine Learning-based NLP for Emotion Classification on a Cholera X Dataset
por: Jideani, Paul, et al.
Publicado: (2024)
por: Jideani, Paul, et al.
Publicado: (2024)
Code2Doc: A Quality-First Curated Dataset for Code Documentation
por: Karaman, Recep Kaan, et al.
Publicado: (2025)
por: Karaman, Recep Kaan, et al.
Publicado: (2025)
ExpertQA: Expert-Curated Questions and Attributed Answers
por: Malaviya, Chaitanya, et al.
Publicado: (2023)
por: Malaviya, Chaitanya, et al.
Publicado: (2023)
The Role of Data Curation in Image Captioning
por: Li, Wenyan, et al.
Publicado: (2023)
por: Li, Wenyan, et al.
Publicado: (2023)
Preference Consistency Matters: Enhancing Preference Learning in Language Models with Automated Self-Curation of Training Corpora
por: Lee, JoonHo, et al.
Publicado: (2024)
por: Lee, JoonHo, et al.
Publicado: (2024)
Bridging the Gap in Bangla Healthcare: Machine Learning Based Disease Prediction Using a Symptoms-Disease Dataset
por: Zannat, Rowzatul, et al.
Publicado: (2026)
por: Zannat, Rowzatul, et al.
Publicado: (2026)
SciCode: A Research Coding Benchmark Curated by Scientists
por: Tian, Minyang, et al.
Publicado: (2024)
por: Tian, Minyang, et al.
Publicado: (2024)
MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following
por: Lou, Renze, et al.
Publicado: (2023)
por: Lou, Renze, et al.
Publicado: (2023)
A Survey of Machine Learning Models and Datasets for the Multi-label Classification of Textual Hate Speech in English
por: Bäumler, Julian, et al.
Publicado: (2025)
por: Bäumler, Julian, et al.
Publicado: (2025)
Improving Data Efficiency via Curating LLM-Driven Rating Systems
por: Pang, Jinlong, et al.
Publicado: (2024)
por: Pang, Jinlong, et al.
Publicado: (2024)
From Hallucinations to Facts: Enhancing Language Models with Curated Knowledge Graphs
por: Joshi, Ratnesh Kumar, et al.
Publicado: (2024)
por: Joshi, Ratnesh Kumar, et al.
Publicado: (2024)
Efficient Continual Learning in Neural Machine Translation: A Low-Rank Adaptation Approach
por: Carrión, Salvador, et al.
Publicado: (2025)
por: Carrión, Salvador, et al.
Publicado: (2025)
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
por: Pombal, José, et al.
Publicado: (2025)
por: Pombal, José, et al.
Publicado: (2025)
Playing with words: Comparing the vocabulary and lexical diversity of ChatGPT and humans
por: Reviriego, Pedro, et al.
Publicado: (2023)
por: Reviriego, Pedro, et al.
Publicado: (2023)
UrduLLaMA 1.0: Dataset Curation, Preprocessing, and Evaluation in Low-Resource Settings
por: Fiaz, Layba, et al.
Publicado: (2025)
por: Fiaz, Layba, et al.
Publicado: (2025)
Beyond Agreement: Scoring Panel-Surfaced Biomedical Entity Candidates for Curator Triage
por: Cao, Shuheng, et al.
Publicado: (2026)
por: Cao, Shuheng, et al.
Publicado: (2026)
Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning
por: Ming, Xiaoyang, et al.
Publicado: (2026)
por: Ming, Xiaoyang, et al.
Publicado: (2026)
Explainable cognitive decline detection in free dialogues with a Machine Learning approach based on pre-trained Large Language Models
por: de Arriba-Pérez, Francisco, et al.
Publicado: (2024)
por: de Arriba-Pérez, Francisco, et al.
Publicado: (2024)
Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
por: Lupidi, Alisia, et al.
Publicado: (2024)
por: Lupidi, Alisia, et al.
Publicado: (2024)
Interpretability of the Intent Detection Problem: A New Approach
por: Sanchez-Karhunen, Eduardo, et al.
Publicado: (2026)
por: Sanchez-Karhunen, Eduardo, et al.
Publicado: (2026)
Can ChatGPT Learn to Count Letters?
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
CUICurate: A GraphRAG-based Framework for Automated Clinical Concept Curation for NLP applications
por: Blake, Victoria, et al.
Publicado: (2026)
por: Blake, Victoria, et al.
Publicado: (2026)
Intellecta Cognitiva: A Comprehensive Dataset for Advancing Academic Knowledge and Machine Reasoning
por: PS, Ajmal, et al.
Publicado: (2024)
por: PS, Ajmal, et al.
Publicado: (2024)
Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
por: Ferrando, Raquel, et al.
Publicado: (2025)
por: Ferrando, Raquel, et al.
Publicado: (2025)
Evidence-Grounded Subspecialty Reasoning: Evaluating a Curated Clinical Intelligence Layer on the 2025 Endocrinology Board-Style Examination
por: Hosseinian, Amir, et al.
Publicado: (2026)
por: Hosseinian, Amir, et al.
Publicado: (2026)
TECCI: Tricky Edits of Collected and Curated Images
por: Agrawal, Aishwarya, et al.
Publicado: (2026)
por: Agrawal, Aishwarya, et al.
Publicado: (2026)
FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity
por: Jourdan, Fanny, et al.
Publicado: (2025)
por: Jourdan, Fanny, et al.
Publicado: (2025)
Transfer Learning for Automated Feedback Generation on Small Datasets
por: Morris, Oscar
Publicado: (2025)
por: Morris, Oscar
Publicado: (2025)
Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service
por: Merx, Raphael, et al.
Publicado: (2024)
por: Merx, Raphael, et al.
Publicado: (2024)
SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification
por: Posada, Jose D., et al.
Publicado: (2026)
por: Posada, Jose D., et al.
Publicado: (2026)
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties
por: Yu, Keunwoo Peter, et al.
Publicado: (2023)
por: Yu, Keunwoo Peter, et al.
Publicado: (2023)
Exploring the Potential of Machine Translation for Generating Named Entity Datasets: A Case Study between Persian and English
por: Sartipi, Amir, et al.
Publicado: (2023)
por: Sartipi, Amir, et al.
Publicado: (2023)
Ejemplares similares
-
AI Thinking as a Meaning-Centered Framework: Reimagining Language Technologies Through Community Agency
por: Quesada, Jose F
Publicado: (2025) -
Machine-Assisted Script Curation
por: Ciosici, Manuel R., et al.
Publicado: (2021) -
SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
por: Chopra, Muskaan, et al.
Publicado: (2025) -
LLM Unlearning Without an Expert Curated Dataset
por: Zhu, Xiaoyuan, et al.
Publicado: (2025) -
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
por: Djuhera, Aladin, et al.
Publicado: (2025)