To Err Is Human, but Llamas Can Learn It Too
Fuente:
arXiv
Saved in:
| Main Authors: | Luhtaru, Agnes, Purason, Taido, Vainikko, Martin, Del, Maksym, Fishel, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Teaching Llama a New Language Through Cross-Lingual Knowledge Transfer
by: Kuulmets, Hele-Andra, et al.
Published: (2024)
by: Kuulmets, Hele-Andra, et al.
Published: (2024)
LLMs for Extremely Low-Resource Finno-Ugric Languages
by: Purason, Taido, et al.
Published: (2024)
by: Purason, Taido, et al.
Published: (2024)
Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models
by: Purason, Taido, et al.
Published: (2025)
by: Purason, Taido, et al.
Published: (2025)
Autocorrect for Estonian texts: final report from project EKTB25
by: Luhtaru, Agnes, et al.
Published: (2024)
by: Luhtaru, Agnes, et al.
Published: (2024)
Prune or Retrain: Optimizing the Vocabulary of Multilingual Models for Estonian
by: Dorkin, Aleksei, et al.
Published: (2025)
by: Dorkin, Aleksei, et al.
Published: (2025)
EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training
by: Dorkin, Aleksei, et al.
Published: (2026)
by: Dorkin, Aleksei, et al.
Published: (2026)
How Uncertainty Estimation Scales with Sampling in Reasoning Models
by: Del, Maksym, et al.
Published: (2026)
by: Del, Maksym, et al.
Published: (2026)
To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation
by: Cheng, Xiang, et al.
Published: (2024)
by: Cheng, Xiang, et al.
Published: (2024)
VariErr NLI: Separating Annotation Error from Human Label Variation
by: Weber-Genzel, Leon, et al.
Published: (2024)
by: Weber-Genzel, Leon, et al.
Published: (2024)
Too Helpful, Too Harmless, Too Honest or Just Right?
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
by: Bianchi, Federico, et al.
Published: (2025)
by: Bianchi, Federico, et al.
Published: (2025)
ChocoLlama: Lessons Learned From Teaching Llamas Dutch
by: Meeus, Matthieu, et al.
Published: (2024)
by: Meeus, Matthieu, et al.
Published: (2024)
Limited Linguistic Diversity in Embodied AI Datasets
by: Wanna, Selma, et al.
Published: (2026)
by: Wanna, Selma, et al.
Published: (2026)
Quantity vs. Quality of Monolingual Source Data in Automatic Text Translation: Can It Be Too Little If It Is Too Good?
by: Abdulmumin, Idris, et al.
Published: (2024)
by: Abdulmumin, Idris, et al.
Published: (2024)
Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning Models
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Fearful Falcons and Angry Llamas: Emotion Category Annotations of Arguments by Humans and LLMs
by: Greschner, Lynn, et al.
Published: (2024)
by: Greschner, Lynn, et al.
Published: (2024)
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
by: Niu, Jingcheng, et al.
Published: (2025)
by: Niu, Jingcheng, et al.
Published: (2025)
MedErrBench: A Fine-Grained Multilingual Benchmark for Medical Error Detection and Correction with Clinical Expert Annotations
by: Ma, Congbo, et al.
Published: (2026)
by: Ma, Congbo, et al.
Published: (2026)
BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B
by: Gade, Pranav, et al.
Published: (2023)
by: Gade, Pranav, et al.
Published: (2023)
Integral Transformer: Denoising Attention, Not Too Much Not Too Little
by: Kobyzev, Ivan, et al.
Published: (2025)
by: Kobyzev, Ivan, et al.
Published: (2025)
Llama-Mob: Instruction-Tuning Llama-3-8B Excels in City-Scale Mobility Prediction
by: Tang, Peizhi, et al.
Published: (2024)
by: Tang, Peizhi, et al.
Published: (2024)
Code Llama: Open Foundation Models for Code
by: Rozière, Baptiste, et al.
Published: (2023)
by: Rozière, Baptiste, et al.
Published: (2023)
Feedback Indicators: The Alignment between Llama and a Teacher in Language Learning
by: Rüdian, Sylvio, et al.
Published: (2025)
by: Rüdian, Sylvio, et al.
Published: (2025)
MGH Radiology Llama: A Llama 3 70B Model for Radiology
by: Shi, Yucheng, et al.
Published: (2024)
by: Shi, Yucheng, et al.
Published: (2024)
Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
by: Chi, Jianfeng, et al.
Published: (2024)
by: Chi, Jianfeng, et al.
Published: (2024)
How Do Humans Write Code? Large Models Do It the Same Way Too
by: Li, Long, et al.
Published: (2024)
by: Li, Long, et al.
Published: (2024)
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
by: Lawrence, Logan, et al.
Published: (2025)
by: Lawrence, Logan, et al.
Published: (2025)
CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law
by: Zhao, Ethan, et al.
Published: (2026)
by: Zhao, Ethan, et al.
Published: (2026)
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
by: He, Zhengfu, et al.
Published: (2024)
by: He, Zhengfu, et al.
Published: (2024)
Extending Llama-3's Context Ten-Fold Overnight
by: Zhang, Peitian, et al.
Published: (2024)
by: Zhang, Peitian, et al.
Published: (2024)
SinLlama -- A Large Language Model for Sinhala
by: Aravinda, H. W. K., et al.
Published: (2025)
by: Aravinda, H. W. K., et al.
Published: (2025)
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
by: Tang, Bangsheng, et al.
Published: (2025)
by: Tang, Bangsheng, et al.
Published: (2025)
Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast
by: Shi, Chufan, et al.
Published: (2024)
by: Shi, Chufan, et al.
Published: (2024)
The Llama 3 Herd of Models
by: Grattafiori, Aaron, et al.
Published: (2024)
by: Grattafiori, Aaron, et al.
Published: (2024)
Enabling High-Sparsity Foundational Llama Models with Efficient Pretraining and Deployment
by: Agarwalla, Abhinav, et al.
Published: (2024)
by: Agarwalla, Abhinav, et al.
Published: (2024)
MobiLlama: Towards Accurate and Lightweight Fully Transparent GPT
by: Thawakar, Omkar, et al.
Published: (2024)
by: Thawakar, Omkar, et al.
Published: (2024)
Llama-Mimi: Exploring the Limits of Flattened Speech Language Modeling
by: Sugiura, Issa, et al.
Published: (2025)
by: Sugiura, Issa, et al.
Published: (2025)
Too Late to Train, Too Early To Use? A Study on Necessity and Viability of Low-Resource Bengali LLMs
by: Mahfuz, Tamzeed, et al.
Published: (2024)
by: Mahfuz, Tamzeed, et al.
Published: (2024)
Do Llamas Work in English? On the Latent Language of Multilingual Transformers
by: Wendler, Chris, et al.
Published: (2024)
by: Wendler, Chris, et al.
Published: (2024)
Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts
by: Chochlakis, Georgios, et al.
Published: (2025)
by: Chochlakis, Georgios, et al.
Published: (2025)
Similar Items
-
Teaching Llama a New Language Through Cross-Lingual Knowledge Transfer
by: Kuulmets, Hele-Andra, et al.
Published: (2024) -
LLMs for Extremely Low-Resource Finno-Ugric Languages
by: Purason, Taido, et al.
Published: (2024) -
Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models
by: Purason, Taido, et al.
Published: (2025) -
Autocorrect for Estonian texts: final report from project EKTB25
by: Luhtaru, Agnes, et al.
Published: (2024) -
Prune or Retrain: Optimizing the Vocabulary of Multilingual Models for Estonian
by: Dorkin, Aleksei, et al.
Published: (2025)