PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin
Fuente:
arXiv
Salvato in:
| Autori principali: | Bothwell, Stephen, DuSell, Brian, Chiang, David, Krostenko, Brian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
di: Collado-Montañez, Jaime, et al.
Pubblicazione: (2025)
di: Collado-Montañez, Jaime, et al.
Pubblicazione: (2025)
Dialect Matters: Cross-Lingual ASR Transfer for Low-Resource Indic Language Varieties
di: Dhasmana, Akriti, et al.
Pubblicazione: (2026)
di: Dhasmana, Akriti, et al.
Pubblicazione: (2026)
Graphemic Normalization of the Perso-Arabic Script
di: Doctor, Raiomond, et al.
Pubblicazione: (2022)
di: Doctor, Raiomond, et al.
Pubblicazione: (2022)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
di: Gutkin, Alexander, et al.
Pubblicazione: (2023)
di: Gutkin, Alexander, et al.
Pubblicazione: (2023)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
Linguistic Interpretability of Transformer-based Language Models: a systematic review
di: López-Otal, Miguel, et al.
Pubblicazione: (2025)
di: López-Otal, Miguel, et al.
Pubblicazione: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
di: Smădu, Răzvan-Alexandru, et al.
Pubblicazione: (2025)
di: Smădu, Răzvan-Alexandru, et al.
Pubblicazione: (2025)
Synthetic Voice Data for Automatic Speech Recognition in African Languages
di: DeRenzi, Brian, et al.
Pubblicazione: (2025)
di: DeRenzi, Brian, et al.
Pubblicazione: (2025)
Lisbon Computational Linguists at SemEval-2024 Task 2: Using A Mistral 7B Model and Data Augmentation
di: Guimarães, Artur, et al.
Pubblicazione: (2024)
di: Guimarães, Artur, et al.
Pubblicazione: (2024)
Historical Ink: 19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction
di: Manrique-Gómez, Laura, et al.
Pubblicazione: (2024)
di: Manrique-Gómez, Laura, et al.
Pubblicazione: (2024)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
di: Costa-jussà, Marta R., et al.
Pubblicazione: (2024)
di: Costa-jussà, Marta R., et al.
Pubblicazione: (2024)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
di: Toker, Michael, et al.
Pubblicazione: (2024)
di: Toker, Michael, et al.
Pubblicazione: (2024)
ML-Promise: A Multilingual Dataset for Corporate Promise Verification
di: Seki, Yohei, et al.
Pubblicazione: (2024)
di: Seki, Yohei, et al.
Pubblicazione: (2024)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
di: Bouchekif, Abdessalam, et al.
Pubblicazione: (2026)
di: Bouchekif, Abdessalam, et al.
Pubblicazione: (2026)
EMO-KNOW: A Large Scale Dataset on Emotion and Emotion-cause
di: Nguyen, Mia Huong, et al.
Pubblicazione: (2024)
di: Nguyen, Mia Huong, et al.
Pubblicazione: (2024)
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis
di: Bose, Joy
Pubblicazione: (2026)
di: Bose, Joy
Pubblicazione: (2026)
Locations of Characters in Narratives: Andersen and Persuasion Datasets
di: Ozyurt, Batuhan, et al.
Pubblicazione: (2025)
di: Ozyurt, Batuhan, et al.
Pubblicazione: (2025)
Tracking Semantic Change in Slovene: A Novel Dataset and Optimal Transport-Based Distance
di: Pranjić, Marko, et al.
Pubblicazione: (2024)
di: Pranjić, Marko, et al.
Pubblicazione: (2024)
DimStance: Multilingual Datasets for Dimensional Stance Analysis
di: Becker, Jonas, et al.
Pubblicazione: (2026)
di: Becker, Jonas, et al.
Pubblicazione: (2026)
I run as fast as a rabbit, can you? A Multilingual Simile Dialogue Dataset
di: Ma, Longxuan, et al.
Pubblicazione: (2023)
di: Ma, Longxuan, et al.
Pubblicazione: (2023)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2023)
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2023)
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
di: Sileo, Damien
Pubblicazione: (2024)
di: Sileo, Damien
Pubblicazione: (2024)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
di: Costa-jussà, Marta R., et al.
Pubblicazione: (2024)
di: Costa-jussà, Marta R., et al.
Pubblicazione: (2024)
DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
di: Lee, Lung-Hao, et al.
Pubblicazione: (2026)
di: Lee, Lung-Hao, et al.
Pubblicazione: (2026)
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
di: Lequeu, Pierre-Antoine, et al.
Pubblicazione: (2026)
di: Lequeu, Pierre-Antoine, et al.
Pubblicazione: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
A Stochastic Analysis of the Linguistic Provenance of English Place Names
di: Dalvean, Michael
Pubblicazione: (2023)
di: Dalvean, Michael
Pubblicazione: (2023)
Double Triangle Annotation: A Scalable Human-in-the-Loop Framework for High-Precision Historical Document Annotation
di: Ren, Yi
Pubblicazione: (2026)
di: Ren, Yi
Pubblicazione: (2026)
CR-LT-KGQA: A Knowledge Graph Question Answering Dataset Requiring Commonsense Reasoning and Long-Tail Knowledge
di: Guo, Willis, et al.
Pubblicazione: (2024)
di: Guo, Willis, et al.
Pubblicazione: (2024)
Blocks Architecture (BloArk): Efficient, Cost-Effective, and Incremental Dataset Architecture for Wikipedia Revision History
di: Li, Lingxi, et al.
Pubblicazione: (2024)
di: Li, Lingxi, et al.
Pubblicazione: (2024)
How Do Large Language Models Acquire Factual Knowledge During Pretraining?
di: Chang, Hoyeon, et al.
Pubblicazione: (2024)
di: Chang, Hoyeon, et al.
Pubblicazione: (2024)
Integrating Emotional and Linguistic Models for Ethical Compliance in Large Language Models
di: Chang, Edward Y.
Pubblicazione: (2024)
di: Chang, Edward Y.
Pubblicazione: (2024)
Linguistically-Informed Multilingual Instruction Tuning: Is There an Optimal Set of Languages to Tune?
di: Soykan, Gürkan, et al.
Pubblicazione: (2024)
di: Soykan, Gürkan, et al.
Pubblicazione: (2024)
Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus
di: Alzetta, Chiara, et al.
Pubblicazione: (2025)
di: Alzetta, Chiara, et al.
Pubblicazione: (2025)
Encoder-Decoder Framework for Interactive Free Verses with Generation with Controllable High-Quality Rhyming
di: Pasini, Tommaso, et al.
Pubblicazione: (2024)
di: Pasini, Tommaso, et al.
Pubblicazione: (2024)
MIMIC-SR-ICD11: A Dataset for Narrative-Based Diagnosis
di: Wu, Yuexin, et al.
Pubblicazione: (2025)
di: Wu, Yuexin, et al.
Pubblicazione: (2025)
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
di: Nieth, Björn, et al.
Pubblicazione: (2026)
di: Nieth, Björn, et al.
Pubblicazione: (2026)
A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
di: Peláez-González, Carlos, et al.
Pubblicazione: (2025)
di: Peláez-González, Carlos, et al.
Pubblicazione: (2025)
Building and Aligning Comparable Corpora
di: Saad, Motaz, et al.
Pubblicazione: (2025)
di: Saad, Motaz, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
di: Collado-Montañez, Jaime, et al.
Pubblicazione: (2025) -
Dialect Matters: Cross-Lingual ASR Transfer for Low-Resource Indic Language Varieties
di: Dhasmana, Akriti, et al.
Pubblicazione: (2026) -
Graphemic Normalization of the Perso-Arabic Script
di: Doctor, Raiomond, et al.
Pubblicazione: (2022) -
Beyond Arabic: Software for Perso-Arabic Script Manipulation
di: Gutkin, Alexander, et al.
Pubblicazione: (2023) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)