A multi-level multi-label text classification dataset of 19th century Ottoman and Russian literary and critical texts
Fuente:
arXiv
Guardado en:
| Autores principales: | Gokceoglu, Gokcen, Cavusoglu, Devrim, Akbas, Emre, Dolcerocca, Özen Nergis |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AlleNoise: large-scale text classification benchmark dataset with real-world label noise
por: Rączkowska, Alicja, et al.
Publicado: (2024)
por: Rączkowska, Alicja, et al.
Publicado: (2024)
DisGeM: Distractor Generation for Multiple Choice Questions with Span Masking
por: Cavusoglu, Devrim, et al.
Publicado: (2024)
por: Cavusoglu, Devrim, et al.
Publicado: (2024)
Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
por: Bourne, Jonathan
Publicado: (2025)
por: Bourne, Jonathan
Publicado: (2025)
KeNet:Knowledge-enhanced Doc-Label Attention Network for Multi-label text classification
por: Li, Bo, et al.
Publicado: (2024)
por: Li, Bo, et al.
Publicado: (2024)
Strategies for political-statement segmentation and labelling in unstructured text
por: Nikolaev, Dmitry, et al.
Publicado: (2025)
por: Nikolaev, Dmitry, et al.
Publicado: (2025)
Jury: A Comprehensive Evaluation Toolkit
por: Cavusoglu, Devrim, et al.
Publicado: (2023)
por: Cavusoglu, Devrim, et al.
Publicado: (2023)
Identifying social isolation themes in NVDRS text narratives using topic modeling and text-classification methods
por: Walker, Drew, et al.
Publicado: (2025)
por: Walker, Drew, et al.
Publicado: (2025)
Spider4SSC & S2CLite: A text-to-multi-query-language dataset using lightweight ontology-agnostic SPARQL to Cypher parser
por: Vejvar, Martin, et al.
Publicado: (2025)
por: Vejvar, Martin, et al.
Publicado: (2025)
Detecting value-expressive text posts in Russian social media
por: Milkova, Maria, et al.
Publicado: (2023)
por: Milkova, Maria, et al.
Publicado: (2023)
A centroid based framework for text classification in itsm environments
por: Mohanna, Hossein, et al.
Publicado: (2025)
por: Mohanna, Hossein, et al.
Publicado: (2025)
Comparing energy consumption and accuracy in text classification inference
por: Zschache, Johannes, et al.
Publicado: (2025)
por: Zschache, Johannes, et al.
Publicado: (2025)
Dependency Annotation of Ottoman Turkish with Multilingual BERT
por: Özateş, Şaziye Betül, et al.
Publicado: (2024)
por: Özateş, Şaziye Betül, et al.
Publicado: (2024)
Synthetically generated text for supervised text analysis
por: Halterman, Andrew
Publicado: (2023)
por: Halterman, Andrew
Publicado: (2023)
Domain-specific long text classification from sparse relevant information
por: D'Cruz, Célia, et al.
Publicado: (2024)
por: D'Cruz, Célia, et al.
Publicado: (2024)
A large-scale image-text dataset benchmark for farmland segmentation
por: Tao, Chao, et al.
Publicado: (2025)
por: Tao, Chao, et al.
Publicado: (2025)
Explainable machine learning multi-label classification of Spanish legal judgements
por: de Arriba-Pérez, Francisco, et al.
Publicado: (2024)
por: de Arriba-Pérez, Francisco, et al.
Publicado: (2024)
Explanation sensitivity to the randomness of large language models: the case of journalistic text classification
por: Bogaert, Jeremie, et al.
Publicado: (2024)
por: Bogaert, Jeremie, et al.
Publicado: (2024)
ProText: A benchmark dataset for measuring (mis)gendering in long-form texts
por: Kotek, Hadas, et al.
Publicado: (2026)
por: Kotek, Hadas, et al.
Publicado: (2026)
Specialized text classification: an approach to classifying Open Banking transactions
por: TA, Duc Tuyen, et al.
Publicado: (2025)
por: TA, Duc Tuyen, et al.
Publicado: (2025)
How do we measure privacy in text? A survey of text anonymization metrics
por: Ren, Yaxuan, et al.
Publicado: (2025)
por: Ren, Yaxuan, et al.
Publicado: (2025)
BeanCounter: A low-toxicity, large-scale, and open dataset of business-oriented text
por: Wang, Siyan, et al.
Publicado: (2024)
por: Wang, Siyan, et al.
Publicado: (2024)
Analysis of child development facts and myths using text mining techniques and classification models
por: Tajrian, Mehedi, et al.
Publicado: (2024)
por: Tajrian, Mehedi, et al.
Publicado: (2024)
Can human clinical rationales improve the performance and explainability of clinical text classification models?
por: Metzner, Christoph, et al.
Publicado: (2025)
por: Metzner, Christoph, et al.
Publicado: (2025)
Exploring space efficiency in a tree-based linear model for extreme multi-label classification
por: Lin, He-Zhe, et al.
Publicado: (2024)
por: Lin, He-Zhe, et al.
Publicado: (2024)
Reshaping MOFs text mining with a dynamic multi-agents framework of large language model
por: Lin, Zuhong, et al.
Publicado: (2025)
por: Lin, Zuhong, et al.
Publicado: (2025)
Scalable and consistent few-shot classification of survey responses using text embeddings
por: Mjaaland, Jonas Timmann, et al.
Publicado: (2025)
por: Mjaaland, Jonas Timmann, et al.
Publicado: (2025)
Reading Between the Lines: A dataset and a study on why some texts are tougher than others
por: Khallaf, Nouran, et al.
Publicado: (2025)
por: Khallaf, Nouran, et al.
Publicado: (2025)
Comparing effectiveness of regularization methods on text classification: Simple and complex model in data shortage situation
por: Lee, Jongga, et al.
Publicado: (2024)
por: Lee, Jongga, et al.
Publicado: (2024)
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
por: He, Jie, et al.
Publicado: (2025)
por: He, Jie, et al.
Publicado: (2025)
Detecting text level intellectual influence with knowledge graph embeddings
por: Li, Lucian, et al.
Publicado: (2024)
por: Li, Lucian, et al.
Publicado: (2024)
A thorough benchmark of automatic text classification: From traditional approaches to large language models
por: Cunha, Washington, et al.
Publicado: (2025)
por: Cunha, Washington, et al.
Publicado: (2025)
Approaches to the classification of complex systems: Words, texts, and more
por: Rovenchak, Andrij
Publicado: (2022)
por: Rovenchak, Andrij
Publicado: (2022)
TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models
por: Feng, Jiarui, et al.
Publicado: (2024)
por: Feng, Jiarui, et al.
Publicado: (2024)
Domain-specific or Uncertainty-aware models: Does it really make a difference for biomedical text classification?
por: Sinha, Aman, et al.
Publicado: (2024)
por: Sinha, Aman, et al.
Publicado: (2024)
Qwen it detect machine-generated text?
por: Marchitan, Teodor-George, et al.
Publicado: (2025)
por: Marchitan, Teodor-George, et al.
Publicado: (2025)
Identifying attributions of causality in political text
por: Garcia-Corral, Paulina
Publicado: (2025)
por: Garcia-Corral, Paulina
Publicado: (2025)
MaterioMiner -- An ontology-based text mining dataset for extraction of process-structure-property entities
por: Durmaz, Ali Riza, et al.
Publicado: (2024)
por: Durmaz, Ali Riza, et al.
Publicado: (2024)
An energy-based comparative analysis of common approaches to text classification in the Legal domain
por: Gultekin, Sinan, et al.
Publicado: (2023)
por: Gultekin, Sinan, et al.
Publicado: (2023)
LLMs can hide text in other text of the same length
por: Norelli, Antonio, et al.
Publicado: (2025)
por: Norelli, Antonio, et al.
Publicado: (2025)
Effects of term weighting approach with and without stop words removing on Arabic text classification
por: Alhenawi, Esra'a, et al.
Publicado: (2024)
por: Alhenawi, Esra'a, et al.
Publicado: (2024)
Ejemplares similares
-
AlleNoise: large-scale text classification benchmark dataset with real-world label noise
por: Rączkowska, Alicja, et al.
Publicado: (2024) -
DisGeM: Distractor Generation for Multiple Choice Questions with Span Masking
por: Cavusoglu, Devrim, et al.
Publicado: (2024) -
Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
por: Bourne, Jonathan
Publicado: (2025) -
KeNet:Knowledge-enhanced Doc-Label Attention Network for Multi-label text classification
por: Li, Bo, et al.
Publicado: (2024) -
Strategies for political-statement segmentation and labelling in unstructured text
por: Nikolaev, Dmitry, et al.
Publicado: (2025)