BeanCounter: A low-toxicity, large-scale, and open dataset of business-oriented text
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Siyan, Levy, Bradford |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A large-scale image-text dataset benchmark for farmland segmentation
di: Tao, Chao, et al.
Pubblicazione: (2025)
di: Tao, Chao, et al.
Pubblicazione: (2025)
AlleNoise: large-scale text classification benchmark dataset with real-world label noise
di: Rączkowska, Alicja, et al.
Pubblicazione: (2024)
di: Rączkowska, Alicja, et al.
Pubblicazione: (2024)
Divergence Decoding: Inference-Time Unlearning via Auxiliary Models
di: Merchant, Humzah, et al.
Pubblicazione: (2026)
di: Merchant, Humzah, et al.
Pubblicazione: (2026)
A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs
di: Merchant, Humzah, et al.
Pubblicazione: (2025)
di: Merchant, Humzah, et al.
Pubblicazione: (2025)
Mapping the Web of Science, a large-scale graph and text-based dataset with LLM embeddings
di: Kunt, Tim, et al.
Pubblicazione: (2026)
di: Kunt, Tim, et al.
Pubblicazione: (2026)
GLeMM: A large-scale multilingual dataset for morphological research
di: Nabil, Hathout, et al.
Pubblicazione: (2026)
di: Nabil, Hathout, et al.
Pubblicazione: (2026)
Private prediction for large-scale synthetic text generation
di: Amin, Kareem, et al.
Pubblicazione: (2024)
di: Amin, Kareem, et al.
Pubblicazione: (2024)
TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models
di: Feng, Jiarui, et al.
Pubblicazione: (2024)
di: Feng, Jiarui, et al.
Pubblicazione: (2024)
CrowdCounter: A benchmark type-specific multi-target counterspeech dataset
di: Saha, Punyajoy, et al.
Pubblicazione: (2024)
di: Saha, Punyajoy, et al.
Pubblicazione: (2024)
Designing large language model prompts to extract scores from messy text: A shared dataset and challenge
di: Thelwall, Mike
Pubblicazione: (2026)
di: Thelwall, Mike
Pubblicazione: (2026)
CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models
di: Wang, Liangdong, et al.
Pubblicazione: (2024)
di: Wang, Liangdong, et al.
Pubblicazione: (2024)
ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining
di: Malik, Haq Nawaz
Pubblicazione: (2026)
di: Malik, Haq Nawaz
Pubblicazione: (2026)
A multi-level multi-label text classification dataset of 19th century Ottoman and Russian literary and critical texts
di: Gokceoglu, Gokcen, et al.
Pubblicazione: (2024)
di: Gokceoglu, Gokcen, et al.
Pubblicazione: (2024)
ProText: A benchmark dataset for measuring (mis)gendering in long-form texts
di: Kotek, Hadas, et al.
Pubblicazione: (2026)
di: Kotek, Hadas, et al.
Pubblicazione: (2026)
TANQ: An open domain dataset of table answered questions
di: Akhtar, Mubashara, et al.
Pubblicazione: (2024)
di: Akhtar, Mubashara, et al.
Pubblicazione: (2024)
Reading Between the Lines: A dataset and a study on why some texts are tougher than others
di: Khallaf, Nouran, et al.
Pubblicazione: (2025)
di: Khallaf, Nouran, et al.
Pubblicazione: (2025)
A safety realignment framework via subspace-oriented model fusion for large language models
di: Yi, Xin, et al.
Pubblicazione: (2024)
di: Yi, Xin, et al.
Pubblicazione: (2024)
Figurative Archive: an open dataset and web-based application for the study of metaphor
di: Bressler, Maddalena, et al.
Pubblicazione: (2025)
di: Bressler, Maddalena, et al.
Pubblicazione: (2025)
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
di: He, Jie, et al.
Pubblicazione: (2025)
di: He, Jie, et al.
Pubblicazione: (2025)
Prompting open-source and commercial language models for grammatical error correction of English learner text
di: Davis, Christopher, et al.
Pubblicazione: (2024)
di: Davis, Christopher, et al.
Pubblicazione: (2024)
Explanation sensitivity to the randomness of large language models: the case of journalistic text classification
di: Bogaert, Jeremie, et al.
Pubblicazione: (2024)
di: Bogaert, Jeremie, et al.
Pubblicazione: (2024)
600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script
di: Malik, Haq Nawaz
Pubblicazione: (2026)
di: Malik, Haq Nawaz
Pubblicazione: (2026)
OpenStaxQA: A multilingual dataset based on open-source college textbooks
di: Gupta, Pranav
Pubblicazione: (2025)
di: Gupta, Pranav
Pubblicazione: (2025)
The signal is the ceiling: Measurement limits of LLM-predicted experience ratings from open-ended survey text
di: Hong, Andrew, et al.
Pubblicazione: (2026)
di: Hong, Andrew, et al.
Pubblicazione: (2026)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
di: Li, Kunning, et al.
Pubblicazione: (2025)
di: Li, Kunning, et al.
Pubblicazione: (2025)
A scale of conceptual orality and literacy: Automatic text categorization in the tradition of "Nähe und Distanz"
di: Emmrich, Volker
Pubblicazione: (2025)
di: Emmrich, Volker
Pubblicazione: (2025)
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
di: Pletenev, Sergey, et al.
Pubblicazione: (2025)
di: Pletenev, Sergey, et al.
Pubblicazione: (2025)
ARC-Encoder: learning compressed text representations for large language models
di: Pilchen, Hippolyte, et al.
Pubblicazione: (2025)
di: Pilchen, Hippolyte, et al.
Pubblicazione: (2025)
MaterioMiner -- An ontology-based text mining dataset for extraction of process-structure-property entities
di: Durmaz, Ali Riza, et al.
Pubblicazione: (2024)
di: Durmaz, Ali Riza, et al.
Pubblicazione: (2024)
SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text
di: Ghosh, Reshmi, et al.
Pubblicazione: (2024)
di: Ghosh, Reshmi, et al.
Pubblicazione: (2024)
Synthetically generated text for supervised text analysis
di: Halterman, Andrew
Pubblicazione: (2023)
di: Halterman, Andrew
Pubblicazione: (2023)
Transferable speech-to-text large language model alignment module
di: Wu, Boyong, et al.
Pubblicazione: (2024)
di: Wu, Boyong, et al.
Pubblicazione: (2024)
A thorough benchmark of automatic text classification: From traditional approaches to large language models
di: Cunha, Washington, et al.
Pubblicazione: (2025)
di: Cunha, Washington, et al.
Pubblicazione: (2025)
Spider4SSC & S2CLite: A text-to-multi-query-language dataset using lightweight ontology-agnostic SPARQL to Cypher parser
di: Vejvar, Martin, et al.
Pubblicazione: (2025)
di: Vejvar, Martin, et al.
Pubblicazione: (2025)
EDEN: Empathetic Dialogues for English learning
di: Siyan, Li, et al.
Pubblicazione: (2024)
di: Siyan, Li, et al.
Pubblicazione: (2024)
Using Adaptive Empathetic Responses for Teaching English
di: Siyan, Li, et al.
Pubblicazione: (2024)
di: Siyan, Li, et al.
Pubblicazione: (2024)
DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning
di: Siyan, Li, et al.
Pubblicazione: (2026)
di: Siyan, Li, et al.
Pubblicazione: (2026)
PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech
di: Rahman, Hanif
Pubblicazione: (2026)
di: Rahman, Hanif
Pubblicazione: (2026)
How do we measure privacy in text? A survey of text anonymization metrics
di: Ren, Yaxuan, et al.
Pubblicazione: (2025)
di: Ren, Yaxuan, et al.
Pubblicazione: (2025)
Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language
di: Jimoh, Toheeb Aduramomi, et al.
Pubblicazione: (2026)
di: Jimoh, Toheeb Aduramomi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A large-scale image-text dataset benchmark for farmland segmentation
di: Tao, Chao, et al.
Pubblicazione: (2025) -
AlleNoise: large-scale text classification benchmark dataset with real-world label noise
di: Rączkowska, Alicja, et al.
Pubblicazione: (2024) -
Divergence Decoding: Inference-Time Unlearning via Auxiliary Models
di: Merchant, Humzah, et al.
Pubblicazione: (2026) -
A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs
di: Merchant, Humzah, et al.
Pubblicazione: (2025) -
Mapping the Web of Science, a large-scale graph and text-based dataset with LLM embeddings
di: Kunt, Tim, et al.
Pubblicazione: (2026)