BeanCounter: A low-toxicity, large-scale, and open dataset of business-oriented text
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Siyan, Levy, Bradford |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A large-scale image-text dataset benchmark for farmland segmentation
by: Tao, Chao, et al.
Published: (2025)
by: Tao, Chao, et al.
Published: (2025)
AlleNoise: large-scale text classification benchmark dataset with real-world label noise
by: Rączkowska, Alicja, et al.
Published: (2024)
by: Rączkowska, Alicja, et al.
Published: (2024)
Divergence Decoding: Inference-Time Unlearning via Auxiliary Models
by: Merchant, Humzah, et al.
Published: (2026)
by: Merchant, Humzah, et al.
Published: (2026)
A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs
by: Merchant, Humzah, et al.
Published: (2025)
by: Merchant, Humzah, et al.
Published: (2025)
Mapping the Web of Science, a large-scale graph and text-based dataset with LLM embeddings
by: Kunt, Tim, et al.
Published: (2026)
by: Kunt, Tim, et al.
Published: (2026)
GLeMM: A large-scale multilingual dataset for morphological research
by: Nabil, Hathout, et al.
Published: (2026)
by: Nabil, Hathout, et al.
Published: (2026)
Private prediction for large-scale synthetic text generation
by: Amin, Kareem, et al.
Published: (2024)
by: Amin, Kareem, et al.
Published: (2024)
TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models
by: Feng, Jiarui, et al.
Published: (2024)
by: Feng, Jiarui, et al.
Published: (2024)
CrowdCounter: A benchmark type-specific multi-target counterspeech dataset
by: Saha, Punyajoy, et al.
Published: (2024)
by: Saha, Punyajoy, et al.
Published: (2024)
Designing large language model prompts to extract scores from messy text: A shared dataset and challenge
by: Thelwall, Mike
Published: (2026)
by: Thelwall, Mike
Published: (2026)
CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models
by: Wang, Liangdong, et al.
Published: (2024)
by: Wang, Liangdong, et al.
Published: (2024)
ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining
by: Malik, Haq Nawaz
Published: (2026)
by: Malik, Haq Nawaz
Published: (2026)
A multi-level multi-label text classification dataset of 19th century Ottoman and Russian literary and critical texts
by: Gokceoglu, Gokcen, et al.
Published: (2024)
by: Gokceoglu, Gokcen, et al.
Published: (2024)
ProText: A benchmark dataset for measuring (mis)gendering in long-form texts
by: Kotek, Hadas, et al.
Published: (2026)
by: Kotek, Hadas, et al.
Published: (2026)
TANQ: An open domain dataset of table answered questions
by: Akhtar, Mubashara, et al.
Published: (2024)
by: Akhtar, Mubashara, et al.
Published: (2024)
Reading Between the Lines: A dataset and a study on why some texts are tougher than others
by: Khallaf, Nouran, et al.
Published: (2025)
by: Khallaf, Nouran, et al.
Published: (2025)
A safety realignment framework via subspace-oriented model fusion for large language models
by: Yi, Xin, et al.
Published: (2024)
by: Yi, Xin, et al.
Published: (2024)
Figurative Archive: an open dataset and web-based application for the study of metaphor
by: Bressler, Maddalena, et al.
Published: (2025)
by: Bressler, Maddalena, et al.
Published: (2025)
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
Prompting open-source and commercial language models for grammatical error correction of English learner text
by: Davis, Christopher, et al.
Published: (2024)
by: Davis, Christopher, et al.
Published: (2024)
Explanation sensitivity to the randomness of large language models: the case of journalistic text classification
by: Bogaert, Jeremie, et al.
Published: (2024)
by: Bogaert, Jeremie, et al.
Published: (2024)
600k-ks-ocr: a large-scale synthetic dataset for optical character recognition in kashmiri script
by: Malik, Haq Nawaz
Published: (2026)
by: Malik, Haq Nawaz
Published: (2026)
OpenStaxQA: A multilingual dataset based on open-source college textbooks
by: Gupta, Pranav
Published: (2025)
by: Gupta, Pranav
Published: (2025)
The signal is the ceiling: Measurement limits of LLM-predicted experience ratings from open-ended survey text
by: Hong, Andrew, et al.
Published: (2026)
by: Hong, Andrew, et al.
Published: (2026)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
by: Li, Kunning, et al.
Published: (2025)
by: Li, Kunning, et al.
Published: (2025)
A scale of conceptual orality and literacy: Automatic text categorization in the tradition of "Nähe und Distanz"
by: Emmrich, Volker
Published: (2025)
by: Emmrich, Volker
Published: (2025)
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
by: Pletenev, Sergey, et al.
Published: (2025)
by: Pletenev, Sergey, et al.
Published: (2025)
ARC-Encoder: learning compressed text representations for large language models
by: Pilchen, Hippolyte, et al.
Published: (2025)
by: Pilchen, Hippolyte, et al.
Published: (2025)
MaterioMiner -- An ontology-based text mining dataset for extraction of process-structure-property entities
by: Durmaz, Ali Riza, et al.
Published: (2024)
by: Durmaz, Ali Riza, et al.
Published: (2024)
SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text
by: Ghosh, Reshmi, et al.
Published: (2024)
by: Ghosh, Reshmi, et al.
Published: (2024)
Synthetically generated text for supervised text analysis
by: Halterman, Andrew
Published: (2023)
by: Halterman, Andrew
Published: (2023)
Transferable speech-to-text large language model alignment module
by: Wu, Boyong, et al.
Published: (2024)
by: Wu, Boyong, et al.
Published: (2024)
A thorough benchmark of automatic text classification: From traditional approaches to large language models
by: Cunha, Washington, et al.
Published: (2025)
by: Cunha, Washington, et al.
Published: (2025)
Spider4SSC & S2CLite: A text-to-multi-query-language dataset using lightweight ontology-agnostic SPARQL to Cypher parser
by: Vejvar, Martin, et al.
Published: (2025)
by: Vejvar, Martin, et al.
Published: (2025)
EDEN: Empathetic Dialogues for English learning
by: Siyan, Li, et al.
Published: (2024)
by: Siyan, Li, et al.
Published: (2024)
Using Adaptive Empathetic Responses for Teaching English
by: Siyan, Li, et al.
Published: (2024)
by: Siyan, Li, et al.
Published: (2024)
DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning
by: Siyan, Li, et al.
Published: (2026)
by: Siyan, Li, et al.
Published: (2026)
PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech
by: Rahman, Hanif
Published: (2026)
by: Rahman, Hanif
Published: (2026)
How do we measure privacy in text? A survey of text anonymization metrics
by: Ren, Yaxuan, et al.
Published: (2025)
by: Ren, Yaxuan, et al.
Published: (2025)
Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language
by: Jimoh, Toheeb Aduramomi, et al.
Published: (2026)
by: Jimoh, Toheeb Aduramomi, et al.
Published: (2026)
Similar Items
-
A large-scale image-text dataset benchmark for farmland segmentation
by: Tao, Chao, et al.
Published: (2025) -
AlleNoise: large-scale text classification benchmark dataset with real-world label noise
by: Rączkowska, Alicja, et al.
Published: (2024) -
Divergence Decoding: Inference-Time Unlearning via Auxiliary Models
by: Merchant, Humzah, et al.
Published: (2026) -
A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs
by: Merchant, Humzah, et al.
Published: (2025) -
Mapping the Web of Science, a large-scale graph and text-based dataset with LLM embeddings
by: Kunt, Tim, et al.
Published: (2026)