AustroTox: A Dataset for Target-Based Austrian German Offensive Language Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pachinger, Pia, Goldzycher, Janis, Planitzer, Anna Maria, Kusa, Wojciech, Hanbury, Allan, Neidhardt, Julia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PL-Guard: Benchmarking Language Model Safety for Polish
von: Krasnodębska, Aleksandra, et al.
Veröffentlicht: (2025)
von: Krasnodębska, Aleksandra, et al.
Veröffentlicht: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
German also Hallucinates! Inconsistency Detection in News Summaries with the Absinth Dataset
von: Mascarell, Laura, et al.
Veröffentlicht: (2024)
von: Mascarell, Laura, et al.
Veröffentlicht: (2024)
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
von: Toker, Michael, et al.
Veröffentlicht: (2024)
von: Toker, Michael, et al.
Veröffentlicht: (2024)
Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus
von: Boratyn, Daria, et al.
Veröffentlicht: (2026)
von: Boratyn, Daria, et al.
Veröffentlicht: (2026)
AI-assisted German Employment Contract Review: A Benchmark Dataset
von: Wardas, Oliver, et al.
Veröffentlicht: (2025)
von: Wardas, Oliver, et al.
Veröffentlicht: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
von: Klöser, Lars, et al.
Veröffentlicht: (2024)
von: Klöser, Lars, et al.
Veröffentlicht: (2024)
Mobile Phone Sensor-based Nigerian Driving Dataset to Detect Alcohol-influenced Behaviours
von: Thompson, Iniakpokeikiye Peter, et al.
Veröffentlicht: (2025)
von: Thompson, Iniakpokeikiye Peter, et al.
Veröffentlicht: (2025)
Locations of Characters in Narratives: Andersen and Persuasion Datasets
von: Ozyurt, Batuhan, et al.
Veröffentlicht: (2025)
von: Ozyurt, Batuhan, et al.
Veröffentlicht: (2025)
DimStance: Multilingual Datasets for Dimensional Stance Analysis
von: Becker, Jonas, et al.
Veröffentlicht: (2026)
von: Becker, Jonas, et al.
Veröffentlicht: (2026)
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin
von: Bothwell, Stephen, et al.
Veröffentlicht: (2024)
von: Bothwell, Stephen, et al.
Veröffentlicht: (2024)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2026)
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2026)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
ML-Promise: A Multilingual Dataset for Corporate Promise Verification
von: Seki, Yohei, et al.
Veröffentlicht: (2024)
von: Seki, Yohei, et al.
Veröffentlicht: (2024)
Clinical Document Corpora -- Real Ones, Translated and Synthetic Substitutes, and Assorted Domain Proxies: A Survey of Diversity in Corpus Design, with Focus on German Text Data
von: Hahn, Udo
Veröffentlicht: (2024)
von: Hahn, Udo
Veröffentlicht: (2024)
EMO-KNOW: A Large Scale Dataset on Emotion and Emotion-cause
von: Nguyen, Mia Huong, et al.
Veröffentlicht: (2024)
von: Nguyen, Mia Huong, et al.
Veröffentlicht: (2024)
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
von: Sileo, Damien
Veröffentlicht: (2024)
von: Sileo, Damien
Veröffentlicht: (2024)
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis
von: Bose, Joy
Veröffentlicht: (2026)
von: Bose, Joy
Veröffentlicht: (2026)
Blocks Architecture (BloArk): Efficient, Cost-Effective, and Incremental Dataset Architecture for Wikipedia Revision History
von: Li, Lingxi, et al.
Veröffentlicht: (2024)
von: Li, Lingxi, et al.
Veröffentlicht: (2024)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis
von: Lee, Lung-Hao, et al.
Veröffentlicht: (2026)
von: Lee, Lung-Hao, et al.
Veröffentlicht: (2026)
Tracking Semantic Change in Slovene: A Novel Dataset and Optimal Transport-Based Distance
von: Pranjić, Marko, et al.
Veröffentlicht: (2024)
von: Pranjić, Marko, et al.
Veröffentlicht: (2024)
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
von: Lequeu, Pierre-Antoine, et al.
Veröffentlicht: (2026)
von: Lequeu, Pierre-Antoine, et al.
Veröffentlicht: (2026)
MedHal: An Evaluation Dataset for Medical Hallucination Detection
von: Mehenni, Gaya, et al.
Veröffentlicht: (2025)
von: Mehenni, Gaya, et al.
Veröffentlicht: (2025)
I run as fast as a rabbit, can you? A Multilingual Simile Dialogue Dataset
von: Ma, Longxuan, et al.
Veröffentlicht: (2023)
von: Ma, Longxuan, et al.
Veröffentlicht: (2023)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
LLMs for Legal Subsumption in German Employment Contracts
von: Wardas, Oliver, et al.
Veröffentlicht: (2025)
von: Wardas, Oliver, et al.
Veröffentlicht: (2025)
Lightweight Connective Detection Using Gradient Boosting
von: Er, Mustafa Erolcan, et al.
Veröffentlicht: (2024)
von: Er, Mustafa Erolcan, et al.
Veröffentlicht: (2024)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
von: Ge, Danying, et al.
Veröffentlicht: (2025)
von: Ge, Danying, et al.
Veröffentlicht: (2025)
RedHerring Attack: Testing the Reliability of Attack Detection
von: Rusert, Jonathan
Veröffentlicht: (2025)
von: Rusert, Jonathan
Veröffentlicht: (2025)
Profiling German Text Simplification with Interpretable Model-Fingerprints
von: Klöser, Lars, et al.
Veröffentlicht: (2026)
von: Klöser, Lars, et al.
Veröffentlicht: (2026)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
von: Gaim, Fitsum, et al.
Veröffentlicht: (2025)
von: Gaim, Fitsum, et al.
Veröffentlicht: (2025)
RuOpinionNE-2024: Extraction of Opinion Tuples from Russian News Texts
von: Loukachevitch, Natalia, et al.
Veröffentlicht: (2025)
von: Loukachevitch, Natalia, et al.
Veröffentlicht: (2025)
Synthetic Voice Data for Automatic Speech Recognition in African Languages
von: DeRenzi, Brian, et al.
Veröffentlicht: (2025)
von: DeRenzi, Brian, et al.
Veröffentlicht: (2025)
Boosting Accuracy and Interpretability in Multilingual Hate Speech Detection Through Layer Freezing and Explainable AI
von: Bilehsavar, Meysam Shirdel, et al.
Veröffentlicht: (2026)
von: Bilehsavar, Meysam Shirdel, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PL-Guard: Benchmarking Language Model Safety for Polish
von: Krasnodębska, Aleksandra, et al.
Veröffentlicht: (2025) -
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025) -
German also Hallucinates! Inconsistency Detection in News Summaries with the Absinth Dataset
von: Mascarell, Laura, et al.
Veröffentlicht: (2024) -
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024) -
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
von: Toker, Michael, et al.
Veröffentlicht: (2024)