TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation
Fuente:
arXiv
Saved in:
| Main Authors: | Hutiri, Wiebke, Cimpoi, Mircea, Scheuerman, Morgan, Matthews, Victoria, Xiang, Alice |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators
by: Hutiri, Wiebke, et al.
Published: (2024)
by: Hutiri, Wiebke, et al.
Published: (2024)
As Biased as You Measure: Methodological Pitfalls of Bias Evaluations in Speaker Verification Research
by: Hutiri, Wiebke, et al.
Published: (2024)
by: Hutiri, Wiebke, et al.
Published: (2024)
Sound Check: Auditing Audio Datasets
by: Agnew, William, et al.
Published: (2024)
by: Agnew, William, et al.
Published: (2024)
How to Evaluate Automatic Speech Recognition: Comparing Different Performance and Bias Measures
by: Patel, Tanvina, et al.
Published: (2025)
by: Patel, Tanvina, et al.
Published: (2025)
Yes, But Not Always. Generative AI Needs Nuanced Opt-in
by: Hutiri, Wiebke, et al.
Published: (2026)
by: Hutiri, Wiebke, et al.
Published: (2026)
FairLENS: Assessing Fairness in Law Enforcement Speech Recognition
by: Wang, Yicheng, et al.
Published: (2024)
by: Wang, Yicheng, et al.
Published: (2024)
MusGO: A Community-Driven Framework For Assessing Openness in Music-Generative AI
by: Batlle-Roca, Roser, et al.
Published: (2025)
by: Batlle-Roca, Roser, et al.
Published: (2025)
Voice EHR: Introducing Multimodal Audio Data for Health
by: Anibal, James, et al.
Published: (2024)
by: Anibal, James, et al.
Published: (2024)
IndieFake Dataset: A Benchmark Dataset for Audio Deepfake Detection
by: Kumar, Abhay, et al.
Published: (2025)
by: Kumar, Abhay, et al.
Published: (2025)
Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain
by: Le-Duc, Khai
Published: (2024)
by: Le-Duc, Khai
Published: (2024)
SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
by: Pandey, Prabhat, et al.
Published: (2025)
by: Pandey, Prabhat, et al.
Published: (2025)
Audio Atlas: Visualizing and Exploring Audio Datasets
by: Lanzendörfer, Luca A., et al.
Published: (2024)
by: Lanzendörfer, Luca A., et al.
Published: (2024)
Audio Deepfake Attribution: An Initial Dataset and Investigation
by: Yan, Xinrui, et al.
Published: (2022)
by: Yan, Xinrui, et al.
Published: (2022)
The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio
by: Xie, Yuankun, et al.
Published: (2024)
by: Xie, Yuankun, et al.
Published: (2024)
Heterogeneous sound classification with the Broad Sound Taxonomy and Dataset
by: Anastasopoulou, Panagiota, et al.
Published: (2024)
by: Anastasopoulou, Panagiota, et al.
Published: (2024)
Cross-Domain Audio Deepfake Detection: Dataset and Analysis
by: Li, Yuang, et al.
Published: (2024)
by: Li, Yuang, et al.
Published: (2024)
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
by: Toyin, Hawau Olamide, et al.
Published: (2025)
by: Toyin, Hawau Olamide, et al.
Published: (2025)
Codecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
DroneAudioset: An Audio Dataset for Drone-based Search and Rescue
by: Gupta, Chitralekha, et al.
Published: (2025)
by: Gupta, Chitralekha, et al.
Published: (2025)
CryCeleb: A Speaker Verification Dataset Based on Infant Cry Sounds
by: Budaghyan, David, et al.
Published: (2023)
by: Budaghyan, David, et al.
Published: (2023)
NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription
by: Vinnikov, Alon, et al.
Published: (2024)
by: Vinnikov, Alon, et al.
Published: (2024)
Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
by: Gedeon, Máté, et al.
Published: (2025)
by: Gedeon, Máté, et al.
Published: (2025)
QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
by: Wang, Siyin, et al.
Published: (2025)
by: Wang, Siyin, et al.
Published: (2025)
Effects of Dataset Sampling Rate for Noise Cancellation through Deep Learning
by: Colelough, Brandon, et al.
Published: (2024)
by: Colelough, Brandon, et al.
Published: (2024)
GOAT: A Large Dataset of Paired Guitar Audio Recordings and Tablatures
by: Loth, Jackson, et al.
Published: (2025)
by: Loth, Jackson, et al.
Published: (2025)
Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset
by: Shah, Neil, et al.
Published: (2024)
by: Shah, Neil, et al.
Published: (2024)
A Novel Labeled Human Voice Signal Dataset for Misbehavior Detection
by: Raza, Ali, et al.
Published: (2024)
by: Raza, Ali, et al.
Published: (2024)
Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
by: Ji, Zhoulin, et al.
Published: (2024)
by: Ji, Zhoulin, et al.
Published: (2024)
Multi-Speaker Conversational Audio Deepfake: Taxonomy, Dataset and Pilot Study
by: Ahmed, Alabi, et al.
Published: (2026)
by: Ahmed, Alabi, et al.
Published: (2026)
AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models
by: Li, Kai, et al.
Published: (2025)
by: Li, Kai, et al.
Published: (2025)
Brilla AI: AI Contestant for the National Science and Maths Quiz
by: Boateng, George, et al.
Published: (2024)
by: Boateng, George, et al.
Published: (2024)
The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege
by: He, Luxi, et al.
Published: (2025)
by: He, Luxi, et al.
Published: (2025)
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
Quranic Audio Dataset: Crowdsourced and Labeled Recitation from Non-Arabic Speakers
by: Salameh, Raghad, et al.
Published: (2024)
by: Salameh, Raghad, et al.
Published: (2024)
BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics
by: Rauch, Lukas, et al.
Published: (2024)
by: Rauch, Lukas, et al.
Published: (2024)
BanglaFake: Constructing and Evaluating a Specialized Bengali Deepfake Audio Dataset
by: Fahad, Istiaq Ahmed, et al.
Published: (2025)
by: Fahad, Istiaq Ahmed, et al.
Published: (2025)
MOSA: Music Motion with Semantic Annotation Dataset for Cross-Modal Music Processing
by: Huang, Yu-Fen, et al.
Published: (2024)
by: Huang, Yu-Fen, et al.
Published: (2024)
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
Task-Lens: Cross-Task Utility Based Speech Dataset Profiling for Low-Resource Indian Languages
by: Sharma, Swati, et al.
Published: (2026)
by: Sharma, Swati, et al.
Published: (2026)
Similar Items
-
Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators
by: Hutiri, Wiebke, et al.
Published: (2024) -
As Biased as You Measure: Methodological Pitfalls of Bias Evaluations in Speaker Verification Research
by: Hutiri, Wiebke, et al.
Published: (2024) -
Sound Check: Auditing Audio Datasets
by: Agnew, William, et al.
Published: (2024) -
How to Evaluate Automatic Speech Recognition: Comparing Different Performance and Bias Measures
by: Patel, Tanvina, et al.
Published: (2025) -
Yes, But Not Always. Generative AI Needs Nuanced Opt-in
by: Hutiri, Wiebke, et al.
Published: (2026)