Classification errors distort findings in automated speech processing: examples and solutions from child-development research
Fuente:
arXiv
Saved in:
| Main Authors: | Gautheron, Lucas, Kidd, Evan, Malko, Anton, Lavechin, Marvin, Cristia, Alejandrina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Context-aware child-directed speech detection from long-form recordings
by: Charlot, Théo, et al.
Published: (2026)
by: Charlot, Théo, et al.
Published: (2026)
Fifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to Validity
by: Peurey, Loann, et al.
Published: (2025)
by: Peurey, Loann, et al.
Published: (2025)
Dilemmas and trade-offs in the diffusion of conventions
by: Gautheron, Lucas
Published: (2025)
by: Gautheron, Lucas
Published: (2025)
Balancing Specialization and Adaptation in a Transforming Scientific Landscape
by: Gautheron, Lucas
Published: (2023)
by: Gautheron, Lucas
Published: (2023)
An open-source voice type classifier for child-centered daylong recordings
by: Lavechin, Marvin, et al.
Published: (2020)
by: Lavechin, Marvin, et al.
Published: (2020)
BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
by: Lavechin, Marvin, et al.
Published: (2023)
by: Lavechin, Marvin, et al.
Published: (2023)
BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings
by: Charlot, Théo, et al.
Published: (2025)
by: Charlot, Théo, et al.
Published: (2025)
Employing self-supervised learning models for cross-linguistic child speech maturity classification
by: Zhang, Theo, et al.
Published: (2025)
by: Zhang, Theo, et al.
Published: (2025)
Deep literature reviews: an application of fine-tuned language models to migration research
by: Iacus, Stefano M., et al.
Published: (2025)
by: Iacus, Stefano M., et al.
Published: (2025)
Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
by: Miller, Evan
Published: (2024)
by: Miller, Evan
Published: (2024)
Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier
by: Kunze, Tarek, et al.
Published: (2025)
by: Kunze, Tarek, et al.
Published: (2025)
Beyond the Hype: Embeddings vs. Prompting for Multiclass Classification Tasks
by: Kokkodis, Marios, et al.
Published: (2025)
by: Kokkodis, Marios, et al.
Published: (2025)
Uncertainty quantification in automated valuation models with spatially weighted conformal prediction
by: Hjort, Anders, et al.
Published: (2023)
by: Hjort, Anders, et al.
Published: (2023)
Context-Alignment: Activating and Enhancing LLM Capabilities in Time Series
by: Hu, Yuxiao, et al.
Published: (2025)
by: Hu, Yuxiao, et al.
Published: (2025)
Extracting Emotion Phrases from Tweets using BART
by: Rezapour, Mahdi
Published: (2024)
by: Rezapour, Mahdi
Published: (2024)
Subjective Perspectives within Learned Representations Predict High-Impact Innovation
by: Cao, Likun, et al.
Published: (2025)
by: Cao, Likun, et al.
Published: (2025)
Statistical multi-metric evaluation and visualization of LLM system predictive performance
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
The Proxy Presumption: From Semantic Embeddings to Valid Social Measures
by: Li, Baishi, et al.
Published: (2026)
by: Li, Baishi, et al.
Published: (2026)
A meta-analysis on the performance of machine-learning based language models for sentiment analysis
by: Rohde, Elena, et al.
Published: (2025)
by: Rohde, Elena, et al.
Published: (2025)
How to Correctly Report LLM-as-a-Judge Evaluations
by: Lee, Chungpa, et al.
Published: (2025)
by: Lee, Chungpa, et al.
Published: (2025)
Dynamic Topic Language Model on Heterogeneous Children's Mental Health Clinical Notes
by: Ye, Hanwen, et al.
Published: (2023)
by: Ye, Hanwen, et al.
Published: (2023)
Specific language impairment (SLI) detection pipeline from transcriptions of spontaneous narratives
by: Arena, Santiago, et al.
Published: (2024)
by: Arena, Santiago, et al.
Published: (2024)
Causal Representation Learning with Generative Artificial Intelligence: Application to Texts as Treatments
by: Imai, Kosuke, et al.
Published: (2024)
by: Imai, Kosuke, et al.
Published: (2024)
Advanced Crash Causation Analysis for Freeway Safety: A Large Language Model Approach to Identifying Key Contributing Factors
by: Abdelrahman, Ahmed S., et al.
Published: (2025)
by: Abdelrahman, Ahmed S., et al.
Published: (2025)
Ensemble Kalman filter for uncertainty in human language comprehension
by: Bhandari, Diksha, et al.
Published: (2025)
by: Bhandari, Diksha, et al.
Published: (2025)
Bayesian Evaluation of Large Language Model Behavior
by: Longjohn, Rachel, et al.
Published: (2025)
by: Longjohn, Rachel, et al.
Published: (2025)
Explainable Automatic Grading with Neural Additive Models
by: Condor, Aubrey, et al.
Published: (2024)
by: Condor, Aubrey, et al.
Published: (2024)
Detecting LLM-Generated Text with Performance Guarantees
by: Zhou, Hongyi, et al.
Published: (2026)
by: Zhou, Hongyi, et al.
Published: (2026)
Deep R Programming
by: Gagolewski, Marek
Published: (2022)
by: Gagolewski, Marek
Published: (2022)
A Design-based Solution for Causal Inference with Text: Can a Language Model Be Too Large?
by: Tierney, Graham, et al.
Published: (2025)
by: Tierney, Graham, et al.
Published: (2025)
Variational phylogenetic inference with products over bipartitions
by: Sidrow, Evan, et al.
Published: (2025)
by: Sidrow, Evan, et al.
Published: (2025)
Improving Probabilistic Models in Text Classification via Active Learning
by: Bosley, Mitchell, et al.
Published: (2022)
by: Bosley, Mitchell, et al.
Published: (2022)
Masked Mineral Modeling: Continent-Scale Mineral Prospecting via Geospatial Infilling
by: Nair, Sujay, et al.
Published: (2025)
by: Nair, Sujay, et al.
Published: (2025)
ICE-ID: A Novel Historical Census Dataset for Longitudinal Identity Resolution
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Predicting First Year Dropout from Pre Enrolment Motivation Statements Using Text Mining
by: Soppe, K. F. B., et al.
Published: (2025)
by: Soppe, K. F. B., et al.
Published: (2025)
Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing
by: Wasi, Azmine Toushik, et al.
Published: (2024)
by: Wasi, Azmine Toushik, et al.
Published: (2024)
How to Choose a Threshold for an Evaluation Metric for Large Language Models
by: Sarmah, Bhaskarjit, et al.
Published: (2024)
by: Sarmah, Bhaskarjit, et al.
Published: (2024)
Unified Representation of Genomic and Biomedical Concepts through Multi-Task, Multi-Source Contrastive Learning
by: Yuan, Hongyi, et al.
Published: (2024)
by: Yuan, Hongyi, et al.
Published: (2024)
Domain-Shift-Aware Conformal Prediction for Large Language Models
by: Lin, Zhexiao, et al.
Published: (2025)
by: Lin, Zhexiao, et al.
Published: (2025)
Similar Items
-
Context-aware child-directed speech detection from long-form recordings
by: Charlot, Théo, et al.
Published: (2026) -
Fifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to Validity
by: Peurey, Loann, et al.
Published: (2025) -
Dilemmas and trade-offs in the diffusion of conventions
by: Gautheron, Lucas
Published: (2025) -
Balancing Specialization and Adaptation in a Transforming Scientific Landscape
by: Gautheron, Lucas
Published: (2023) -
An open-source voice type classifier for child-centered daylong recordings
by: Lavechin, Marvin, et al.
Published: (2020)