Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khorrami, Khazar, Räsänen, Okko |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Computational modeling of early language learning from acoustic speech and audiovisual input without linguistic priors
von: Räsänen, Okko
Veröffentlicht: (2026)
von: Räsänen, Okko
Veröffentlicht: (2026)
NAAQA: A Neural Architecture for Acoustic Question Answering
von: Abdelnour, Jerome, et al.
Veröffentlicht: (2021)
von: Abdelnour, Jerome, et al.
Veröffentlicht: (2021)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
von: Nfissi, Alaa, et al.
Veröffentlicht: (2025)
von: Nfissi, Alaa, et al.
Veröffentlicht: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
Unveiling Hidden Factors: Explainable AI for Feature Boosting in Speech Emotion Recognition
von: Nfissi, Alaa, et al.
Veröffentlicht: (2024)
von: Nfissi, Alaa, et al.
Veröffentlicht: (2024)
Iterative Feature Boosting for Explainable Speech Emotion Recognition
von: Nfissi, Alaa, et al.
Veröffentlicht: (2024)
von: Nfissi, Alaa, et al.
Veröffentlicht: (2024)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
von: Simões, Lucca Emmanuel Pineli, et al.
Veröffentlicht: (2024)
von: Simões, Lucca Emmanuel Pineli, et al.
Veröffentlicht: (2024)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
Quantifying the effect of speech pathology on automatic and human speaker verification
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
Measuring the Accuracy of Automatic Speech Recognition Solutions
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
Everyday Speech in the Indian Subcontinent
von: P, Utkarsh
Veröffentlicht: (2024)
von: P, Utkarsh
Veröffentlicht: (2024)
Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
Beyond Levenshtein: Leveraging Multiple Algorithms for Robust Word Error Rate Computations And Granular Error Classifications
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026)
von: Yang, Shan
Veröffentlicht: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
A Voice-based Triage for Type 2 Diabetes using a Conversational Virtual Assistant in the Home Environment
von: Summoogum, Kelvin, et al.
Veröffentlicht: (2024)
von: Summoogum, Kelvin, et al.
Veröffentlicht: (2024)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
von: Dai, Song, et al.
Veröffentlicht: (2025)
von: Dai, Song, et al.
Veröffentlicht: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
The formation of perceptual space in early phonetic acquisition: a cross-linguistic modeling approach
von: Tan, Frank Lihui, et al.
Veröffentlicht: (2024)
von: Tan, Frank Lihui, et al.
Veröffentlicht: (2024)
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset
von: Marie, Ambre, et al.
Veröffentlicht: (2025)
von: Marie, Ambre, et al.
Veröffentlicht: (2025)
Communication Access Real-Time Translation Through Collaborative Correction of Automatic Speech Recognition
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2025)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2025)
A Domain Knowledge Informed Approach for Anomaly Detection of Electric Vehicle Interior Sounds
von: Kunte, Deepti, et al.
Veröffentlicht: (2025)
von: Kunte, Deepti, et al.
Veröffentlicht: (2025)
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
Developing Acoustic Models for Automatic Speech Recognition in Swedish
von: Salvi, Giampiero
Veröffentlicht: (2024)
von: Salvi, Giampiero
Veröffentlicht: (2024)
Hearing Anything Anywhere
von: Wang, Mason, et al.
Veröffentlicht: (2024)
von: Wang, Mason, et al.
Veröffentlicht: (2024)
Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device
von: Kozak, Nazar
Veröffentlicht: (2026)
von: Kozak, Nazar
Veröffentlicht: (2026)
Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages
von: Zaragozá, Lucía Gómez, et al.
Veröffentlicht: (2024)
von: Zaragozá, Lucía Gómez, et al.
Veröffentlicht: (2024)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
von: Papyan, Narek, et al.
Veröffentlicht: (2024)
von: Papyan, Narek, et al.
Veröffentlicht: (2024)
A Baseline Multimodal Approach to Emotion Recognition in Conversations
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation
von: Yang, Dingyi, et al.
Veröffentlicht: (2024)
von: Yang, Dingyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Computational modeling of early language learning from acoustic speech and audiovisual input without linguistic priors
von: Räsänen, Okko
Veröffentlicht: (2026) -
NAAQA: A Neural Architecture for Acoustic Question Answering
von: Abdelnour, Jerome, et al.
Veröffentlicht: (2021) -
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
von: Adelson, Trevor, et al.
Veröffentlicht: (2026) -
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
von: Nfissi, Alaa, et al.
Veröffentlicht: (2025) -
Learning the meanings of function words from grounded language using a visual question answering model
von: Portelance, Eva, et al.
Veröffentlicht: (2023)