NAAQA: A Neural Architecture for Acoustic Question Answering
Fuente:
arXiv
Guardado en:
| Autores principales: | Abdelnour, Jerome, Rouat, Jean, Salvi, Giampiero |
|---|---|
| Formato: | Preprint |
| Publicado: |
2021
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Developing Acoustic Models for Automatic Speech Recognition in Swedish
por: Salvi, Giampiero
Publicado: (2024)
por: Salvi, Giampiero
Publicado: (2024)
Segment Boundary Detection via Class Entropy Measurements in Connectionist Phoneme Recognition
por: Salvi, Giampiero
Publicado: (2024)
por: Salvi, Giampiero
Publicado: (2024)
Dynamic Behaviour of Connectionist Speech Recognition with Strong Latency Constraints
por: Salvi, Giampiero
Publicado: (2024)
por: Salvi, Giampiero
Publicado: (2024)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
por: Papyan, Narek, et al.
Publicado: (2024)
por: Papyan, Narek, et al.
Publicado: (2024)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
por: Adelson, Trevor, et al.
Publicado: (2026)
por: Adelson, Trevor, et al.
Publicado: (2026)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
por: Viswanathan, Janaki, et al.
Publicado: (2025)
por: Viswanathan, Janaki, et al.
Publicado: (2025)
Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
por: Khorrami, Khazar, et al.
Publicado: (2021)
por: Khorrami, Khazar, et al.
Publicado: (2021)
Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
por: Lasbordes, Maxence, et al.
Publicado: (2025)
por: Lasbordes, Maxence, et al.
Publicado: (2025)
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
por: Hori, Takaaki, et al.
Publicado: (2025)
por: Hori, Takaaki, et al.
Publicado: (2025)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
por: Kim, Minu, et al.
Publicado: (2025)
por: Kim, Minu, et al.
Publicado: (2025)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
por: Wang, Hsuan-Yu, et al.
Publicado: (2025)
por: Wang, Hsuan-Yu, et al.
Publicado: (2025)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
por: Alamr, Meshal, et al.
Publicado: (2026)
por: Alamr, Meshal, et al.
Publicado: (2026)
Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device
por: Kozak, Nazar
Publicado: (2026)
por: Kozak, Nazar
Publicado: (2026)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
por: Cheng, Zhuangfei, et al.
Publicado: (2025)
por: Cheng, Zhuangfei, et al.
Publicado: (2025)
Quantifying the effect of speech pathology on automatic and human speaker verification
por: Halpern, Bence Mark, et al.
Publicado: (2024)
por: Halpern, Bence Mark, et al.
Publicado: (2024)
Deep Feed-Forward Neural Network for Bangla Isolated Speech Recognition
por: Bhadra, Dipayan, et al.
Publicado: (2025)
por: Bhadra, Dipayan, et al.
Publicado: (2025)
Generation of Musical Timbres using a Text-Guided Diffusion Model
por: Yuan, Weixuan, et al.
Publicado: (2025)
por: Yuan, Weixuan, et al.
Publicado: (2025)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
por: Wang, Shaowen, et al.
Publicado: (2025)
por: Wang, Shaowen, et al.
Publicado: (2025)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
por: Li, Pengcheng, et al.
Publicado: (2024)
por: Li, Pengcheng, et al.
Publicado: (2024)
Impact of Phonetics on Speaker Identity in Adversarial Voice Attack
por: Dar, Daniyal Kabir, et al.
Publicado: (2025)
por: Dar, Daniyal Kabir, et al.
Publicado: (2025)
Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing
por: Cui, Hanfang, et al.
Publicado: (2025)
por: Cui, Hanfang, et al.
Publicado: (2025)
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
por: Nfissi, Alaa, et al.
Publicado: (2025)
por: Nfissi, Alaa, et al.
Publicado: (2025)
Quantization for OpenAI's Whisper Models: A Comparative Analysis
por: Andreyev, Allison
Publicado: (2025)
por: Andreyev, Allison
Publicado: (2025)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction
por: Cheripally, Sowmya
Publicado: (2024)
por: Cheripally, Sowmya
Publicado: (2024)
Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages
por: Zaragozá, Lucía Gómez, et al.
Publicado: (2024)
por: Zaragozá, Lucía Gómez, et al.
Publicado: (2024)
An End-to-End Approach for Korean Wakeword Systems with Speaker Authentication
por: Seo, Geonwoo
Publicado: (2025)
por: Seo, Geonwoo
Publicado: (2025)
Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset
por: Marie, Ambre, et al.
Publicado: (2025)
por: Marie, Ambre, et al.
Publicado: (2025)
Window Size Versus Accuracy Experiments in Voice Activity Detectors
por: McKinnon, Max, et al.
Publicado: (2026)
por: McKinnon, Max, et al.
Publicado: (2026)
Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit
por: Soni, Aniket Abhishek
Publicado: (2025)
por: Soni, Aniket Abhishek
Publicado: (2025)
A Voice-based Triage for Type 2 Diabetes using a Conversational Virtual Assistant in the Home Environment
por: Summoogum, Kelvin, et al.
Publicado: (2024)
por: Summoogum, Kelvin, et al.
Publicado: (2024)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
por: Chen, Kuan-Yu, et al.
Publicado: (2025)
por: Chen, Kuan-Yu, et al.
Publicado: (2025)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
por: Donepudi, Dharma Teja
Publicado: (2025)
por: Donepudi, Dharma Teja
Publicado: (2025)
Taming Audio VAEs via Target-KL Regularization
por: Seetharaman, Prem, et al.
Publicado: (2026)
por: Seetharaman, Prem, et al.
Publicado: (2026)
AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
por: Maben, Leander Melroy, et al.
Publicado: (2025)
por: Maben, Leander Melroy, et al.
Publicado: (2025)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
por: Sharma, Manali, et al.
Publicado: (2026)
por: Sharma, Manali, et al.
Publicado: (2026)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
Ejemplares similares
-
Developing Acoustic Models for Automatic Speech Recognition in Swedish
por: Salvi, Giampiero
Publicado: (2024) -
Segment Boundary Detection via Class Entropy Measurements in Connectionist Phoneme Recognition
por: Salvi, Giampiero
Publicado: (2024) -
Dynamic Behaviour of Connectionist Speech Recognition with Strong Latency Constraints
por: Salvi, Giampiero
Publicado: (2024) -
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
por: Papyan, Narek, et al.
Publicado: (2024) -
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
por: Adelson, Trevor, et al.
Publicado: (2026)