Literary and Colloquial Tamil Dialect Identification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nanmalar, M., Vijayalakshmi, P., Nagarajan, T. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Literary and Colloquial Dialect Identification for Tamil using Acoustic Features
von: Nanmalar, M., et al.
Veröffentlicht: (2024)
von: Nanmalar, M., et al.
Veröffentlicht: (2024)
A Feature Engineering Approach for Literary and Colloquial Tamil Speech Classification using 1D-CNN
von: Nanmalar, M., et al.
Veröffentlicht: (2024)
von: Nanmalar, M., et al.
Veröffentlicht: (2024)
Spontaneous Informal Speech Dataset for Punctuation Restoration
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing
von: Zhang, Yixiao, et al.
Veröffentlicht: (2023)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2023)
A conversational gesture synthesis system based on emotions and semantics
von: Hoang-Minh, Thanh
Veröffentlicht: (2025)
von: Hoang-Minh, Thanh
Veröffentlicht: (2025)
LLAMAPIE: Proactive In-Ear Conversation Assistants
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
VoXtream2: Full-stream TTS with dynamic speaking rate control
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
Efficient Ensemble for Multimodal Punctuation Restoration using Time-Delay Neural Network
von: Liu, Xing Yi, et al.
Veröffentlicht: (2023)
von: Liu, Xing Yi, et al.
Veröffentlicht: (2023)
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech
von: Ali, Hasmot, et al.
Veröffentlicht: (2024)
von: Ali, Hasmot, et al.
Veröffentlicht: (2024)
Harnessing Smartwatch Microphone Sensors for Cough Detection and Classification
von: Jaiswal, Pranay, et al.
Veröffentlicht: (2024)
von: Jaiswal, Pranay, et al.
Veröffentlicht: (2024)
DOO-RE: A dataset of ambient sensors in a meeting room for activity recognition
von: Kim, Hyunju, et al.
Veröffentlicht: (2024)
von: Kim, Hyunju, et al.
Veröffentlicht: (2024)
Voice Passing : a Non-Binary Voice Gender Prediction System for evaluating Transgender voice transition
von: Doukhan, David, et al.
Veröffentlicht: (2024)
von: Doukhan, David, et al.
Veröffentlicht: (2024)
Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation
von: Garcia, Nelly, et al.
Veröffentlicht: (2026)
von: Garcia, Nelly, et al.
Veröffentlicht: (2026)
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
von: Yuan, Kuang, et al.
Veröffentlicht: (2025)
von: Yuan, Kuang, et al.
Veröffentlicht: (2025)
Improving AI-generated music with user-guided training
von: Singh, Vishwa Mohan, et al.
Veröffentlicht: (2025)
von: Singh, Vishwa Mohan, et al.
Veröffentlicht: (2025)
Human Feedback Driven Dynamic Speech Emotion Recognition
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
von: Sankey-Olsen, Cuno, et al.
Veröffentlicht: (2025)
von: Sankey-Olsen, Cuno, et al.
Veröffentlicht: (2025)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
von: Hui, Macarious, et al.
Veröffentlicht: (2024)
von: Hui, Macarious, et al.
Veröffentlicht: (2024)
MCMChaos: Improvising Rap Music with MCMC Methods and Chaos Theory
von: Kimelman, Robert G.
Veröffentlicht: (2024)
von: Kimelman, Robert G.
Veröffentlicht: (2024)
Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
Are Expressions for Music Emotions the Same Across Cultures?
von: Celen, Elif, et al.
Veröffentlicht: (2025)
von: Celen, Elif, et al.
Veröffentlicht: (2025)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
von: Xie, Zhifei, et al.
Veröffentlicht: (2024)
von: Xie, Zhifei, et al.
Veröffentlicht: (2024)
The language of sound search: Examining User Queries in Audio Search Engines
von: Weck, Benno, et al.
Veröffentlicht: (2024)
von: Weck, Benno, et al.
Veröffentlicht: (2024)
Call2Instruct: Automated Pipeline for Generating Q&A Datasets from Call Center Recordings for LLM Fine-Tuning
von: Echeverria, Alex, et al.
Veröffentlicht: (2025)
von: Echeverria, Alex, et al.
Veröffentlicht: (2025)
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
von: Liu, Tianyun
Veröffentlicht: (2025)
von: Liu, Tianyun
Veröffentlicht: (2025)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
EmoHeal: An End-to-End System for Personalized Therapeutic Music Retrieval from Fine-grained Emotions
von: Wan, Xinchen, et al.
Veröffentlicht: (2025)
von: Wan, Xinchen, et al.
Veröffentlicht: (2025)
Early Detection of Furniture-Infesting Wood-Boring Beetles Using CNN-LSTM Networks and MFCC-Based Acoustic Features
von: Manukalpa, J. M. Chan Sri, et al.
Veröffentlicht: (2025)
von: Manukalpa, J. M. Chan Sri, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Literary and Colloquial Dialect Identification for Tamil using Acoustic Features
von: Nanmalar, M., et al.
Veröffentlicht: (2024) -
A Feature Engineering Approach for Literary and Colloquial Tamil Speech Classification using 1D-CNN
von: Nanmalar, M., et al.
Veröffentlicht: (2024) -
Spontaneous Informal Speech Dataset for Punctuation Restoration
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024) -
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024) -
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)