ISPA: Inter-Species Phonetic Alphabet for Transcribing Animal Sounds
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hagiwara, Masato, Miron, Marius, Liu, Jen-Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Biodenoising: Animal Vocalization Denoising without Access to Clean Data
von: Miron, Marius, et al.
Veröffentlicht: (2024)
von: Miron, Marius, et al.
Veröffentlicht: (2024)
Self-Train Before You Transcribe
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
Conversational Rubert for Detecting Competitive Interruptions in ASR-Transcribed Dialogues
von: Galimzianov, Dmitrii, et al.
Veröffentlicht: (2024)
von: Galimzianov, Dmitrii, et al.
Veröffentlicht: (2024)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
von: Li, Xingyuan, et al.
Veröffentlicht: (2024)
von: Li, Xingyuan, et al.
Veröffentlicht: (2024)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
The ART of Conversation: Measuring Phonetic Convergence and Deliberate Imitation in L2-Speech with a Siamese RNN
von: Yuan, Zheng, et al.
Veröffentlicht: (2023)
von: Yuan, Zheng, et al.
Veröffentlicht: (2023)
The Mason-Alberta Phonetic Segmenter: A forced alignment system based on deep neural networks and interpolation
von: Kelley, Matthew C., et al.
Veröffentlicht: (2023)
von: Kelley, Matthew C., et al.
Veröffentlicht: (2023)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
Synthetic data enables context-aware bioacoustic sound event detection
von: Hoffman, Benjamin, et al.
Veröffentlicht: (2025)
von: Hoffman, Benjamin, et al.
Veröffentlicht: (2025)
TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
TOGGL: Transcribing Overlapping Speech with Staggered Labeling
von: Li, Chak-Fai, et al.
Veröffentlicht: (2024)
von: Li, Chak-Fai, et al.
Veröffentlicht: (2024)
GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
von: Li, Zehua Kcriss, et al.
Veröffentlicht: (2024)
von: Li, Zehua Kcriss, et al.
Veröffentlicht: (2024)
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
von: Robinson, David, et al.
Veröffentlicht: (2024)
von: Robinson, David, et al.
Veröffentlicht: (2024)
Transcribing Rhythmic Patterns of the Guitar Track in Polyphonic Music
von: Lukoianov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Lukoianov, Aleksandr, et al.
Veröffentlicht: (2025)
Robust detection of overlapping bioacoustic sound events
von: Mahon, Louis, et al.
Veröffentlicht: (2025)
von: Mahon, Louis, et al.
Veröffentlicht: (2025)
Advanced Framework for Animal Sound Classification With Features Optimization
von: Yang, Qiang, et al.
Veröffentlicht: (2024)
von: Yang, Qiang, et al.
Veröffentlicht: (2024)
Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
Multimodal Input Aids a Bayesian Model of Phonetic Learning
von: Zhi, Sophia, et al.
Veröffentlicht: (2024)
von: Zhi, Sophia, et al.
Veröffentlicht: (2024)
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
von: Chen, Guoguo, et al.
Veröffentlicht: (2021)
von: Chen, Guoguo, et al.
Veröffentlicht: (2021)
A Technique for Isolating Lexically-Independent Phonetic Dependencies in Generative CNNs
von: Šegedin, Bruno Ferenc
Veröffentlicht: (2025)
von: Šegedin, Bruno Ferenc
Veröffentlicht: (2025)
MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research
von: Li, Song, et al.
Veröffentlicht: (2024)
von: Li, Song, et al.
Veröffentlicht: (2024)
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
von: Parcollet, Titouan, et al.
Veröffentlicht: (2025)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2025)
Phonetic Error Analysis of Raw Waveform Acoustic Models with Parametric and Non-Parametric CNNs
von: Loweimi, Erfan, et al.
Veröffentlicht: (2024)
von: Loweimi, Erfan, et al.
Veröffentlicht: (2024)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2025)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2025)
Pre-Finetuning for Few-Shot Emotional Speech Recognition
von: Chen, Maximillian, et al.
Veröffentlicht: (2023)
von: Chen, Maximillian, et al.
Veröffentlicht: (2023)
A Real-Time Lyrics Alignment System Using Chroma And Phonetic Features For Classical Vocal Performance
von: Park, Jiyun, et al.
Veröffentlicht: (2024)
von: Park, Jiyun, et al.
Veröffentlicht: (2024)
SoundReactor: Frame-level Online Video-to-Audio Generation
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
(SimPhon Speech Test): A Data-Driven Method for In Silico Design and Validation of a Phonetically Balanced Speech Test
von: Bleeck, Stefan
Veröffentlicht: (2025)
von: Bleeck, Stefan
Veröffentlicht: (2025)
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
von: Liu, Alexander H., et al.
Veröffentlicht: (2024)
von: Liu, Alexander H., et al.
Veröffentlicht: (2024)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
von: Eren, Eray, et al.
Veröffentlicht: (2025)
von: Eren, Eray, et al.
Veröffentlicht: (2025)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Biodenoising: Animal Vocalization Denoising without Access to Clean Data
von: Miron, Marius, et al.
Veröffentlicht: (2024) -
Self-Train Before You Transcribe
von: Flynn, Robert, et al.
Veröffentlicht: (2024) -
Conversational Rubert for Detecting Competitive Interruptions in ASR-Transcribed Dialogues
von: Galimzianov, Dmitrii, et al.
Veröffentlicht: (2024) -
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025) -
Phonetic and Lexical Discovery of a Canine Language using HuBERT
von: Li, Xingyuan, et al.
Veröffentlicht: (2024)