Nollywood: Let's Go to the Movies!
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ortega, John E., Ahmad, Ibrahim Said, Chen, William |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Using LLM for Real-Time Transcription and Summarization of Doctor-Patient Interactions into ePuskesmas in Indonesia: A Proof-of-Concept Study
von: Khatim, Nur Ahmad, et al.
Veröffentlicht: (2024)
von: Khatim, Nur Ahmad, et al.
Veröffentlicht: (2024)
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID
von: Handoyo, Ahmad Alfani, et al.
Veröffentlicht: (2024)
von: Handoyo, Ahmad Alfani, et al.
Veröffentlicht: (2024)
Towards Robust Speech Representation Learning for Thousands of Languages
von: Chen, William, et al.
Veröffentlicht: (2024)
von: Chen, William, et al.
Veröffentlicht: (2024)
RADIA -- Radio Advertisement Detection with Intelligent Analytics
von: Álvarez, Jorge, et al.
Veröffentlicht: (2024)
von: Álvarez, Jorge, et al.
Veröffentlicht: (2024)
Luganda Speech Intent Recognition for IoT Applications
von: Katumba, Andrew, et al.
Veröffentlicht: (2024)
von: Katumba, Andrew, et al.
Veröffentlicht: (2024)
VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
Leveraging Audio and Text Modalities in Mental Health: A Study of LLMs Performance
von: Ali, Abdelrahman A., et al.
Veröffentlicht: (2024)
von: Ali, Abdelrahman A., et al.
Veröffentlicht: (2024)
ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering
von: Song, Yakun, et al.
Veröffentlicht: (2024)
von: Song, Yakun, et al.
Veröffentlicht: (2024)
VoiceBench: Benchmarking LLM-Based Voice Assistants
von: Chen, Yiming, et al.
Veröffentlicht: (2024)
von: Chen, Yiming, et al.
Veröffentlicht: (2024)
Optimizing Automatic Speech Assessment: W-RankSim Regularization and Hybrid Feature Fusion Strategies
von: Wu, Chung-Wen, et al.
Veröffentlicht: (2024)
von: Wu, Chung-Wen, et al.
Veröffentlicht: (2024)
Effective Noise-aware Data Simulation for Domain-adaptive Speech Enhancement Leveraging Dynamic Stochastic Perturbation
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment
von: Chao, Fu-An, et al.
Veröffentlicht: (2025)
von: Chao, Fu-An, et al.
Veröffentlicht: (2025)
Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
LLM-Driven Multimodal Opinion Expression Identification
von: Jia, Bonian, et al.
Veröffentlicht: (2024)
von: Jia, Bonian, et al.
Veröffentlicht: (2024)
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
von: Song, Zheshu, et al.
Veröffentlicht: (2024)
von: Song, Zheshu, et al.
Veröffentlicht: (2024)
ALICE: A Multifaceted Evaluation Framework of Large Audio-Language Models' In-Context Learning Ability
von: Piao, Yen-Ting, et al.
Veröffentlicht: (2026)
von: Piao, Yen-Ting, et al.
Veröffentlicht: (2026)
SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
BAT: Learning to Reason about Spatial Sounds with Large Language Models
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2024)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2024)
Multi-class Decoding of Attended Speaker Direction Using Electroencephalogram and Audio Spatial Spectrum
von: Zhang, Yuanming, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanming, et al.
Veröffentlicht: (2024)
ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
PHRASED: Phrase Dictionary Biasing for Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
A unified front-end framework for English text-to-speech synthesis
von: Ying, Zelin, et al.
Veröffentlicht: (2023)
von: Ying, Zelin, et al.
Veröffentlicht: (2023)
KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening
von: Sharma, Rohan, et al.
Veröffentlicht: (2025)
von: Sharma, Rohan, et al.
Veröffentlicht: (2025)
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
von: Zhuo, Jianheng, et al.
Veröffentlicht: (2025)
Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder
von: Dai, Yusheng, et al.
Veröffentlicht: (2023)
von: Dai, Yusheng, et al.
Veröffentlicht: (2023)
All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation
von: Foo, Leonardo Haw-Yang, et al.
Veröffentlicht: (2026)
von: Foo, Leonardo Haw-Yang, et al.
Veröffentlicht: (2026)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
Recent Advances in End-to-End Simultaneous Speech Translation
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2024)
Efficient Dialect-Aware Modeling and Conditioning for Low-Resource Taiwanese Hakka Speech Processing
von: Peng, An-Ci, et al.
Veröffentlicht: (2026)
von: Peng, An-Ci, et al.
Veröffentlicht: (2026)
H-PRM: A Pluggable Hotword Pre-Retrieval Module for Various Speech Recognition Systems
von: Dai, Huangyu, et al.
Veröffentlicht: (2025)
von: Dai, Huangyu, et al.
Veröffentlicht: (2025)
SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition
von: Wang, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Wang, Yi-Cheng, et al.
Veröffentlicht: (2024)
Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models
von: Ieong, Lok-Lam, et al.
Veröffentlicht: (2026)
von: Ieong, Lok-Lam, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Using LLM for Real-Time Transcription and Summarization of Doctor-Patient Interactions into ePuskesmas in Indonesia: A Proof-of-Concept Study
von: Khatim, Nur Ahmad, et al.
Veröffentlicht: (2024) -
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID
von: Handoyo, Ahmad Alfani, et al.
Veröffentlicht: (2024) -
Towards Robust Speech Representation Learning for Thousands of Languages
von: Chen, William, et al.
Veröffentlicht: (2024) -
RADIA -- Radio Advertisement Detection with Intelligent Analytics
von: Álvarez, Jorge, et al.
Veröffentlicht: (2024) -
Luganda Speech Intent Recognition for IoT Applications
von: Katumba, Andrew, et al.
Veröffentlicht: (2024)