Salvato in:
| Autore principale: | Togni, Jimi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2503.15501 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
UniGlyph: A Seven-Segment Script for Universal Language Representation
di: Sherin, G. V. Bency, et al.
Pubblicazione: (2024)
di: Sherin, G. V. Bency, et al.
Pubblicazione: (2024)
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
di: Chen, Zhehuai, et al.
Pubblicazione: (2024)
di: Chen, Zhehuai, et al.
Pubblicazione: (2024)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
di: Mehta, Shivam, et al.
Pubblicazione: (2024)
di: Mehta, Shivam, et al.
Pubblicazione: (2024)
Communication Access Real-Time Translation Through Collaborative Correction of Automatic Speech Recognition
di: Kuhn, Korbinian, et al.
Pubblicazione: (2025)
di: Kuhn, Korbinian, et al.
Pubblicazione: (2025)
Developing Acoustic Models for Automatic Speech Recognition in Swedish
di: Salvi, Giampiero
Pubblicazione: (2024)
di: Salvi, Giampiero
Pubblicazione: (2024)
Fine-Tuning Large Audio-Language Models with LoRA for Precise Temporal Localization of Prolonged Exposure Therapy Elements
di: BN, Suhas, et al.
Pubblicazione: (2025)
di: BN, Suhas, et al.
Pubblicazione: (2025)
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction
di: Cheripally, Sowmya
Pubblicazione: (2024)
di: Cheripally, Sowmya
Pubblicazione: (2024)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
di: Anam, Rizal Khoirul
Pubblicazione: (2025)
di: Anam, Rizal Khoirul
Pubblicazione: (2025)
Matcha-TTS: A fast TTS architecture with conditional flow matching
di: Mehta, Shivam, et al.
Pubblicazione: (2023)
di: Mehta, Shivam, et al.
Pubblicazione: (2023)
Designing Synthetic Discussion Generation Systems: A Case Study for Online Facilitation
di: Tsirmpas, Dimitris, et al.
Pubblicazione: (2025)
di: Tsirmpas, Dimitris, et al.
Pubblicazione: (2025)
Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces
di: Kuhn, Korbinian, et al.
Pubblicazione: (2025)
di: Kuhn, Korbinian, et al.
Pubblicazione: (2025)
Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
di: Lasbordes, Maxence, et al.
Pubblicazione: (2025)
di: Lasbordes, Maxence, et al.
Pubblicazione: (2025)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
di: Otal, Hakan T., et al.
Pubblicazione: (2024)
di: Otal, Hakan T., et al.
Pubblicazione: (2024)
The Transformative Influence of LLMs on Software Development & Developer Productivity
di: Jalil, Sajed
Pubblicazione: (2023)
di: Jalil, Sajed
Pubblicazione: (2023)
Unified speech and gesture synthesis using flow matching
di: Mehta, Shivam, et al.
Pubblicazione: (2023)
di: Mehta, Shivam, et al.
Pubblicazione: (2023)
MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
di: Ali, Zien Sheikh, et al.
Pubblicazione: (2026)
di: Ali, Zien Sheikh, et al.
Pubblicazione: (2026)
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMs
di: Bhatti, Hunzalah Hassan, et al.
Pubblicazione: (2026)
di: Bhatti, Hunzalah Hassan, et al.
Pubblicazione: (2026)
Fake it to make it: Using synthetic data to remedy the data shortage in joint multimodal speech-and-gesture synthesis
di: Mehta, Shivam, et al.
Pubblicazione: (2024)
di: Mehta, Shivam, et al.
Pubblicazione: (2024)
Reviewriter: AI-Generated Instructions For Peer Review Writing
di: Su, Xiaotian, et al.
Pubblicazione: (2025)
di: Su, Xiaotian, et al.
Pubblicazione: (2025)
CognitiveArm: Enabling Real-Time EEG-Controlled Prosthetic Arm Using Embodied Machine Learning
di: Basit, Abdul, et al.
Pubblicazione: (2025)
di: Basit, Abdul, et al.
Pubblicazione: (2025)
Quantization for OpenAI's Whisper Models: A Comparative Analysis
di: Andreyev, Allison
Pubblicazione: (2025)
di: Andreyev, Allison
Pubblicazione: (2025)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
di: Jain, Sarthak, et al.
Pubblicazione: (2024)
di: Jain, Sarthak, et al.
Pubblicazione: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
di: Maben, Leander Melroy, et al.
Pubblicazione: (2025)
di: Maben, Leander Melroy, et al.
Pubblicazione: (2025)
Quantitative Assessment of Intersectional Empathetic Bias and Understanding
di: Formanek, Vojtech, et al.
Pubblicazione: (2024)
di: Formanek, Vojtech, et al.
Pubblicazione: (2024)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
di: Papyan, Narek, et al.
Pubblicazione: (2024)
di: Papyan, Narek, et al.
Pubblicazione: (2024)
Advancing Explainability in Neural Machine Translation: Analytical Metrics for Attention and Alignment Consistency
di: Mishra, Anurag
Pubblicazione: (2024)
di: Mishra, Anurag
Pubblicazione: (2024)
M6(GPT)3: Generating Multitrack Modifiable Multi-Minute MIDI Music from Text using Genetic algorithms, Probabilistic methods and GPT Models in any Progression and Time Signature
di: Poćwiardowski, Jakub, et al.
Pubblicazione: (2024)
di: Poćwiardowski, Jakub, et al.
Pubblicazione: (2024)
Toward Low-Latency End-to-End Voice Agents for Telecommunications Using Streaming ASR, Quantized LLMs, and Real-Time TTS
di: Ethiraj, Vignesh, et al.
Pubblicazione: (2025)
di: Ethiraj, Vignesh, et al.
Pubblicazione: (2025)
NLP-Based Review for Toxic Comment Detection Tailored to the Chinese Cyberspace
di: Ren, Ruixing, et al.
Pubblicazione: (2026)
di: Ren, Ruixing, et al.
Pubblicazione: (2026)
Beyond Speech and More: Investigating the Emergent Ability of Speech Foundation Models for Classifying Physiological Time-Series Signals
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
di: Chen, Kuan-Yu, et al.
Pubblicazione: (2025)
di: Chen, Kuan-Yu, et al.
Pubblicazione: (2025)
Detecting Check-Worthy Claims in Political Debates, Speeches, and Interviews Using Audio Data
di: Ivanov, Petar, et al.
Pubblicazione: (2023)
di: Ivanov, Petar, et al.
Pubblicazione: (2023)
Multilingual Standalone Trustworthy Voice-Based Social Network for Disaster Situations
di: Behravan, Majid, et al.
Pubblicazione: (2024)
di: Behravan, Majid, et al.
Pubblicazione: (2024)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
emg2qwerty: A Large Dataset with Baselines for Touch Typing using Surface Electromyography
di: Sivakumar, Viswanath, et al.
Pubblicazione: (2024)
di: Sivakumar, Viswanath, et al.
Pubblicazione: (2024)
Documenti analoghi
-
UniGlyph: A Seven-Segment Script for Universal Language Representation
di: Sherin, G. V. Bency, et al.
Pubblicazione: (2024) -
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
di: Chen, Zhehuai, et al.
Pubblicazione: (2024) -
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
di: Mehta, Shivam, et al.
Pubblicazione: (2024) -
Communication Access Real-Time Translation Through Collaborative Correction of Automatic Speech Recognition
di: Kuhn, Korbinian, et al.
Pubblicazione: (2025) -
Developing Acoustic Models for Automatic Speech Recognition in Swedish
di: Salvi, Giampiero
Pubblicazione: (2024)