Sequence-to-sequence models in peer-to-peer learning: A practical application
Fuente:
arXiv
Guardado en:
| Autores principales: | Šajina, Robert, Ipšić, Ivo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
por: Wan, Zhen, et al.
Publicado: (2026)
por: Wan, Zhen, et al.
Publicado: (2026)
Spoken Conversational Agents with Large Language Models
por: Yang, Chao-Han Huck, et al.
Publicado: (2025)
por: Yang, Chao-Han Huck, et al.
Publicado: (2025)
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation
por: Rong, Yan, et al.
Publicado: (2025)
por: Rong, Yan, et al.
Publicado: (2025)
A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration
por: Selvamani, Shaja Arul, et al.
Publicado: (2025)
por: Selvamani, Shaja Arul, et al.
Publicado: (2025)
PodAgent: A Comprehensive Framework for Podcast Generation
por: Xiao, Yujia, et al.
Publicado: (2025)
por: Xiao, Yujia, et al.
Publicado: (2025)
TuneGenie: Reasoning-based LLM agents for preferential music generation
por: Pandey, Amitesh, et al.
Publicado: (2025)
por: Pandey, Amitesh, et al.
Publicado: (2025)
Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation
por: Zhou, Jinxing, et al.
Publicado: (2025)
por: Zhou, Jinxing, et al.
Publicado: (2025)
Musical Agent Systems: MACAT and MACataRT
por: Lee, Keon Ju M., et al.
Publicado: (2025)
por: Lee, Keon Ju M., et al.
Publicado: (2025)
A Pilot Study of Applying Sequence-to-Sequence Voice Conversion to Evaluate the Intelligibility of L2 Speech Using a Native Speaker's Shadowings
por: Geng, Haopeng, et al.
Publicado: (2024)
por: Geng, Haopeng, et al.
Publicado: (2024)
Streaming Sequence Transduction through Dynamic Compression
por: Tan, Weiting, et al.
Publicado: (2024)
por: Tan, Weiting, et al.
Publicado: (2024)
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
por: Liu, Oli Danyi, et al.
Publicado: (2024)
por: Liu, Oli Danyi, et al.
Publicado: (2024)
Revival: Collaborative Artistic Creation through Human-AI Interactions in Musical Creativity
por: Lee, Keon Ju M., et al.
Publicado: (2025)
por: Lee, Keon Ju M., et al.
Publicado: (2025)
Towards continually learning new languages
por: Pham, Ngoc-Quan, et al.
Publicado: (2022)
por: Pham, Ngoc-Quan, et al.
Publicado: (2022)
LastResort at SemEval-2024 Task 3: Exploring Multimodal Emotion Cause Pair Extraction as Sequence Labelling Task
por: Mathur, Suyash Vardhan, et al.
Publicado: (2024)
por: Mathur, Suyash Vardhan, et al.
Publicado: (2024)
Revisiting speech segmentation and lexicon learning with better features
por: Kamper, Herman, et al.
Publicado: (2024)
por: Kamper, Herman, et al.
Publicado: (2024)
Can Whisper perform speech-based in-context learning?
por: Wang, Siyin, et al.
Publicado: (2023)
por: Wang, Siyin, et al.
Publicado: (2023)
A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
por: Laperrière, Gaëlle, et al.
Publicado: (2024)
por: Laperrière, Gaëlle, et al.
Publicado: (2024)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
por: Slabbert, Danel, et al.
Publicado: (2025)
por: Slabbert, Danel, et al.
Publicado: (2025)
ZIPA: A family of efficient models for multilingual phone recognition
por: Zhu, Jian, et al.
Publicado: (2025)
por: Zhu, Jian, et al.
Publicado: (2025)
XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection
por: Dat, Phuong Tuan, et al.
Publicado: (2025)
por: Dat, Phuong Tuan, et al.
Publicado: (2025)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
por: Wang, Hsuan-Fu, et al.
Publicado: (2024)
por: Wang, Hsuan-Fu, et al.
Publicado: (2024)
How Much Context Does My Attention-Based ASR System Need?
por: Flynn, Robert, et al.
Publicado: (2023)
por: Flynn, Robert, et al.
Publicado: (2023)
AfriHuBERT: A self-supervised speech representation model for African languages
por: Alabi, Jesujoba O., et al.
Publicado: (2024)
por: Alabi, Jesujoba O., et al.
Publicado: (2024)
Word-wise intonation model for cross-language TTS systems
por: A., Tomilov A., et al.
Publicado: (2024)
por: A., Tomilov A., et al.
Publicado: (2024)
Transferable speech-to-text large language model alignment module
por: Wu, Boyong, et al.
Publicado: (2024)
por: Wu, Boyong, et al.
Publicado: (2024)
Towards a dynamical model of English vowels. Evidence from diphthongisation
por: Strycharczuk, Patrycja, et al.
Publicado: (2024)
por: Strycharczuk, Patrycja, et al.
Publicado: (2024)
Robustness assessment of large audio language models in multiple-choice evaluation
por: López, Fernando, et al.
Publicado: (2025)
por: López, Fernando, et al.
Publicado: (2025)
Exploring the limits of decoder-only models trained on public speech recognition corpora
por: Gupta, Ankit, et al.
Publicado: (2024)
por: Gupta, Ankit, et al.
Publicado: (2024)
Asymmetric and trial-dependent modeling: the contribution of LIA to SdSV Challenge Task 2
por: Bousquet, Pierre-Michel, et al.
Publicado: (2024)
por: Bousquet, Pierre-Michel, et al.
Publicado: (2024)
Streaming Bilingual End-to-End ASR model using Attention over Multiple Softmax
por: Patil, Aditya, et al.
Publicado: (2024)
por: Patil, Aditya, et al.
Publicado: (2024)
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
por: Sanchez, Ariadna, et al.
Publicado: (2025)
por: Sanchez, Ariadna, et al.
Publicado: (2025)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
por: Paraskevopoulos, Georgios, et al.
Publicado: (2024)
por: Paraskevopoulos, Georgios, et al.
Publicado: (2024)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
por: Wright, George August, et al.
Publicado: (2023)
por: Wright, George August, et al.
Publicado: (2023)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
por: Gowda, Harshavardhana T., et al.
Publicado: (2025)
por: Gowda, Harshavardhana T., et al.
Publicado: (2025)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
por: Kesiraju, Santosh, et al.
Publicado: (2023)
por: Kesiraju, Santosh, et al.
Publicado: (2023)
Linguists should learn to love speech-based deep learning models
por: Kloots, Marianne de Heer, et al.
Publicado: (2025)
por: Kloots, Marianne de Heer, et al.
Publicado: (2025)
PRODIS -- a speech database and a phoneme-based language model for the study of predictability effects in Polish
por: Malisz, Zofia, et al.
Publicado: (2024)
por: Malisz, Zofia, et al.
Publicado: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
por: Okocha, Chibuzor, et al.
Publicado: (2025)
por: Okocha, Chibuzor, et al.
Publicado: (2025)
Establishing degrees of closeness between audio recordings along different dimensions using large-scale cross-lingual models
por: Fily, Maxime, et al.
Publicado: (2024)
por: Fily, Maxime, et al.
Publicado: (2024)
A Benchmark for Multi-speaker Anonymization
por: Miao, Xiaoxiao, et al.
Publicado: (2024)
por: Miao, Xiaoxiao, et al.
Publicado: (2024)
Ejemplares similares
-
Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
por: Wan, Zhen, et al.
Publicado: (2026) -
Spoken Conversational Agents with Large Language Models
por: Yang, Chao-Han Huck, et al.
Publicado: (2025) -
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation
por: Rong, Yan, et al.
Publicado: (2025) -
A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration
por: Selvamani, Shaja Arul, et al.
Publicado: (2025) -
PodAgent: A Comprehensive Framework for Podcast Generation
por: Xiao, Yujia, et al.
Publicado: (2025)