Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Meng, Chutong, Koehn, Philipp |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task
by: Meng, Chutong, et al.
Published: (2025)
by: Meng, Chutong, et al.
Published: (2025)
DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation
by: Tan, Weiting, et al.
Published: (2024)
by: Tan, Weiting, et al.
Published: (2024)
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
by: Zhao, Xiutian, et al.
Published: (2026)
by: Zhao, Xiutian, et al.
Published: (2026)
SpeechAlign: Aligning Speech Generation to Human Preferences
by: Zhang, Dong, et al.
Published: (2024)
by: Zhang, Dong, et al.
Published: (2024)
SpeechAlign: a Framework for Speech Translation Alignment Evaluation
by: Alastruey, Belen, et al.
Published: (2023)
by: Alastruey, Belen, et al.
Published: (2023)
Text Style Transfer with Parameter-efficient LLM Finetuning and Round-trip Translation
by: Liu, Ruoxi, et al.
Published: (2026)
by: Liu, Ruoxi, et al.
Published: (2026)
SpeechMapper: Speech-to-text Embedding Projector for LLMs
by: Mohapatra, Biswesh, et al.
Published: (2026)
by: Mohapatra, Biswesh, et al.
Published: (2026)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
by: Tan, Weiting, et al.
Published: (2025)
by: Tan, Weiting, et al.
Published: (2025)
Textless Speech-to-Speech Translation With Limited Parallel Data
by: Diwan, Anuj, et al.
Published: (2023)
by: Diwan, Anuj, et al.
Published: (2023)
RosettaSpeech: Zero-Shot Speech-to-Speech Translation without Parallel Speech
by: Zheng, Zhisheng, et al.
Published: (2025)
by: Zheng, Zhisheng, et al.
Published: (2025)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
by: Tseng, Liang-Hsuan, et al.
Published: (2025)
by: Tseng, Liang-Hsuan, et al.
Published: (2025)
Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs
by: Mukherjee, Anjishnu, et al.
Published: (2026)
by: Mukherjee, Anjishnu, et al.
Published: (2026)
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
Simultaneous Speech-to-Speech Translation Without Aligned Data
by: Labiausse, Tom, et al.
Published: (2026)
by: Labiausse, Tom, et al.
Published: (2026)
Connecting Voices: LoReSpeech as a Low-Resource Speech Parallel Corpus
by: Ouzerrout, Samy
Published: (2025)
by: Ouzerrout, Samy
Published: (2025)
TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
by: Tseng, Liang-Hsuan, et al.
Published: (2026)
by: Tseng, Liang-Hsuan, et al.
Published: (2026)
Pointer-Generator Networks for Low-Resource Machine Translation: Don't Copy That!
by: Bafna, Niyati, et al.
Published: (2024)
by: Bafna, Niyati, et al.
Published: (2024)
Recovering document annotations for sentence-level bitext
by: Wicks, Rachel, et al.
Published: (2024)
by: Wicks, Rachel, et al.
Published: (2024)
DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
by: Tan, Chao-Hong, et al.
Published: (2025)
by: Tan, Chao-Hong, et al.
Published: (2025)
AlignCap: Aligning Speech Emotion Captioning to Human Preferences
by: Liang, Ziqi, et al.
Published: (2024)
by: Liang, Ziqi, et al.
Published: (2024)
Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects
by: Chang, Kalvin, et al.
Published: (2026)
by: Chang, Kalvin, et al.
Published: (2026)
Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically
by: Shim, Ryan Soh-Eun, et al.
Published: (2025)
by: Shim, Ryan Soh-Eun, et al.
Published: (2025)
SPACER: A Parallel Dataset of Speech Production And Comprehension of Error Repairs
by: Upadhye, Shiva, et al.
Published: (2025)
by: Upadhye, Shiva, et al.
Published: (2025)
SinFoS: A Parallel Dataset for Translating Sinhala Figures of Speech
by: Sofalas, Johan, et al.
Published: (2026)
by: Sofalas, Johan, et al.
Published: (2026)
WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
by: Zhang, Binbin, et al.
Published: (2025)
by: Zhang, Binbin, et al.
Published: (2025)
Instituto de Telecomunicações at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning
by: Attanasio, Giuseppe, et al.
Published: (2025)
by: Attanasio, Giuseppe, et al.
Published: (2025)
RepCodec: A Speech Representation Codec for Speech Tokenization
by: Huang, Zhichao, et al.
Published: (2023)
by: Huang, Zhichao, et al.
Published: (2023)
Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
by: Rashidi, Sina, et al.
Published: (2025)
by: Rashidi, Sina, et al.
Published: (2025)
Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
by: Azad, Asif, et al.
Published: (2026)
by: Azad, Asif, et al.
Published: (2026)
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving
by: Shankar, Bhavani, et al.
Published: (2024)
by: Shankar, Bhavani, et al.
Published: (2024)
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
by: Spohn, Philipp, et al.
Published: (2025)
by: Spohn, Philipp, et al.
Published: (2025)
FASA: a Flexible and Automatic Speech Aligner for Extracting High-quality Aligned Children Speech Data
by: Liu, Dancheng, et al.
Published: (2024)
by: Liu, Dancheng, et al.
Published: (2024)
On Importance of Code-Mixed Embeddings for Hate Speech Identification
by: Jagdale, Shruti, et al.
Published: (2024)
by: Jagdale, Shruti, et al.
Published: (2024)
Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
by: Wang, Yuejiao, et al.
Published: (2024)
by: Wang, Yuejiao, et al.
Published: (2024)
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
by: Akti, Seymanur, et al.
Published: (2026)
by: Akti, Seymanur, et al.
Published: (2026)
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
by: Eilertsen, Brage, et al.
Published: (2025)
by: Eilertsen, Brage, et al.
Published: (2025)
SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models
by: Yang, Wanqi, et al.
Published: (2025)
by: Yang, Wanqi, et al.
Published: (2025)
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
by: Fang, Qingkai, et al.
Published: (2024)
by: Fang, Qingkai, et al.
Published: (2024)
SEAL: Speech Embedding Alignment Learning for Speech Large Language Model with Retrieval-Augmented Generation
by: Sun, Chunyu, et al.
Published: (2025)
by: Sun, Chunyu, et al.
Published: (2025)
Similar Items
-
GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task
by: Meng, Chutong, et al.
Published: (2025) -
DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation
by: Tan, Weiting, et al.
Published: (2024) -
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
by: Zhao, Xiutian, et al.
Published: (2026) -
SpeechAlign: Aligning Speech Generation to Human Preferences
by: Zhang, Dong, et al.
Published: (2024) -
SpeechAlign: a Framework for Speech Translation Alignment Evaluation
by: Alastruey, Belen, et al.
Published: (2023)