Learning Multiple Utterance-Level Attribute Representations with a Unified Speech Encoder
Fuente:
arXiv
Saved in:
| Main Authors: | Bouziane, Maryem, Mdhaffar, Salima, Estève, Yannick |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
by: Mdhaffar, Salima, et al.
Published: (2024)
by: Mdhaffar, Salima, et al.
Published: (2024)
TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English
by: Bougares, Fethi, et al.
Published: (2025)
by: Bougares, Fethi, et al.
Published: (2025)
Using Multimodal and Language-Agnostic Sentence Embeddings for Abstractive Summarization
by: Chellaf, Chaimae, et al.
Published: (2026)
by: Chellaf, Chaimae, et al.
Published: (2026)
SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding
by: Elleuch, Haroun, et al.
Published: (2026)
by: Elleuch, Haroun, et al.
Published: (2026)
ADI-20: Arabic Dialect Identification dataset and models
by: Elleuch, Haroun, et al.
Published: (2025)
by: Elleuch, Haroun, et al.
Published: (2025)
Ara-Best-RQ: Multi Dialectal Arabic SSL
by: Elleuch, Haroun, et al.
Published: (2026)
by: Elleuch, Haroun, et al.
Published: (2026)
ELYADATA & LIA at NADI 2025: ASR and ADI Subtasks
by: Elleuch, Haroun, et al.
Published: (2025)
by: Elleuch, Haroun, et al.
Published: (2025)
SENSE models: an open source solution for multilingual and multimodal semantic-based tasks
by: Mdhaffar, Salima, et al.
Published: (2025)
by: Mdhaffar, Salima, et al.
Published: (2025)
Pantagruel: Unified Self-Supervised Encoders for French Text and Speech
by: Le, Phuong-Hang, et al.
Published: (2026)
by: Le, Phuong-Hang, et al.
Published: (2026)
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants
by: Sekkat, Chloé, et al.
Published: (2024)
by: Sekkat, Chloé, et al.
Published: (2024)
In-domain SSL pre-training and streaming ASR
by: Duret, Jarod, et al.
Published: (2025)
by: Duret, Jarod, et al.
Published: (2025)
An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding
by: Su, Tianhui, et al.
Published: (2026)
by: Su, Tianhui, et al.
Published: (2026)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
by: Duret, Jarod, et al.
Published: (2024)
by: Duret, Jarod, et al.
Published: (2024)
Open Implementation and Study of BEST-RQ for Speech Processing
by: Whetten, Ryan, et al.
Published: (2024)
by: Whetten, Ryan, et al.
Published: (2024)
Utterance-Level Methods for Identifying Reliable ASR-Output for Child Speech
by: Lathouwers, Gus, et al.
Published: (2026)
by: Lathouwers, Gus, et al.
Published: (2026)
LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
by: Parcollet, Titouan, et al.
Published: (2023)
by: Parcollet, Titouan, et al.
Published: (2023)
Simultaneous Speech-to-Speech Translation Without Aligned Data
by: Labiausse, Tom, et al.
Published: (2026)
by: Labiausse, Tom, et al.
Published: (2026)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
by: Huang, Kuan-Po, et al.
Published: (2023)
by: Huang, Kuan-Po, et al.
Published: (2023)
Learning Multiplex Representations on Text-Attributed Graphs with One Language Model Encoder
by: Jin, Bowen, et al.
Published: (2023)
by: Jin, Bowen, et al.
Published: (2023)
Cross-Utterance Conditioned VAE for Speech Generation
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
ViToSA: Audio-Based Toxic Spans Detection on Vietnamese Speech Utterances
by: Do, Huy Ba, et al.
Published: (2025)
by: Do, Huy Ba, et al.
Published: (2025)
In-Context Learning with Reinforcement Learning for Incomplete Utterance Rewriting
by: Du, Haowei, et al.
Published: (2024)
by: Du, Haowei, et al.
Published: (2024)
Incomplete Utterance Rewriting with Editing Operation Guidance and Utterance Augmentation
by: Cao, Zhiyu, et al.
Published: (2025)
by: Cao, Zhiyu, et al.
Published: (2025)
Towards Early Prediction of Self-Supervised Speech Model Performance
by: Whetten, Ryan, et al.
Published: (2025)
by: Whetten, Ryan, et al.
Published: (2025)
KULCQ: An Unsupervised Keyword-based Utterance Level Clustering Quality Metric
by: Guruprasad, Pranav, et al.
Published: (2024)
by: Guruprasad, Pranav, et al.
Published: (2024)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
by: Zhang, Enshi, et al.
Published: (2024)
by: Zhang, Enshi, et al.
Published: (2024)
Investigating Low-Cost LLM Annotation for~Spoken Dialogue Understanding Datasets
by: Druart, Lucas, et al.
Published: (2024)
by: Druart, Lucas, et al.
Published: (2024)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
by: Zink, Oswald, et al.
Published: (2024)
by: Zink, Oswald, et al.
Published: (2024)
Natural Language Decomposition and Interpretation of Complex Utterances
by: Jhamtani, Harsh, et al.
Published: (2023)
by: Jhamtani, Harsh, et al.
Published: (2023)
AEyeDE: An Attention-Based Attribution Framework for AI-Generated Text Detection
by: Nourbakhsh, Aria, et al.
Published: (2026)
by: Nourbakhsh, Aria, et al.
Published: (2026)
From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality
by: Kim, Yoonsang, et al.
Published: (2026)
by: Kim, Yoonsang, et al.
Published: (2026)
Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models
by: Kang, Jaehoon, et al.
Published: (2026)
by: Kang, Jaehoon, et al.
Published: (2026)
A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
by: Laperrière, Gaëlle, et al.
Published: (2024)
by: Laperrière, Gaëlle, et al.
Published: (2024)
Evaluating Explainable AI Attribution Methods in Neural Machine Translation via Attention-Guided Knowledge Distillation
by: Nourbakhsh, Aria, et al.
Published: (2026)
by: Nourbakhsh, Aria, et al.
Published: (2026)
Is one brick enough to break the wall of spoken dialogue state tracking?
by: Druart, Lucas, et al.
Published: (2023)
by: Druart, Lucas, et al.
Published: (2023)
TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment
by: Kim, Taesoo, et al.
Published: (2025)
by: Kim, Taesoo, et al.
Published: (2025)
Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic Parsing
by: Kang, Deokhyung, et al.
Published: (2024)
by: Kang, Deokhyung, et al.
Published: (2024)
SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition
by: Wu, Yihan, et al.
Published: (2024)
by: Wu, Yihan, et al.
Published: (2024)
Segment-Level Attribution for Selective Learning of Long Reasoning Traces
by: Wang, Siyuan, et al.
Published: (2026)
by: Wang, Siyuan, et al.
Published: (2026)
Is Information Density Uniform when Utterances are Grounded on Perception and Discourse?
by: Gay, Matteo, et al.
Published: (2026)
by: Gay, Matteo, et al.
Published: (2026)
Similar Items
-
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
by: Mdhaffar, Salima, et al.
Published: (2024) -
TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English
by: Bougares, Fethi, et al.
Published: (2025) -
Using Multimodal and Language-Agnostic Sentence Embeddings for Abstractive Summarization
by: Chellaf, Chaimae, et al.
Published: (2026) -
SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding
by: Elleuch, Haroun, et al.
Published: (2026) -
ADI-20: Arabic Dialect Identification dataset and models
by: Elleuch, Haroun, et al.
Published: (2025)