Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models
Fuente:
arXiv
Saved in:
| Main Authors: | Nakai, Toshiki, Suresh, Varsha, Demberg, Vera |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation Framework
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Synthetic Data Augmentation for Cross-domain Implicit Discourse Relation Recognition
by: Yung, Frances, et al.
Published: (2025)
by: Yung, Frances, et al.
Published: (2025)
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues
by: Suresh, Varsha, et al.
Published: (2025)
by: Suresh, Varsha, et al.
Published: (2025)
Modeling Turn-Taking with Semantically Informed Gestures
by: Suresh, Varsha, et al.
Published: (2025)
by: Suresh, Varsha, et al.
Published: (2025)
MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
by: Saha, Anisha, et al.
Published: (2025)
by: Saha, Anisha, et al.
Published: (2025)
System-Mediated Attention Imbalances Make Vision-Language Models Say Yes
by: Chan, Tsan Tsai, et al.
Published: (2026)
by: Chan, Tsan Tsai, et al.
Published: (2026)
ChatGPT vs Human-authored Text: Insights into Controllable Text Summarization and Sentence Style Transfer
by: Liu, Dongqi, et al.
Published: (2023)
by: Liu, Dongqi, et al.
Published: (2023)
What Language(s) Does Aya-23 Think In? How Multilinguality Affects Internal Language Representations
by: Trinley, Katharina, et al.
Published: (2025)
by: Trinley, Katharina, et al.
Published: (2025)
Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures
by: Suresh, Varsha, et al.
Published: (2026)
by: Suresh, Varsha, et al.
Published: (2026)
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization
by: Saha, Anisha, et al.
Published: (2026)
by: Saha, Anisha, et al.
Published: (2026)
Human Speech Perception in Noise: Can Large Language Models Paraphrase to Improve It?
by: Chingacham, Anupama, et al.
Published: (2024)
by: Chingacham, Anupama, et al.
Published: (2024)
Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
by: Omnilingual SONAR Team, et al.
Published: (2026)
by: Omnilingual SONAR Team, et al.
Published: (2026)
RST-LoRA: A Discourse-Aware Low-Rank Adaptation for Long Document Abstractive Summarization
by: Liu, Dongqi, et al.
Published: (2024)
by: Liu, Dongqi, et al.
Published: (2024)
How do Multimodal Foundation Models Encode Text and Speech? An Analysis of Cross-Lingual and Cross-Modal Representations
by: Lee, Hyunji, et al.
Published: (2024)
by: Lee, Hyunji, et al.
Published: (2024)
On Crowdsourcing Task Design for Discourse Relation Annotation
by: Yung, Frances, et al.
Published: (2024)
by: Yung, Frances, et al.
Published: (2024)
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
by: Lou, Haowei, et al.
Published: (2025)
by: Lou, Haowei, et al.
Published: (2025)
Is Cross-Lingual Transfer in Bilingual Models Human-Like? A Study with Overlapping Word Forms in Dutch and English
by: Škrjanec, Iza, et al.
Published: (2026)
by: Škrjanec, Iza, et al.
Published: (2026)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
by: Zhang, Pei, et al.
Published: (2025)
by: Zhang, Pei, et al.
Published: (2025)
Implicit Discourse Relation Classification For Nigerian Pidgin
by: Saeed, Muhammed, et al.
Published: (2024)
by: Saeed, Muhammed, et al.
Published: (2024)
LLMs syntactically adapt their language use to their conversational partner
by: Kandra, Florian, et al.
Published: (2025)
by: Kandra, Florian, et al.
Published: (2025)
Planning Ahead with RSA: Efficient Signalling in Dynamic Environments by Projecting User Awareness across Future Timesteps
by: Das, Anwesha, et al.
Published: (2025)
by: Das, Anwesha, et al.
Published: (2025)
CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation
by: Du, Yexing, et al.
Published: (2025)
by: Du, Yexing, et al.
Published: (2025)
Cross-lingual Matryoshka Representation Learning across Speech and Text
by: Sy, Yaya, et al.
Published: (2026)
by: Sy, Yaya, et al.
Published: (2026)
Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"?
by: Liu, Tong, et al.
Published: (2023)
by: Liu, Tong, et al.
Published: (2023)
Cross-Modal Robustness Transfer (CMRT): Training Robust Speech Translation Models Using Adversarial Text
by: Issam, Abderrahmane, et al.
Published: (2026)
by: Issam, Abderrahmane, et al.
Published: (2026)
Pragmatic Reasoning improves LLM Code Generation
by: Cao, Zhuchen, et al.
Published: (2025)
by: Cao, Zhuchen, et al.
Published: (2025)
mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
by: Singh, Jaskaran, et al.
Published: (2025)
by: Singh, Jaskaran, et al.
Published: (2025)
Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs
by: Wang, Zhizhi, et al.
Published: (2026)
by: Wang, Zhizhi, et al.
Published: (2026)
Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders
by: Mazumder, Debajyoti, et al.
Published: (2026)
by: Mazumder, Debajyoti, et al.
Published: (2026)
AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
by: Yang, Bo, et al.
Published: (2025)
by: Yang, Bo, et al.
Published: (2025)
Incorporating Distributions of Discourse Structure for Long Document Abstractive Summarization
by: Liu, Dongqi, et al.
Published: (2023)
by: Liu, Dongqi, et al.
Published: (2023)
SciNews: From Scholarly Complexities to Public Narratives -- A Dataset for Scientific News Report Generation
by: Liu, Dongqi, et al.
Published: (2024)
by: Liu, Dongqi, et al.
Published: (2024)
Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin
by: Lin, Pin-Jie, et al.
Published: (2024)
by: Lin, Pin-Jie, et al.
Published: (2024)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
by: Li, Xuanchen, et al.
Published: (2025)
by: Li, Xuanchen, et al.
Published: (2025)
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
GestureCoach: Rehearsing for Engaging Talks with LLM-Driven Gesture Recommendations
by: Ram, Ashwin, et al.
Published: (2025)
by: Ram, Ashwin, et al.
Published: (2025)
The Design of Informative Take-Over Requests for Semi-Autonomous Cyber-Physical Systems: Combining Spoken Language and Visual Icons in a Drone-Controller Setting
by: Gundappa, Ashwini, et al.
Published: (2024)
by: Gundappa, Ashwini, et al.
Published: (2024)
Multilingual Extraction and Recognition of Implicit Discourse Relations in Speech and Text
by: Ruby, Ahmed, et al.
Published: (2026)
by: Ruby, Ahmed, et al.
Published: (2026)
Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion
by: Du, Yexing, et al.
Published: (2026)
by: Du, Yexing, et al.
Published: (2026)
Similar Items
-
RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation Framework
by: Wang, Yifan, et al.
Published: (2024) -
Synthetic Data Augmentation for Cross-domain Implicit Discourse Relation Recognition
by: Yung, Frances, et al.
Published: (2025) -
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues
by: Suresh, Varsha, et al.
Published: (2025) -
Modeling Turn-Taking with Semantically Informed Gestures
by: Suresh, Varsha, et al.
Published: (2025) -
MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
by: Saha, Anisha, et al.
Published: (2025)