SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Buess, Lukas, Geier, Jan, Bani-Harouni, David, Pellegrini, Chantal, Keicher, Matthias, Perez-Toro, Paula Andrea, Navab, Nassir, Maier, Andreas, Arias-Vergara, Tomas |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
EHR2Path: Scalable Modeling of Longitudinal Patient Pathways from Multimodal Electronic Health Records
par: Pellegrini, Chantal, et autres
Publié: (2025)
par: Pellegrini, Chantal, et autres
Publié: (2025)
Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning
par: Bani-Harouni, David, et autres
Publié: (2025)
par: Bani-Harouni, David, et autres
Publié: (2025)
MAGDA: Multi-agent guideline-driven diagnostic assistance
par: Bani-Harouni, David, et autres
Publié: (2024)
par: Bani-Harouni, David, et autres
Publié: (2024)
Specialized Foundation Models for Intelligent Operating Rooms
par: Özsoy, Ege, et autres
Publié: (2025)
par: Özsoy, Ege, et autres
Publié: (2025)
Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models
par: Bani-Harouni, David, et autres
Publié: (2025)
par: Bani-Harouni, David, et autres
Publié: (2025)
From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
par: Buess, Lukas, et autres
Publié: (2025)
par: Buess, Lukas, et autres
Publié: (2025)
ORacle: Large Vision-Language Models for Knowledge-Guided Holistic OR Domain Modeling
par: Özsoy, Ege, et autres
Publié: (2024)
par: Özsoy, Ege, et autres
Publié: (2024)
MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments
par: Özsoy, Ege, et autres
Publié: (2025)
par: Özsoy, Ege, et autres
Publié: (2025)
Learning Diagnostic Reasoning for Decision Support in Toxicology
par: Oberländer, Nico, et autres
Publié: (2026)
par: Oberländer, Nico, et autres
Publié: (2026)
The Impact of Speech Anonymization on Pathology and Its Limits
par: Arasteh, Soroosh Tayebi, et autres
Publié: (2024)
par: Arasteh, Soroosh Tayebi, et autres
Publié: (2024)
Detecting Spoof Voices in Asian Non-Native Speech: An Indonesian and Thai Case Study
par: Adila, Aulia, et autres
Publié: (2024)
par: Adila, Aulia, et autres
Publié: (2024)
Calibrated Confidence Expression for Radiology Report Generation
par: Bani-Harouni, David, et autres
Publié: (2026)
par: Bani-Harouni, David, et autres
Publié: (2026)
Prototype-Based Knowledge Guidance for Fine-Grained Structured Radiology Reporting
par: Pellegrini, Chantal, et autres
Publié: (2026)
par: Pellegrini, Chantal, et autres
Publié: (2026)
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
par: Pellegrini, Chantal, et autres
Publié: (2023)
par: Pellegrini, Chantal, et autres
Publié: (2023)
VoiceX: A Text-To-Speech Framework for Custom Voices
par: Mertes, Silvan, et autres
Publié: (2024)
par: Mertes, Silvan, et autres
Publié: (2024)
Speech to Speech Synthesis for Voice Impersonation
par: Johnson, Bjorn, et autres
Publié: (2026)
par: Johnson, Bjorn, et autres
Publié: (2026)
LearnerVoice: A Dataset of Non-Native English Learners' Spontaneous Speech
par: Kim, Haechan, et autres
Publié: (2024)
par: Kim, Haechan, et autres
Publié: (2024)
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
par: Kim, Minsu, et autres
Publié: (2025)
par: Kim, Minsu, et autres
Publié: (2025)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
par: Park, Nohil, et autres
Publié: (2024)
par: Park, Nohil, et autres
Publié: (2024)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
par: Peng, Puyuan, et autres
Publié: (2024)
par: Peng, Puyuan, et autres
Publié: (2024)
VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
par: Kim, Heeseung, et autres
Publié: (2024)
par: Kim, Heeseung, et autres
Publié: (2024)
A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings
par: Mondal, Anindita, et autres
Publié: (2024)
par: Mondal, Anindita, et autres
Publié: (2024)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
par: Geng, Haopeng, et autres
Publié: (2024)
par: Geng, Haopeng, et autres
Publié: (2024)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
par: Zheng, Zhisheng, et autres
Publié: (2025)
par: Zheng, Zhisheng, et autres
Publié: (2025)
Enhancing Multilingual Voice Toxicity Detection with Speech-Text Alignment
par: Liu, Joseph, et autres
Publié: (2024)
par: Liu, Joseph, et autres
Publié: (2024)
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
par: Peng, Yifan, et autres
Publié: (2024)
par: Peng, Yifan, et autres
Publié: (2024)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
par: Zhu, Han, et autres
Publié: (2025)
par: Zhu, Han, et autres
Publié: (2025)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
par: Gudmalwar, Ashishkumar, et autres
Publié: (2024)
par: Gudmalwar, Ashishkumar, et autres
Publié: (2024)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
par: Yang, Yifan, et autres
Publié: (2024)
par: Yang, Yifan, et autres
Publié: (2024)
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
par: Sahipjohn, Neha, et autres
Publié: (2024)
par: Sahipjohn, Neha, et autres
Publié: (2024)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
par: Kuhlmann, Michael, et autres
Publié: (2026)
par: Kuhlmann, Michael, et autres
Publié: (2026)
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
par: Ma, Linhan, et autres
Publié: (2025)
par: Ma, Linhan, et autres
Publié: (2025)
Robust Speech Activity Detection in the Presence of Singing Voice
par: Grundhuber, Philipp, et autres
Publié: (2025)
par: Grundhuber, Philipp, et autres
Publié: (2025)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
par: Huo, Mingyue, et autres
Publié: (2025)
par: Huo, Mingyue, et autres
Publié: (2025)
Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment
par: Cai, Huanchen, et autres
Publié: (2026)
par: Cai, Huanchen, et autres
Publié: (2026)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
par: Anastassiou, Philip, et autres
Publié: (2024)
par: Anastassiou, Philip, et autres
Publié: (2024)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
par: Yang, Guanrou, et autres
Publié: (2025)
par: Yang, Guanrou, et autres
Publié: (2025)
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
par: Zhu, Han, et autres
Publié: (2026)
par: Zhu, Han, et autres
Publié: (2026)
Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
par: Lau, Hok-Shing, et autres
Publié: (2024)
par: Lau, Hok-Shing, et autres
Publié: (2024)
DM-Codec: Distilling Multimodal Representations for Speech Tokenization
par: Ahasan, Md Mubtasim, et autres
Publié: (2024)
par: Ahasan, Md Mubtasim, et autres
Publié: (2024)
Documents similaires
-
EHR2Path: Scalable Modeling of Longitudinal Patient Pathways from Multimodal Electronic Health Records
par: Pellegrini, Chantal, et autres
Publié: (2025) -
Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning
par: Bani-Harouni, David, et autres
Publié: (2025) -
MAGDA: Multi-agent guideline-driven diagnostic assistance
par: Bani-Harouni, David, et autres
Publié: (2024) -
Specialized Foundation Models for Intelligent Operating Rooms
par: Özsoy, Ege, et autres
Publié: (2025) -
Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models
par: Bani-Harouni, David, et autres
Publié: (2025)