Gespeichert in:
| Hauptverfasser: | Jung, Chaeyoung, Ki, Hojoon, Kim, Ji-Hoon, Kim, Junmo, Chung, Joon Son |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2506.03020 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Probing Cross-modal Information Hubs in Audio-Visual LLMs
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
FxSearcher: gradient-free text-driven audio transformation
von: Ki, Hojoon, et al.
Veröffentlicht: (2025)
von: Ki, Hojoon, et al.
Veröffentlicht: (2025)
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2025)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2025)
Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting
von: Kim, Youkyum, et al.
Veröffentlicht: (2024)
von: Kim, Youkyum, et al.
Veröffentlicht: (2024)
TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2025)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2025)
Lightweight Audio Segmentation for Long-form Speech Translation
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
von: Kim, Jongsuk, et al.
Veröffentlicht: (2025)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2025)
ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2026)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2026)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2023)
SEED: Speaker Embedding Enhancement Diffusion Model
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
von: Jung, Jihoo, et al.
Veröffentlicht: (2026)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
SCORE: Scaling audio generation using Standardized COmposite REwards
von: Jung, Jaemin, et al.
Veröffentlicht: (2025)
von: Jung, Jaemin, et al.
Veröffentlicht: (2025)
VoxSim: A perceptual voice similarity dataset
von: Ahn, Junseok, et al.
Veröffentlicht: (2024)
von: Ahn, Junseok, et al.
Veröffentlicht: (2024)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
From Coarse to Fine: Efficient Training for Audio Spectrogram Transformers
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
MATE: Matryoshka Audio-Text Embeddings for Open-Vocabulary Keyword Spotting
von: Jung, Youngmoon, et al.
Veröffentlicht: (2026)
von: Jung, Youngmoon, et al.
Veröffentlicht: (2026)
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2024)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2024)
Adapting a Text-to-Audio Model for Room Impulse Response Generation
von: Kim, Kirak, et al.
Veröffentlicht: (2026)
von: Kim, Kirak, et al.
Veröffentlicht: (2026)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
AdaptVC: High Quality Voice Conversion with Adaptive Learning
von: Kim, Jaehun, et al.
Veröffentlicht: (2025)
von: Kim, Jaehun, et al.
Veröffentlicht: (2025)
Relational Proxy Loss for Audio-Text based Keyword Spotting
von: Jung, Youngmoon, et al.
Veröffentlicht: (2024)
von: Jung, Youngmoon, et al.
Veröffentlicht: (2024)
Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
von: Chung, HaeChun
Veröffentlicht: (2025)
von: Chung, HaeChun
Veröffentlicht: (2025)
Ähnliche Einträge
-
Probing Cross-modal Information Hubs in Audio-Visual LLMs
von: Jung, Jihoo, et al.
Veröffentlicht: (2026) -
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024) -
FxSearcher: gradient-free text-driven audio transformation
von: Ki, Hojoon, et al.
Veröffentlicht: (2025) -
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2025) -
Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting
von: Kim, Youkyum, et al.
Veröffentlicht: (2024)