Saved in:
| Main Authors: | R, Ezhini Rasendiran, Maurya, Chandresh Kumar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.18334 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automated Classification of Phonetic Segments in Child Speech Using Raw Ultrasound Imaging
by: Ani, Saja Al, et al.
Published: (2024)
by: Ani, Saja Al, et al.
Published: (2024)
End-to-End Speech-to-Text Translation: A Survey
by: Sethiya, Nivedita, et al.
Published: (2023)
by: Sethiya, Nivedita, et al.
Published: (2023)
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters
by: Zhang, Haomin, et al.
Published: (2025)
by: Zhang, Haomin, et al.
Published: (2025)
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
by: Kim, Ji-Hoon, et al.
Published: (2025)
by: Kim, Ji-Hoon, et al.
Published: (2025)
From Vision to Sound: Advancing Audio Anomaly Detection with Vision-Based Algorithms
by: Barusco, Manuel, et al.
Published: (2025)
by: Barusco, Manuel, et al.
Published: (2025)
BGM2Pose: Active 3D Human Pose Estimation with Non-Stationary Sounds
by: Shibata, Yuto, et al.
Published: (2025)
by: Shibata, Yuto, et al.
Published: (2025)
Unified Cross-modal Translation of Score Images, Symbolic Music, and Performance Audio
by: Jung, Jongmin, et al.
Published: (2025)
by: Jung, Jongmin, et al.
Published: (2025)
See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement
by: Wang, Jinting, et al.
Published: (2025)
by: Wang, Jinting, et al.
Published: (2025)
Shushing! Let's Imagine an Authentic Speech from the Silent Video
by: Ye, Jiaxin, et al.
Published: (2025)
by: Ye, Jiaxin, et al.
Published: (2025)
Learning Separable Hidden Unit Contributions for Speaker-Adaptive Lip-Reading
by: Luo, Songtao, et al.
Published: (2023)
by: Luo, Songtao, et al.
Published: (2023)
Towards Unconstrained Audio Splicing Detection and Localization with Neural Networks
by: Moussa, Denise, et al.
Published: (2022)
by: Moussa, Denise, et al.
Published: (2022)
Enriching Multimodal Sentiment Analysis through Textual Emotional Descriptions of Visual-Audio Content
by: Wu, Sheng, et al.
Published: (2024)
by: Wu, Sheng, et al.
Published: (2024)
Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity
by: Pascual, Santiago, et al.
Published: (2024)
by: Pascual, Santiago, et al.
Published: (2024)
SAVE: Segment Audio-Visual Easy way using Segment Anything Model
by: Nguyen, Khanh-Binh, et al.
Published: (2024)
by: Nguyen, Khanh-Binh, et al.
Published: (2024)
Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-driven Diffusion
by: Ma, Jian, et al.
Published: (2024)
by: Ma, Jian, et al.
Published: (2024)
ICASSP 2024 Speech Signal Improvement Challenge
by: Ristea, Nicolae Catalin, et al.
Published: (2024)
by: Ristea, Nicolae Catalin, et al.
Published: (2024)
CleanUMamba: A Compact Mamba Network for Speech Denoising using Channel Pruning
by: Groot, Sjoerd, et al.
Published: (2024)
by: Groot, Sjoerd, et al.
Published: (2024)
GIRAFE: Glottal Imaging Dataset for Advanced Segmentation, Analysis, and Facilitative Playbacks Evaluation
by: Andrade-Miranda, G., et al.
Published: (2024)
by: Andrade-Miranda, G., et al.
Published: (2024)
Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual Segmentation
by: Yang, Qi, et al.
Published: (2023)
by: Yang, Qi, et al.
Published: (2023)
Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion
by: Rong, Yan, et al.
Published: (2024)
by: Rong, Yan, et al.
Published: (2024)
Improving Acoustic Scene Classification with City Features
by: Cai, Yiqiang, et al.
Published: (2025)
by: Cai, Yiqiang, et al.
Published: (2025)
Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities
by: Wang, Siyin, et al.
Published: (2024)
by: Wang, Siyin, et al.
Published: (2024)
The Effect of Perceptual Metrics on Music Representation Learning for Genre Classification
by: Namgyal, Tashi, et al.
Published: (2024)
by: Namgyal, Tashi, et al.
Published: (2024)
Spectral and Rhythm Features for Audio Classification with Deep Convolutional Neural Networks
by: Wolf-Monheim, Friedrich
Published: (2024)
by: Wolf-Monheim, Friedrich
Published: (2024)
Direct Speech-to-Speech Neural Machine Translation: A Survey
by: Gupta, Mahendra, et al.
Published: (2024)
by: Gupta, Mahendra, et al.
Published: (2024)
Spectral and Rhythm Feature Performance Evaluation for Category and Class Level Audio Classification with Deep Convolutional Neural Networks
by: Wolf-Monheim, Friedrich
Published: (2025)
by: Wolf-Monheim, Friedrich
Published: (2025)
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
ECHO: Environmental Sound Classification with Hierarchical Ontology-guided Semi-Supervised Learning
by: Gupta, Pranav, et al.
Published: (2024)
by: Gupta, Pranav, et al.
Published: (2024)
Automated Bioacoustic Monitoring for South African Bird Species on Unlabeled Data
by: Doell, Michael, et al.
Published: (2024)
by: Doell, Michael, et al.
Published: (2024)
NBM: an Open Dataset for the Acoustic Monitoring of Nocturnal Migratory Birds in Europe
by: Airale, Louis, et al.
Published: (2024)
by: Airale, Louis, et al.
Published: (2024)
Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization
by: Jia, Yanhao, et al.
Published: (2025)
by: Jia, Yanhao, et al.
Published: (2025)
VGGSounder: Audio-Visual Evaluations for Foundation Models
by: Zverev, Daniil, et al.
Published: (2025)
by: Zverev, Daniil, et al.
Published: (2025)
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
by: Choi, Hahyeon, et al.
Published: (2025)
by: Choi, Hahyeon, et al.
Published: (2025)
DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition
by: Jiang, Peiyuan, et al.
Published: (2025)
by: Jiang, Peiyuan, et al.
Published: (2025)
Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature Representation
by: Fan, Congyi, et al.
Published: (2025)
by: Fan, Congyi, et al.
Published: (2025)
Vision-to-Music Generation: A Survey
by: Wang, Zhaokai, et al.
Published: (2025)
by: Wang, Zhaokai, et al.
Published: (2025)
Latent Swap Joint Diffusion for 2D Long-Form Latent Generation
by: Dai, Yusheng, et al.
Published: (2025)
by: Dai, Yusheng, et al.
Published: (2025)
VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning
by: Hu, Qianyue, et al.
Published: (2025)
by: Hu, Qianyue, et al.
Published: (2025)
Similar Items
-
Automated Classification of Phonetic Segments in Child Speech Using Raw Ultrasound Imaging
by: Ani, Saja Al, et al.
Published: (2024) -
End-to-End Speech-to-Text Translation: A Survey
by: Sethiya, Nivedita, et al.
Published: (2023) -
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
by: Liu, Kai, et al.
Published: (2025) -
LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters
by: Zhang, Haomin, et al.
Published: (2025) -
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
by: Kim, Ji-Hoon, et al.
Published: (2025)