Gespeichert in:
| Hauptverfasser: | Xu, Xiran, Yan, Yujie, Wu, Xihong, Chen, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.23960 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
von: Zhu, Haolin, et al.
Veröffentlicht: (2024)
von: Zhu, Haolin, et al.
Veröffentlicht: (2024)
The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models
von: You, Yuhuan, et al.
Veröffentlicht: (2026)
von: You, Yuhuan, et al.
Veröffentlicht: (2026)
Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording
von: Wang, Bo, et al.
Veröffentlicht: (2024)
von: Wang, Bo, et al.
Veröffentlicht: (2024)
MEBM-Speech: Multi-scale Enhanced BrainMagic for Robust MEG Speech Detection
von: Songyi, Li, et al.
Veröffentlicht: (2026)
von: Songyi, Li, et al.
Veröffentlicht: (2026)
MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification
von: Jinghua, Liang, et al.
Veröffentlicht: (2026)
von: Jinghua, Liang, et al.
Veröffentlicht: (2026)
Unifying EEG and Speech for Emotion Recognition: A Two-Step Joint Learning Framework for Handling Missing EEG Data During Inference
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
MindMelody: A Closed-Loop EEG-Driven System for Personalized Music Intervention
von: Zhang, Yimeng, et al.
Veröffentlicht: (2026)
von: Zhang, Yimeng, et al.
Veröffentlicht: (2026)
Hyperbolic Additive Margin Softmax with Hierarchical Information for Speaker Verification
von: Fang, Zhihua, et al.
Veröffentlicht: (2026)
von: Fang, Zhihua, et al.
Veröffentlicht: (2026)
Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs
von: Xue, Jun, et al.
Veröffentlicht: (2026)
von: Xue, Jun, et al.
Veröffentlicht: (2026)
Switchable deep beamformer for high-quality and real-time passive acoustic mapping
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
A DenseNet-based method for decoding auditory spatial attention with EEG
von: Xu, Xiran, et al.
Veröffentlicht: (2023)
von: Xu, Xiran, et al.
Veröffentlicht: (2023)
Reject Threshold Adaptation for Open-Set Model Attribution of Deepfake Audio
von: Yan, Xinrui, et al.
Veröffentlicht: (2024)
von: Yan, Xinrui, et al.
Veröffentlicht: (2024)
CIPHER: Conformer-based Inference of Phonemes from High-density EEG
von: Madishetty, Varshith
Veröffentlicht: (2026)
von: Madishetty, Varshith
Veröffentlicht: (2026)
SymphonyGen: 3D Hierarchical Orchestral Generation with Controllable Harmony Skeleton
von: He, Xuzheng, et al.
Veröffentlicht: (2026)
von: He, Xuzheng, et al.
Veröffentlicht: (2026)
RAMoEA-QA: Hierarchical Specialization for Robust Respiratory Audio Question Answering
von: Bertolino, Gaia A., et al.
Veröffentlicht: (2026)
von: Bertolino, Gaia A., et al.
Veröffentlicht: (2026)
Smark: A Watermark for Text-to-Speech Diffusion Models via Discrete Wavelet Transform
von: Zhang, Yichuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yichuan, et al.
Veröffentlicht: (2025)
Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG
von: Yang, Haoyun, et al.
Veröffentlicht: (2026)
von: Yang, Haoyun, et al.
Veröffentlicht: (2026)
SWIM: Short-Window CNN Integrated with Mamba for EEG-Based Auditory Spatial Attention Decoding
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
Dynamic Fusion Multimodal Network for SpeechWellness Detection
von: Sun, Wenqiang, et al.
Veröffentlicht: (2025)
von: Sun, Wenqiang, et al.
Veröffentlicht: (2025)
Efficient Long-Sequence Diffusion Modeling for Symbolic Music Generation
von: Xu, Jinhan, et al.
Veröffentlicht: (2026)
von: Xu, Jinhan, et al.
Veröffentlicht: (2026)
EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction
von: Jing, Chong, et al.
Veröffentlicht: (2026)
von: Jing, Chong, et al.
Veröffentlicht: (2026)
CompLex: Music Theory Lexicon Constructed by Autonomous Agents for Automatic Music Generation
von: Hu, Zhejing, et al.
Veröffentlicht: (2025)
von: Hu, Zhejing, et al.
Veröffentlicht: (2025)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
GeHirNet: A Gender-Aware Hierarchical Model for Voice Pathology Classification
von: Wu, Fan, et al.
Veröffentlicht: (2025)
von: Wu, Fan, et al.
Veröffentlicht: (2025)
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
von: Zuo, Heda, et al.
Veröffentlicht: (2025)
von: Zuo, Heda, et al.
Veröffentlicht: (2025)
RAS: a Reliability Oriented Metric for Automatic Speech Recognition
von: Huang, Wenbin, et al.
Veröffentlicht: (2026)
von: Huang, Wenbin, et al.
Veröffentlicht: (2026)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
Towards Unified Neural Decoding of Perceived, Spoken and Imagined Speech from EEG Signals
von: Lee, Jung-Sun, et al.
Veröffentlicht: (2024)
von: Lee, Jung-Sun, et al.
Veröffentlicht: (2024)
Hierarchical Graph Neural Network for Compressed Speech Steganalysis
von: Hemis, Mustapha, et al.
Veröffentlicht: (2025)
von: Hemis, Mustapha, et al.
Veröffentlicht: (2025)
Hear: Hierarchically Enhanced Aesthetic Representations For Multidimensional Music Evaluation
von: Liu, Shuyang, et al.
Veröffentlicht: (2025)
von: Liu, Shuyang, et al.
Veröffentlicht: (2025)
Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
von: Chen, Wei, et al.
Veröffentlicht: (2025)
von: Chen, Wei, et al.
Veröffentlicht: (2025)
NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control
von: Wen, Yufan, et al.
Veröffentlicht: (2026)
von: Wen, Yufan, et al.
Veröffentlicht: (2026)
MuseCPBench: an Empirical Study of Music Editing Methods through Music Context Preservation
von: Vishe, Yash, et al.
Veröffentlicht: (2025)
von: Vishe, Yash, et al.
Veröffentlicht: (2025)
UniSRCodec: Unified and Low-Bitrate Single Codebook Codec with Sub-Band Reconstruction
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
von: Luo, Dan, et al.
Veröffentlicht: (2025)
von: Luo, Dan, et al.
Veröffentlicht: (2025)
Toward Complex-Valued Neural Networks for Waveform Generation
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2026)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2026)
Evaluating Neural Networks Architectures for Spring Reverb Modelling
von: Papaleo, Francesco, et al.
Veröffentlicht: (2024)
von: Papaleo, Francesco, et al.
Veröffentlicht: (2024)
Evaluating Semantic Fragility in Text-to-Audio Generation Systems Under Controlled Prompt Perturbations
von: Wu, Jiahui
Veröffentlicht: (2026)
von: Wu, Jiahui
Veröffentlicht: (2026)
NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations
von: Liao, Huan, et al.
Veröffentlicht: (2025)
von: Liao, Huan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
von: Zhu, Haolin, et al.
Veröffentlicht: (2024) -
The World is Not Mono: Enabling Spatial Understanding in Large Audio-Language Models
von: You, Yuhuan, et al.
Veröffentlicht: (2026) -
Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording
von: Wang, Bo, et al.
Veröffentlicht: (2024) -
MEBM-Speech: Multi-scale Enhanced BrainMagic for Robust MEG Speech Detection
von: Songyi, Li, et al.
Veröffentlicht: (2026) -
MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification
von: Jinghua, Liang, et al.
Veröffentlicht: (2026)