Decoding Linguistic Representations of Human Brain
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yu, Liu, Heyang, Wang, Yuhao, Xuan, Chuan, Hou, Yixuan, Feng, Sheng, Liu, Hongcheng, Liao, Yusheng, Wang, Yanfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards an End-to-End Framework for Invasive Brain Signal Decoding with Large Language Models
von: Feng, Sheng, et al.
Veröffentlicht: (2024)
von: Feng, Sheng, et al.
Veröffentlicht: (2024)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
von: Wei, Sheng-Lun, et al.
Veröffentlicht: (2026)
von: Wei, Sheng-Lun, et al.
Veröffentlicht: (2026)
Guided by the Plan: Enhancing Faithful Autoregressive Text-to-Audio Generation with Guided Decoding
von: Wang, Juncheng, et al.
Veröffentlicht: (2026)
von: Wang, Juncheng, et al.
Veröffentlicht: (2026)
Investigating Causal Cues: Strengthening Spoofed Audio Detection with Human-Discernible Linguistic Features
von: Khanjani, Zahra, et al.
Veröffentlicht: (2024)
von: Khanjani, Zahra, et al.
Veröffentlicht: (2024)
Personalized Voice Synthesis through Human-in-the-Loop Coordinate Descent
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
von: Tian, Yusheng, et al.
Veröffentlicht: (2024)
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
von: Zhao, Jiahui, et al.
Veröffentlicht: (2024)
von: Zhao, Jiahui, et al.
Veröffentlicht: (2024)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
von: Chen, Qian, et al.
Veröffentlicht: (2023)
von: Chen, Qian, et al.
Veröffentlicht: (2023)
Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Ruoyu, et al.
Veröffentlicht: (2025)
Mobile Recording Device Recognition Based Cross-Scale and Multi-Level Representation Learning
von: Zeng, Chunyan, et al.
Veröffentlicht: (2024)
von: Zeng, Chunyan, et al.
Veröffentlicht: (2024)
BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation
von: Li, Jilong, et al.
Veröffentlicht: (2024)
von: Li, Jilong, et al.
Veröffentlicht: (2024)
Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
LUCY: Linguistic Understanding and Control Yielding Early Stage of Her
von: Gao, Heting, et al.
Veröffentlicht: (2025)
von: Gao, Heting, et al.
Veröffentlicht: (2025)
Task-Agnostic Structured Pruning of Speech Representation Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2023)
von: Wang, Haoyu, et al.
Veröffentlicht: (2023)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
The Third VoicePrivacy Challenge: Preserving Emotional Expressiveness and Linguistic Content in Voice Anonymization
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2026)
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2026)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment
von: Wang, Ke, et al.
Veröffentlicht: (2025)
von: Wang, Ke, et al.
Veröffentlicht: (2025)
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
von: Zhao, Wenbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wenbo, et al.
Veröffentlicht: (2024)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
Rethinking Discrete Speech Representation Tokens for Accent Generation
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2026)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2026)
BATON: Aligning Text-to-Audio Model with Human Preference Feedback
von: Liao, Huan, et al.
Veröffentlicht: (2024)
von: Liao, Huan, et al.
Veröffentlicht: (2024)
Next Tokens Denoising for Speech Synthesis
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
The FruitShell French synthesis system at the Blizzard 2023 Challenge
von: Qi, Xin, et al.
Veröffentlicht: (2023)
von: Qi, Xin, et al.
Veröffentlicht: (2023)
SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
von: Poli, Maxime, et al.
Veröffentlicht: (2025)
von: Poli, Maxime, et al.
Veröffentlicht: (2025)
Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
von: He, Haorui, et al.
Veröffentlicht: (2024)
von: He, Haorui, et al.
Veröffentlicht: (2024)
BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing
von: Wang, Chen, et al.
Veröffentlicht: (2023)
von: Wang, Chen, et al.
Veröffentlicht: (2023)
U-GIFT: Uncertainty-Guided Firewall for Toxic Speech in Few-Shot Scenario
von: Song, Jiaxin, et al.
Veröffentlicht: (2025)
von: Song, Jiaxin, et al.
Veröffentlicht: (2025)
Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards an End-to-End Framework for Invasive Brain Signal Decoding with Large Language Models
von: Feng, Sheng, et al.
Veröffentlicht: (2024) -
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024) -
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
von: Hou, Yixuan, et al.
Veröffentlicht: (2025) -
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025) -
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
von: Wei, Sheng-Lun, et al.
Veröffentlicht: (2026)