SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Peidong, Ma, Zhiming, Dai, Xin, Liu, Yongkang, Feng, Shi, Yang, Xiaocui, Hu, Wenxing, Wang, Zhihao, Pan, Mingjun, Yuan, Li, Wang, Daling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
von: Chen, Yujie, et al.
Veröffentlicht: (2024)
von: Chen, Yujie, et al.
Veröffentlicht: (2024)
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2024)
von: Wang, Peidong, et al.
Veröffentlicht: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
Synthetic Audio Forensics Evaluation (SAFE) Challenge
von: Trapeznikov, Kirill, et al.
Veröffentlicht: (2025)
von: Trapeznikov, Kirill, et al.
Veröffentlicht: (2025)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
audio2chart: End to End Audio Transcription into playable Guitar Hero charts
von: Tripodi, Riccardo
Veröffentlicht: (2025)
von: Tripodi, Riccardo
Veröffentlicht: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
End-to-End Real-World Polyphonic Piano Audio-to-Score Transcription with Hierarchical Decoding
von: Zeng, Wei, et al.
Veröffentlicht: (2024)
von: Zeng, Wei, et al.
Veröffentlicht: (2024)
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation
von: Gao, Zhuoyue, et al.
Veröffentlicht: (2026)
von: Gao, Zhuoyue, et al.
Veröffentlicht: (2026)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
von: Lin, Wan, et al.
Veröffentlicht: (2024)
von: Lin, Wan, et al.
Veröffentlicht: (2024)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
von: Shi, Shuchen, et al.
Veröffentlicht: (2024)
von: Shi, Shuchen, et al.
Veröffentlicht: (2024)
Meta-Learning in Audio and Speech Processing: An End to End Comprehensive Review
von: Raimon, Athul, et al.
Veröffentlicht: (2024)
von: Raimon, Athul, et al.
Veröffentlicht: (2024)
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
von: Yu, Yifeng, et al.
Veröffentlicht: (2024)
von: Yu, Yifeng, et al.
Veröffentlicht: (2024)
End-to-End Speech-to-Text Translation: A Survey
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
End-to-End Diarization utilizing Attractor Deep Clustering
von: Palzer, David, et al.
Veröffentlicht: (2025)
von: Palzer, David, et al.
Veröffentlicht: (2025)
Speaker Adaptation for Quantised End-to-End ASR Models
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
von: Moslem, Yasmin
Veröffentlicht: (2024)
von: Moslem, Yasmin
Veröffentlicht: (2024)
SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
von: Li, Sirui, et al.
Veröffentlicht: (2025)
von: Li, Sirui, et al.
Veröffentlicht: (2025)
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
von: Plaquet, Alexis, et al.
Veröffentlicht: (2025)
von: Plaquet, Alexis, et al.
Veröffentlicht: (2025)
IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
von: Zhuang, Zhuoran, et al.
Veröffentlicht: (2026)
von: Zhuang, Zhuoran, et al.
Veröffentlicht: (2026)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
von: Huang, Wen, et al.
Veröffentlicht: (2025)
von: Huang, Wen, et al.
Veröffentlicht: (2025)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors
von: Landini, Federico, et al.
Veröffentlicht: (2023)
von: Landini, Federico, et al.
Veröffentlicht: (2023)
End-to-End Amp Modeling: From Data to Controllable Guitar Amplifier Models
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2024)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2024)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Representation Purification for End-to-End Speech Translation
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
von: Chen, Yujie, et al.
Veröffentlicht: (2024) -
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2024) -
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024) -
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024) -
Synthetic Audio Forensics Evaluation (SAFE) Challenge
von: Trapeznikov, Kirill, et al.
Veröffentlicht: (2025)