Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and Localization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Junyan, Lu, Wei, Luo, Xiangyang, Yang, Rui, Wang, Qian, Cao, Xiaochun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment
von: Roy, Abhinaba, et al.
Veröffentlicht: (2025)
von: Roy, Abhinaba, et al.
Veröffentlicht: (2025)
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
Faked Speech Detection with Zero Prior Knowledge
von: Ajmi, Sahar Al, et al.
Veröffentlicht: (2022)
von: Ajmi, Sahar Al, et al.
Veröffentlicht: (2022)
Can Sound Replace Vision in LLaVA With Token Substitution?
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
ChordSync: Conformer-Based Alignment of Chord Annotations to Music Audio
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2024)
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2024)
FakeSound: Deepfake General Audio Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction
von: Cheripally, Sowmya
Veröffentlicht: (2024)
von: Cheripally, Sowmya
Veröffentlicht: (2024)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
STAR: Speech-to-Audio Generation via Representation Learning
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
M2D-CLAP: Exploring General-purpose Audio-Language Representations Beyond CLAP
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2025)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2025)
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond
von: Richter-Powell, Jessie, et al.
Veröffentlicht: (2025)
von: Richter-Powell, Jessie, et al.
Veröffentlicht: (2025)
Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning Network
von: Wu, Junyan, et al.
Veröffentlicht: (2025)
von: Wu, Junyan, et al.
Veröffentlicht: (2025)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
SARA: Stress Test Reasoning in Audio Deepfake Detection
von: Nguyen, Binh, et al.
Veröffentlicht: (2026)
von: Nguyen, Binh, et al.
Veröffentlicht: (2026)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Prevailing Research Areas for Music AI in the Era of Foundation Models
von: Wei, Megan, et al.
Veröffentlicht: (2024)
von: Wei, Megan, et al.
Veröffentlicht: (2024)
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
von: Melechovsky, Jan, et al.
Veröffentlicht: (2025)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2025)
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
von: Han, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Han, Zhiyuan, et al.
Veröffentlicht: (2025)
Automatic Album Sequencing
von: Herrmann, Vincent, et al.
Veröffentlicht: (2024)
von: Herrmann, Vincent, et al.
Veröffentlicht: (2024)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
Audio-based Kinship Verification Using Age Domain Conversion
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025)
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025)
Towards Pre-training an Effective Respiratory Audio Foundation Model
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2025)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2025)
Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection
von: Cao, Xinwei, et al.
Veröffentlicht: (2026)
von: Cao, Xinwei, et al.
Veröffentlicht: (2026)
Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2025)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2025)
Graph Connectionist Temporal Classification for Phoneme Recognition
von: Grafé, Henry, et al.
Veröffentlicht: (2025)
von: Grafé, Henry, et al.
Veröffentlicht: (2025)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
von: Khushiyant, et al.
Veröffentlicht: (2026)
von: Khushiyant, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026) -
Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment
von: Roy, Abhinaba, et al.
Veröffentlicht: (2025) -
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024) -
Faked Speech Detection with Zero Prior Knowledge
von: Ajmi, Sahar Al, et al.
Veröffentlicht: (2022) -
Can Sound Replace Vision in LLaVA With Token Substitution?
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)