Semantic-Aware Confidence Calibration for Automated Audio Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dunker, Lucas, Menta, Sai Akshay, Addepalli, Snigdha Mohana, Garapati, Venkata Krishna Rayalu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
von: Takeuchi, Daiki, et al.
Veröffentlicht: (2025)
von: Takeuchi, Daiki, et al.
Veröffentlicht: (2025)
Enhancing Knee Osteoarthritis severity level classification using diffusion augmented images
von: Chowdary, Paleti Nikhil, et al.
Veröffentlicht: (2023)
von: Chowdary, Paleti Nikhil, et al.
Veröffentlicht: (2023)
The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
Aligning Audio Captions with Human Preferences
von: Hegde, Kartik, et al.
Veröffentlicht: (2025)
von: Hegde, Kartik, et al.
Veröffentlicht: (2025)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering
von: Tran, Dinh Phu, et al.
Veröffentlicht: (2026)
von: Tran, Dinh Phu, et al.
Veröffentlicht: (2026)
Audio-Visual Continual Test-Time Adaptation without Forgetting
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2026)
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2026)
Audio Question Answering with GRPO-Based Fine-Tuning and Calibrated Segment-Level Predictions
von: Gibier, Marcel, et al.
Veröffentlicht: (2025)
von: Gibier, Marcel, et al.
Veröffentlicht: (2025)
Improving Text-To-Audio Models with Synthetic Captions
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training
von: Wu, Yanru, et al.
Veröffentlicht: (2026)
von: Wu, Yanru, et al.
Veröffentlicht: (2026)
Rethinking Music Captioning with Music Metadata LLMs
von: Bukey, Irmak, et al.
Veröffentlicht: (2026)
von: Bukey, Irmak, et al.
Veröffentlicht: (2026)
AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait Inference
von: Shinoda, Risa, et al.
Veröffentlicht: (2026)
von: Shinoda, Risa, et al.
Veröffentlicht: (2026)
PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
EmoHRNet: High-Resolution Neural Network Based Speech Emotion Recognition
von: Muppidi, Akshay, et al.
Veröffentlicht: (2025)
von: Muppidi, Akshay, et al.
Veröffentlicht: (2025)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
How to Label Resynthesized Audio: The Dual Role of Neural Audio Codecs in Audio Deepfake Detection
von: Xiao, Yixuan, et al.
Veröffentlicht: (2026)
von: Xiao, Yixuan, et al.
Veröffentlicht: (2026)
ADNAC: Audio Denoiser using Neural Audio Codec
von: Jimon, Daniel, et al.
Veröffentlicht: (2025)
von: Jimon, Daniel, et al.
Veröffentlicht: (2025)
Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning
von: Wu, Daiqing, et al.
Veröffentlicht: (2026)
von: Wu, Daiqing, et al.
Veröffentlicht: (2026)
Cross-Attention with Confidence Weighting for Multi-Channel Audio Alignment
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2025)
EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation
von: Manivannan, Mithun, et al.
Veröffentlicht: (2024)
von: Manivannan, Mithun, et al.
Veröffentlicht: (2024)
ACES: Evaluating Automated Audio Captioning Models on the Semantics of Sounds
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2024)
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2024)
Virtual Consistency for Audio Editing
von: Cervera, Matthieu, et al.
Veröffentlicht: (2025)
von: Cervera, Matthieu, et al.
Veröffentlicht: (2025)
Learning Spatially-Aware Language and Audio Embeddings
von: Devnani, Bhavika, et al.
Veröffentlicht: (2024)
von: Devnani, Bhavika, et al.
Veröffentlicht: (2024)
A Semi-Supervised Framework for Speech Confidence Detection using Whisper
von: Wynn, Adam, et al.
Veröffentlicht: (2026)
von: Wynn, Adam, et al.
Veröffentlicht: (2026)
Segmentwise Pruning in Audio-Language Models
von: Gibier, Marcel, et al.
Veröffentlicht: (2025)
von: Gibier, Marcel, et al.
Veröffentlicht: (2025)
Adapting Neural Audio Codecs to EEG
von: Kastrati, Ard, et al.
Veröffentlicht: (2025)
von: Kastrati, Ard, et al.
Veröffentlicht: (2025)
PACE: Pretrained Audio Continual Learning
von: Li, Chang, et al.
Veröffentlicht: (2026)
von: Li, Chang, et al.
Veröffentlicht: (2026)
Phase-Aware Deep Learning with Complex-Valued CNNs for Audio Signal Applications
von: Agrawal, Naman
Veröffentlicht: (2025)
von: Agrawal, Naman
Veröffentlicht: (2025)
$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
Investigating Modality Contribution in Audio LLMs for Music
von: Morais, Giovana, et al.
Veröffentlicht: (2025)
von: Morais, Giovana, et al.
Veröffentlicht: (2025)
Audio Super-Resolution with Latent Bridge Models
von: Li, Chang, et al.
Veröffentlicht: (2025)
von: Li, Chang, et al.
Veröffentlicht: (2025)
LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
APEX: Audio Prototype EXplanations for Classification Tasks
von: Kawa, Piotr, et al.
Veröffentlicht: (2026)
von: Kawa, Piotr, et al.
Veröffentlicht: (2026)
Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards
von: Fang, Linghan, et al.
Veröffentlicht: (2026)
von: Fang, Linghan, et al.
Veröffentlicht: (2026)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
Training-Free Multimodal Guidance for Video to Audio Generation
von: Grassucci, Eleonora, et al.
Veröffentlicht: (2025)
von: Grassucci, Eleonora, et al.
Veröffentlicht: (2025)
Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
Discrete Audio Representations for Automated Audio Captioning
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
Uncertainty Calibration of Multi-Label Bird Sound Classifiers
von: Schwinger, Raphael, et al.
Veröffentlicht: (2025)
von: Schwinger, Raphael, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
von: Takeuchi, Daiki, et al.
Veröffentlicht: (2025) -
Enhancing Knee Osteoarthritis severity level classification using diffusion augmented images
von: Chowdary, Paleti Nikhil, et al.
Veröffentlicht: (2023) -
The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026) -
Aligning Audio Captions with Human Preferences
von: Hegde, Kartik, et al.
Veröffentlicht: (2025) -
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)