Classifier-Guided Captioning Across Modalities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shaulov, Ariel, Shaharabany, Tal, Shaar, Eitan, Chechik, Gal, Wolf, Lior |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RECAP: Retrieval-Augmented Audio Captioning
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
Text2midi: Generating Symbolic Music from Captions
von: Bhandari, Keshav, et al.
Veröffentlicht: (2024)
von: Bhandari, Keshav, et al.
Veröffentlicht: (2024)
Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
von: Manco, Ilaria, et al.
Veröffentlicht: (2024)
von: Manco, Ilaria, et al.
Veröffentlicht: (2024)
Evaluating the Usefulness of Non-Diagnostic Speech Data for Developing Parkinson's Disease Classifiers
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2025)
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2025)
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2023)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2023)
Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2026)
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2026)
SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
Leveraging Audio and Text Modalities in Mental Health: A Study of LLMs Performance
von: Ali, Abdelrahman A., et al.
Veröffentlicht: (2024)
von: Ali, Abdelrahman A., et al.
Veröffentlicht: (2024)
VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
von: Qiang, Chunyu, et al.
Veröffentlicht: (2024)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2024)
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
von: Wang, Qiaolin, et al.
Veröffentlicht: (2025)
von: Wang, Qiaolin, et al.
Veröffentlicht: (2025)
Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder
von: Dai, Yusheng, et al.
Veröffentlicht: (2023)
von: Dai, Yusheng, et al.
Veröffentlicht: (2023)
MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response
von: Deng, Zihao, et al.
Veröffentlicht: (2023)
von: Deng, Zihao, et al.
Veröffentlicht: (2023)
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
von: Shin, Seungyoun, et al.
Veröffentlicht: (2025)
von: Shin, Seungyoun, et al.
Veröffentlicht: (2025)
Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
von: Yoon, Jinsung, et al.
Veröffentlicht: (2025)
von: Yoon, Jinsung, et al.
Veröffentlicht: (2025)
ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
von: Yang, Cheng-Yeh, et al.
Veröffentlicht: (2026)
Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2026)
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2026)
Cross-Modal Learning for Music-to-Music-Video Description Generation
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
S2Cap: A Benchmark and a Baseline for Singing Style Captioning
von: Ok, Hyunjong, et al.
Veröffentlicht: (2024)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2024)
DiffMoog: a Differentiable Modular Synthesizer for Sound Matching
von: Uzrad, Noy, et al.
Veröffentlicht: (2024)
von: Uzrad, Noy, et al.
Veröffentlicht: (2024)
Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition
von: Radhakrishnan, Srijith, et al.
Veröffentlicht: (2023)
von: Radhakrishnan, Srijith, et al.
Veröffentlicht: (2023)
NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention
von: De Silva, Dashanka, et al.
Veröffentlicht: (2024)
von: De Silva, Dashanka, et al.
Veröffentlicht: (2024)
Factor-Conditioned Speaking-Style Captioning
von: Ando, Atsushi, et al.
Veröffentlicht: (2024)
von: Ando, Atsushi, et al.
Veröffentlicht: (2024)
SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models
von: Wan, Zhen, et al.
Veröffentlicht: (2025)
von: Wan, Zhen, et al.
Veröffentlicht: (2025)
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2025)
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
von: Chou, Cheng-Kang, et al.
Veröffentlicht: (2025)
von: Chou, Cheng-Kang, et al.
Veröffentlicht: (2025)
GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness
von: Chen, Hongjie, et al.
Veröffentlicht: (2025)
von: Chen, Hongjie, et al.
Veröffentlicht: (2025)
Data Augmentation for Spoken Grammatical Error Correction
von: Karanasou, Penny, et al.
Veröffentlicht: (2025)
von: Karanasou, Penny, et al.
Veröffentlicht: (2025)
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
von: Frohmann, Markus, et al.
Veröffentlicht: (2025)
von: Frohmann, Markus, et al.
Veröffentlicht: (2025)
A Chinese Heart Failure Status Speech Database with Universal and Personalised Classification
von: Pan, Yue, et al.
Veröffentlicht: (2025)
von: Pan, Yue, et al.
Veröffentlicht: (2025)
ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
TurboBias: Universal ASR Context-Biasing powered by GPU-accelerated Phrase-Boosting Tree
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2025)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2025)
Fun-Audio-Chat Technical Report
von: Tongyi Fun Team, et al.
Veröffentlicht: (2025)
von: Tongyi Fun Team, et al.
Veröffentlicht: (2025)
Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
von: Srivastav, Vaibhav, et al.
Veröffentlicht: (2025)
von: Srivastav, Vaibhav, et al.
Veröffentlicht: (2025)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RECAP: Retrieval-Augmented Audio Captioning
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023) -
Text2midi: Generating Symbolic Music from Captions
von: Bhandari, Keshav, et al.
Veröffentlicht: (2024) -
Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
von: Manco, Ilaria, et al.
Veröffentlicht: (2024) -
Evaluating the Usefulness of Non-Diagnostic Speech Data for Developing Parkinson's Disease Classifiers
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2025) -
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2023)