AI Safety in Practice: Enhancing Adversarial Robustness in Multimodal Image Captioning
Fuente:
arXiv
Guardado en:
| Autores principales: | Rashid, Maisha Binte, Rivas, Pablo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MGSC: A Multi-granularity Consistency Framework for Robust End-to-end Asr
por: Yang, Xuwen
Publicado: (2025)
por: Yang, Xuwen
Publicado: (2025)
Leveraging OpenFlamingo for Multimodal Embedding Analysis of C2C Car Parts Data
por: Rashid, Maisha Binte, et al.
Publicado: (2025)
por: Rashid, Maisha Binte, et al.
Publicado: (2025)
Dynamic Behaviour of Connectionist Speech Recognition with Strong Latency Constraints
por: Salvi, Giampiero
Publicado: (2024)
por: Salvi, Giampiero
Publicado: (2024)
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
por: Lee, Yoonhyung, et al.
Publicado: (2026)
por: Lee, Yoonhyung, et al.
Publicado: (2026)
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
por: Dai, Ziqi, et al.
Publicado: (2025)
por: Dai, Ziqi, et al.
Publicado: (2025)
An open-source voice type classifier for child-centered daylong recordings
por: Lavechin, Marvin, et al.
Publicado: (2020)
por: Lavechin, Marvin, et al.
Publicado: (2020)
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
por: Fang, Qingkai, et al.
Publicado: (2024)
por: Fang, Qingkai, et al.
Publicado: (2024)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
por: Fang, Qingkai, et al.
Publicado: (2024)
por: Fang, Qingkai, et al.
Publicado: (2024)
MIRFLEX: Music Information Retrieval Feature Library for Extraction
por: Chopra, Anuradha, et al.
Publicado: (2024)
por: Chopra, Anuradha, et al.
Publicado: (2024)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
por: Fang, Qingkai, et al.
Publicado: (2024)
por: Fang, Qingkai, et al.
Publicado: (2024)
Effectiveness of Text, Acoustic, and Lattice-based representations in Spoken Language Understanding tasks
por: Villatoro-Tello, Esaú, et al.
Publicado: (2022)
por: Villatoro-Tello, Esaú, et al.
Publicado: (2022)
A Baseline Multimodal Approach to Emotion Recognition in Conversations
por: Yeste, Víctor, et al.
Publicado: (2026)
por: Yeste, Víctor, et al.
Publicado: (2026)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
por: Cheng, Zhuangfei, et al.
Publicado: (2025)
por: Cheng, Zhuangfei, et al.
Publicado: (2025)
Quantifying the effect of speech pathology on automatic and human speaker verification
por: Halpern, Bence Mark, et al.
Publicado: (2024)
por: Halpern, Bence Mark, et al.
Publicado: (2024)
Community-Informed AI Models for Police Accountability
por: Graham, Benjamin A. T., et al.
Publicado: (2024)
por: Graham, Benjamin A. T., et al.
Publicado: (2024)
Exploring rhythm formant analysis for Indic language classification
por: Gogoi, Parismita, et al.
Publicado: (2024)
por: Gogoi, Parismita, et al.
Publicado: (2024)
A Robust Classification Method using Hybrid Word Embedding for Early Diagnosis of Alzheimer's Disease
por: Li, Yangyang
Publicado: (2025)
por: Li, Yangyang
Publicado: (2025)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
por: Donepudi, Dharma Teja
Publicado: (2025)
por: Donepudi, Dharma Teja
Publicado: (2025)
An End-to-End Approach for Korean Wakeword Systems with Speaker Authentication
por: Seo, Geonwoo
Publicado: (2025)
por: Seo, Geonwoo
Publicado: (2025)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
por: Sharma, Manali, et al.
Publicado: (2026)
por: Sharma, Manali, et al.
Publicado: (2026)
Beyond Levenshtein: Leveraging Multiple Algorithms for Robust Word Error Rate Computations And Granular Error Classifications
por: Kuhn, Korbinian, et al.
Publicado: (2024)
por: Kuhn, Korbinian, et al.
Publicado: (2024)
SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding
por: Bai, Bingsong, et al.
Publicado: (2025)
por: Bai, Bingsong, et al.
Publicado: (2025)
Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation
por: Freisinger, Steffen, et al.
Publicado: (2026)
por: Freisinger, Steffen, et al.
Publicado: (2026)
U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF
por: Song, Xingchen, et al.
Publicado: (2024)
por: Song, Xingchen, et al.
Publicado: (2024)
PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation
por: Yi, Yungang, et al.
Publicado: (2024)
por: Yi, Yungang, et al.
Publicado: (2024)
Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English
por: Zhang, Haoyang, et al.
Publicado: (2025)
por: Zhang, Haoyang, et al.
Publicado: (2025)
Framework for Curating Speech Datasets and Evaluating ASR Systems: A Case Study for Polish
por: Junczyk, Michał
Publicado: (2024)
por: Junczyk, Michał
Publicado: (2024)
Computational modeling of early language learning from acoustic speech and audiovisual input without linguistic priors
por: Räsänen, Okko
Publicado: (2026)
por: Räsänen, Okko
Publicado: (2026)
Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
por: Khorrami, Khazar, et al.
Publicado: (2021)
por: Khorrami, Khazar, et al.
Publicado: (2021)
An accurate and revised version of optical character recognition-based speech synthesis using LabVIEW
por: Mehta, Prateek, et al.
Publicado: (2025)
por: Mehta, Prateek, et al.
Publicado: (2025)
Multimodal Emotion Recognition using Audio-Video Transformer Fusion with Cross Attention
por: R, Joe Dhanith P, et al.
Publicado: (2024)
por: R, Joe Dhanith P, et al.
Publicado: (2024)
SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models
por: Dua, Karan, et al.
Publicado: (2025)
por: Dua, Karan, et al.
Publicado: (2025)
Advancing Voice Cloning for Nepali: Leveraging Transfer Learning in a Low-Resource Language
por: Karki, Manjil, et al.
Publicado: (2024)
por: Karki, Manjil, et al.
Publicado: (2024)
AQUALLM: Audio Question Answering Data Generation Using Large Language Models
por: Behera, Swarup Ranjan, et al.
Publicado: (2023)
por: Behera, Swarup Ranjan, et al.
Publicado: (2023)
Beyond Speech and More: Investigating the Emergent Ability of Speech Foundation Models for Classifying Physiological Time-Series Signals
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
por: Ahn, Taekyung, et al.
Publicado: (2024)
por: Ahn, Taekyung, et al.
Publicado: (2024)
Ejemplares similares
-
MGSC: A Multi-granularity Consistency Framework for Robust End-to-end Asr
por: Yang, Xuwen
Publicado: (2025) -
Leveraging OpenFlamingo for Multimodal Embedding Analysis of C2C Car Parts Data
por: Rashid, Maisha Binte, et al.
Publicado: (2025) -
Dynamic Behaviour of Connectionist Speech Recognition with Strong Latency Constraints
por: Salvi, Giampiero
Publicado: (2024) -
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
por: Lee, Yoonhyung, et al.
Publicado: (2026) -
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
por: Dai, Ziqi, et al.
Publicado: (2025)