MSTS: A Multimodal Safety Test Suite for Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Röttger, Paul, Attanasio, Giuseppe, Friedrich, Felix, Goldzycher, Janis, Parrish, Alicia, Bhardwaj, Rishabh, Di Bonaventura, Chiara, Eng, Roman, Geagea, Gaia El Khoury, Goswami, Sujata, Han, Jieun, Hovy, Dirk, Jeong, Seogyeong, Jeretič, Paloma, Plaza-del-Arco, Flor Miriam, Rooein, Donya, Schramowski, Patrick, Shaitarova, Anastassia, Shen, Xudong, Willats, Richard, Zugarini, Andrea, Vidgen, Bertie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
di: Rooein, Donya, et al.
Pubblicazione: (2024)
di: Rooein, Donya, et al.
Pubblicazione: (2024)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
di: Röttger, Paul, et al.
Pubblicazione: (2024)
di: Röttger, Paul, et al.
Pubblicazione: (2024)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
di: Röttger, Paul, et al.
Pubblicazione: (2023)
di: Röttger, Paul, et al.
Pubblicazione: (2023)
Conversations as a Source for Teaching Scientific Concepts at Different Education Levels
di: Rooein, Donya, et al.
Pubblicazione: (2024)
di: Rooein, Donya, et al.
Pubblicazione: (2024)
Classification is a RAG problem: A case study on hate speech detection
di: Willats, Richard, et al.
Pubblicazione: (2025)
di: Willats, Richard, et al.
Pubblicazione: (2025)
Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models
di: Rooein, Donya, et al.
Pubblicazione: (2025)
di: Rooein, Donya, et al.
Pubblicazione: (2025)
Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset
di: Goldzycher, Janis, et al.
Pubblicazione: (2024)
di: Goldzycher, Janis, et al.
Pubblicazione: (2024)
Biased Tales: Cultural and Topic Bias in Generating Children's Stories
di: Rooein, Donya, et al.
Pubblicazione: (2025)
di: Rooein, Donya, et al.
Pubblicazione: (2025)
Educators' Perceptions of Large Language Models as Tutors: Comparing Human and AI Tutors in a Blind Text-only Setting
di: Chowdhury, Sankalan Pal, et al.
Pubblicazione: (2025)
di: Chowdhury, Sankalan Pal, et al.
Pubblicazione: (2025)
PATS: Personality-Aware Teaching Strategies with Large Language Model Tutors
di: Rooein, Donya, et al.
Pubblicazione: (2026)
di: Rooein, Donya, et al.
Pubblicazione: (2026)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
di: Vidgen, Bertie, et al.
Pubblicazione: (2023)
di: Vidgen, Bertie, et al.
Pubblicazione: (2023)
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
di: Pernisi, Fabio, et al.
Pubblicazione: (2024)
di: Pernisi, Fabio, et al.
Pubblicazione: (2024)
Censorship in Democracy
di: Caesmann, Marcel, et al.
Pubblicazione: (2024)
di: Caesmann, Marcel, et al.
Pubblicazione: (2024)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
di: Ni, Jingwei, et al.
Pubblicazione: (2025)
di: Ni, Jingwei, et al.
Pubblicazione: (2025)
Reseña "Lento en la sombra. Ensayos sobre literatura, arte y cine" de Handke, Peter
di: Alejandro Goldzycher
Pubblicazione: (2014)
di: Alejandro Goldzycher
Pubblicazione: (2014)
Can I introduce my boyfriend to my grandmother? Evaluating Large Language Models Capabilities on Iranian Social Norm Classification
di: Saffari, Hamidreza, et al.
Pubblicazione: (2024)
di: Saffari, Hamidreza, et al.
Pubblicazione: (2024)
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2025)
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2025)
Multilingual Performance Biases of Large Language Models in Education
di: Gupta, Vansh, et al.
Pubblicazione: (2025)
di: Gupta, Vansh, et al.
Pubblicazione: (2025)
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
di: Orlikowski, Matthias, et al.
Pubblicazione: (2023)
di: Orlikowski, Matthias, et al.
Pubblicazione: (2023)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
di: Russo, Giuseppe, et al.
Pubblicazione: (2025)
di: Russo, Giuseppe, et al.
Pubblicazione: (2025)
Twists, Humps, and Pebbles: Multilingual Speech Recognition Models Exhibit Gender Performance Gaps
di: Attanasio, Giuseppe, et al.
Pubblicazione: (2024)
di: Attanasio, Giuseppe, et al.
Pubblicazione: (2024)
Why human-AI relationships need socioaffective alignment
di: Kirk, Hannah Rose, et al.
Pubblicazione: (2025)
di: Kirk, Hannah Rose, et al.
Pubblicazione: (2025)
SpiritRAG: A Q&A System for Religion and Spirituality in the United Nations Archive
di: Gao, Yingqiang, et al.
Pubblicazione: (2025)
di: Gao, Yingqiang, et al.
Pubblicazione: (2025)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
di: Styles, Olly, et al.
Pubblicazione: (2024)
di: Styles, Olly, et al.
Pubblicazione: (2024)
Classist Tools: Social Class Correlates with Performance in NLP
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2024)
di: Curry, Amanda Cercas, et al.
Pubblicazione: (2024)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
di: Zengaffinen, Yanick, et al.
Pubblicazione: (2026)
di: Zengaffinen, Yanick, et al.
Pubblicazione: (2026)
Wisdom of Instruction-Tuned Language Model Crowds. Exploring Model Label Variation
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2023)
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2023)
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
di: Kirk, Hannah Rose, et al.
Pubblicazione: (2024)
di: Kirk, Hannah Rose, et al.
Pubblicazione: (2024)
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
di: de Araujo, Pedro Henrique Luz, et al.
Pubblicazione: (2025)
di: de Araujo, Pedro Henrique Luz, et al.
Pubblicazione: (2025)
AustroTox: A Dataset for Target-Based Austrian German Offensive Language Detection
di: Pachinger, Pia, et al.
Pubblicazione: (2024)
di: Pachinger, Pia, et al.
Pubblicazione: (2024)
The AI Consumer Index (ACE)
di: Benchek, Julien, et al.
Pubblicazione: (2025)
di: Benchek, Julien, et al.
Pubblicazione: (2025)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
di: Kirk, Hannah Rose, et al.
Pubblicazione: (2026)
di: Kirk, Hannah Rose, et al.
Pubblicazione: (2026)
Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
di: Kirk, Hannah Rose, et al.
Pubblicazione: (2025)
di: Kirk, Hannah Rose, et al.
Pubblicazione: (2025)
Diffusion Language Models Are Natively Length-Aware
di: Rossi, Vittorio, et al.
Pubblicazione: (2026)
di: Rossi, Vittorio, et al.
Pubblicazione: (2026)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
di: Orlikowski, Matthias, et al.
Pubblicazione: (2025)
di: Orlikowski, Matthias, et al.
Pubblicazione: (2025)
Pretrained Multilingual Transformers Reveal Quantitative Distance Between Human Languages
di: Zhao, Yue, et al.
Pubblicazione: (2026)
di: Zhao, Yue, et al.
Pubblicazione: (2026)
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
di: Xiong, Chenfei, et al.
Pubblicazione: (2025)
di: Xiong, Chenfei, et al.
Pubblicazione: (2025)
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
di: Salhan, Suchir, et al.
Pubblicazione: (2025)
di: Salhan, Suchir, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
di: Rooein, Donya, et al.
Pubblicazione: (2024) -
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
di: Röttger, Paul, et al.
Pubblicazione: (2024) -
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
di: Röttger, Paul, et al.
Pubblicazione: (2023) -
Conversations as a Source for Teaching Scientific Concepts at Different Education Levels
di: Rooein, Donya, et al.
Pubblicazione: (2024) -
Classification is a RAG problem: A case study on hate speech detection
di: Willats, Richard, et al.
Pubblicazione: (2025)