The Problem with Safety Classification is not just the Models
Fuente:
arXiv
Salvato in:
| Autore principale: | Vajjala, Sowmya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
IndicGEC: Powerful Models, or a Measurement Mirage?
di: Vajjala, Sowmya
Pubblicazione: (2025)
di: Vajjala, Sowmya
Pubblicazione: (2025)
Text Classification in the LLM Era -- Where do we stand?
di: Vajjala, Sowmya, et al.
Pubblicazione: (2025)
di: Vajjala, Sowmya, et al.
Pubblicazione: (2025)
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
Does Synthetic Data Help Named Entity Recognition for Low-Resource Languages?
di: Kamath, Gaurav, et al.
Pubblicazione: (2025)
di: Kamath, Gaurav, et al.
Pubblicazione: (2025)
Dravidian language family through Universal Dependencies lens
di: Rama, Taraka, et al.
Pubblicazione: (2024)
di: Rama, Taraka, et al.
Pubblicazione: (2024)
Annotation Errors and NER: A Study with OntoNotes 5.0
di: Bernier-Colborne, Gabriel, et al.
Pubblicazione: (2024)
di: Bernier-Colborne, Gabriel, et al.
Pubblicazione: (2024)
Scope Ambiguities in Large Language Models
di: Kamath, Gaurav, et al.
Pubblicazione: (2024)
di: Kamath, Gaurav, et al.
Pubblicazione: (2024)
Test Set Quality in Multilingual LLM Evaluation
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
Opportunities and Challenges of LLMs in Education: An NLP Perspective
di: Vajjala, Sowmya, et al.
Pubblicazione: (2025)
di: Vajjala, Sowmya, et al.
Pubblicazione: (2025)
LLMs in Education: Novel Perspectives, Challenges, and Opportunities
di: Alhafni, Bashar, et al.
Pubblicazione: (2024)
di: Alhafni, Bashar, et al.
Pubblicazione: (2024)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
di: Vishnubhotla, Krishnapriya, et al.
Pubblicazione: (2026)
di: Vishnubhotla, Krishnapriya, et al.
Pubblicazione: (2026)
Making Metadata More FAIR Using Large Language Models
di: Sundaram, Sowmya S., et al.
Pubblicazione: (2023)
di: Sundaram, Sowmya S., et al.
Pubblicazione: (2023)
Are Emergent Abilities in Large Language Models just In-Context Learning?
di: Lu, Sheng, et al.
Pubblicazione: (2023)
di: Lu, Sheng, et al.
Pubblicazione: (2023)
CoCoTen: Detecting Adversarial Inputs to Large Language Models through Latent Space Features of Contextual Co-occurrence Tensors
di: Kadali, Sri Durga Sai Sowmya, et al.
Pubblicazione: (2025)
di: Kadali, Sri Durga Sai Sowmya, et al.
Pubblicazione: (2025)
Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models
di: Kadali, Sri Durga Sai Sowmya, et al.
Pubblicazione: (2026)
di: Kadali, Sri Durga Sai Sowmya, et al.
Pubblicazione: (2026)
Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?
di: Kadali, Sri Durga Sai Sowmya, et al.
Pubblicazione: (2025)
di: Kadali, Sri Durga Sai Sowmya, et al.
Pubblicazione: (2025)
Comparative Analysis of Transformer Models in Disaster Tweet Classification for Public Safety
di: Zisad, Sharif Noor, et al.
Pubblicazione: (2025)
di: Zisad, Sharif Noor, et al.
Pubblicazione: (2025)
Lightweight Safety Classification Using Pruned Language Models
di: Sawtell, Mason, et al.
Pubblicazione: (2024)
di: Sawtell, Mason, et al.
Pubblicazione: (2024)
Translating speech with just images
di: Oneata, Dan, et al.
Pubblicazione: (2024)
di: Oneata, Dan, et al.
Pubblicazione: (2024)
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
di: Sun, Yuhong, et al.
Pubblicazione: (2025)
di: Sun, Yuhong, et al.
Pubblicazione: (2025)
SafetyBench: Evaluating the Safety of Large Language Models
di: Zhang, Zhexin, et al.
Pubblicazione: (2023)
di: Zhang, Zhexin, et al.
Pubblicazione: (2023)
An Answer is just the Start: Related Insight Generation for Open-Ended Document-Grounded QA
di: Sharma, Saransh, et al.
Pubblicazione: (2026)
di: Sharma, Saransh, et al.
Pubblicazione: (2026)
Classification of Safety Events at Nuclear Sites using Large Language Models
di: de Costa, Mishca, et al.
Pubblicazione: (2024)
di: de Costa, Mishca, et al.
Pubblicazione: (2024)
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment
di: Imperial, Joseph Marvin, et al.
Pubblicazione: (2025)
di: Imperial, Joseph Marvin, et al.
Pubblicazione: (2025)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
di: Rios-Sialer, Ian
Pubblicazione: (2026)
di: Rios-Sialer, Ian
Pubblicazione: (2026)
Sequential Classification of Aviation Safety Occurrences with Natural Language Processing
di: Nanyonga, Aziida, et al.
Pubblicazione: (2025)
di: Nanyonga, Aziida, et al.
Pubblicazione: (2025)
ThaiSafetyBench: Assessing Language Model Safety in Thai Cultural Contexts
di: Ukarapol, Trapoom, et al.
Pubblicazione: (2026)
di: Ukarapol, Trapoom, et al.
Pubblicazione: (2026)
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models
di: In, Yeonjun, et al.
Pubblicazione: (2025)
di: In, Yeonjun, et al.
Pubblicazione: (2025)
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model
di: Lou, Xinyue, et al.
Pubblicazione: (2025)
di: Lou, Xinyue, et al.
Pubblicazione: (2025)
Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
di: Zhao, Weixiang, et al.
Pubblicazione: (2024)
di: Zhao, Weixiang, et al.
Pubblicazione: (2024)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
di: Zhou, Xuhui, et al.
Pubblicazione: (2024)
di: Zhou, Xuhui, et al.
Pubblicazione: (2024)
KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models
di: Zaghouani, Wajdi, et al.
Pubblicazione: (2026)
di: Zaghouani, Wajdi, et al.
Pubblicazione: (2026)
Safety in Large Reasoning Models: A Survey
di: Wang, Cheng, et al.
Pubblicazione: (2025)
di: Wang, Cheng, et al.
Pubblicazione: (2025)
Mitigating Exaggerated Safety in Large Language Models
di: Ray, Ruchira, et al.
Pubblicazione: (2024)
di: Ray, Ruchira, et al.
Pubblicazione: (2024)
Tracing Facts or just Copies? A critical investigation of the Competitions of Mechanisms in Large Language Models
di: Campregher, Dante, et al.
Pubblicazione: (2025)
di: Campregher, Dante, et al.
Pubblicazione: (2025)
Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
di: Hazra, Rima, et al.
Pubblicazione: (2024)
di: Hazra, Rima, et al.
Pubblicazione: (2024)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
di: Vidgen, Bertie, et al.
Pubblicazione: (2023)
di: Vidgen, Bertie, et al.
Pubblicazione: (2023)
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
di: Bianchi, Federico, et al.
Pubblicazione: (2023)
di: Bianchi, Federico, et al.
Pubblicazione: (2023)
Prompt-Induced Score Variance in Zero-Shot Binary Vision-Language Safety Classification
di: Weng, Charles, et al.
Pubblicazione: (2026)
di: Weng, Charles, et al.
Pubblicazione: (2026)
Documenti analoghi
-
IndicGEC: Powerful Models, or a Measurement Mirage?
di: Vajjala, Sowmya
Pubblicazione: (2025) -
Text Classification in the LLM Era -- Where do we stand?
di: Vajjala, Sowmya, et al.
Pubblicazione: (2025) -
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025) -
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025) -
Does Synthetic Data Help Named Entity Recognition for Low-Resource Languages?
di: Kamath, Gaurav, et al.
Pubblicazione: (2025)