Gespeichert in:
| Hauptverfasser: | Kaneko, Masahiro, Talat, Zeerak, Baldwin, Timothy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2510.17006 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2026)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2026)
Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon
von: Koto, Fajri, et al.
Veröffentlicht: (2024)
von: Koto, Fajri, et al.
Veröffentlicht: (2024)
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
A Little Leak Will Sink a Great Ship: Survey of Transparency for Large Language Models from Start to Finish
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
A Capabilities Approach to Studying Bias and Harm in Language Technologies
von: Nigatu, Hellina Hailu, et al.
Veröffentlicht: (2024)
von: Nigatu, Hellina Hailu, et al.
Veröffentlicht: (2024)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
The Gaps between Pre-train and Downstream Settings in Bias Evaluation and Debiasing
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
Beyond the Resumé: A Rubric-Aware Automatic Interview System for Information Elicitation
von: Stuart, Harry, et al.
Veröffentlicht: (2026)
von: Stuart, Harry, et al.
Veröffentlicht: (2026)
Eagle: Ethical Dataset Given from Real Interactions
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
Subjective $\textit{Isms}$? On the Danger of Conflating Hate and Offence in Abusive Language Detection
von: Curry, Amanda Cercas, et al.
Veröffentlicht: (2024)
von: Curry, Amanda Cercas, et al.
Veröffentlicht: (2024)
Impoverished Language Technology: The Lack of (Social) Class in NLP
von: Curry, Amanda Cercas, et al.
Veröffentlicht: (2024)
von: Curry, Amanda Cercas, et al.
Veröffentlicht: (2024)
Understanding "Democratization" in NLP and ML Research
von: Subramonian, Arjun, et al.
Veröffentlicht: (2024)
von: Subramonian, Arjun, et al.
Veröffentlicht: (2024)
Unrequited Emotions: Investigating the Gaps in Motivation and Practice in Speech Emotion Recognition Research
von: Wong, Taryn, et al.
Veröffentlicht: (2026)
von: Wong, Taryn, et al.
Veröffentlicht: (2026)
FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data
von: Abdelkadir, Nuredin Ali, et al.
Veröffentlicht: (2026)
von: Abdelkadir, Nuredin Ali, et al.
Veröffentlicht: (2026)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
Classist Tools: Social Class Correlates with Performance in NLP
von: Curry, Amanda Cercas, et al.
Veröffentlicht: (2024)
von: Curry, Amanda Cercas, et al.
Veröffentlicht: (2024)
Defense against Prompt Injection Attacks via Mixture of Encodings
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2025)
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
von: Fleisig, Eve, et al.
Veröffentlicht: (2024)
von: Fleisig, Eve, et al.
Veröffentlicht: (2024)
Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
Exploitation All the Way Down: Calling out the Root Cause of Bad Online Experiences for Users of the "Majority World"
von: Nigatu, Hellina Hailu, et al.
Veröffentlicht: (2024)
von: Nigatu, Hellina Hailu, et al.
Veröffentlicht: (2024)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
von: Kaneko, Masahiro
Veröffentlicht: (2026)
von: Kaneko, Masahiro
Veröffentlicht: (2026)
Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education
von: Yi, Xin, et al.
Veröffentlicht: (2025)
von: Yi, Xin, et al.
Veröffentlicht: (2025)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
von: Wang, Yidan, et al.
Veröffentlicht: (2025)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
Defending LLMs against Jailbreaking Attacks via Backtranslation
von: Wang, Yihan, et al.
Veröffentlicht: (2024)
von: Wang, Yihan, et al.
Veröffentlicht: (2024)
Social Bias Evaluation for Large Language Models Requires Prompt Variations
von: Hida, Rem, et al.
Veröffentlicht: (2024)
von: Hida, Rem, et al.
Veröffentlicht: (2024)
IYKYK: Using language models to decode extremist cryptolects
von: de Kock, Christine, et al.
Veröffentlicht: (2025)
von: de Kock, Christine, et al.
Veröffentlicht: (2025)
Exploring the Limitations of Detecting Machine-Generated Text
von: Doughman, Jad, et al.
Veröffentlicht: (2024)
von: Doughman, Jad, et al.
Veröffentlicht: (2024)
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
von: Siska, Charlotte, et al.
Veröffentlicht: (2025)
von: Siska, Charlotte, et al.
Veröffentlicht: (2025)
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
von: Koike, Ryuto, et al.
Veröffentlicht: (2023)
von: Koike, Ryuto, et al.
Veröffentlicht: (2023)
Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
Sampling-based Pseudo-Likelihood for Membership Inference Attacks
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
Personal Attribute Leakage in Federated Speech Models
von: Al-Ali, Hamdan, et al.
Veröffentlicht: (2025)
von: Al-Ali, Hamdan, et al.
Veröffentlicht: (2025)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
von: Zhou, Andy, et al.
Veröffentlicht: (2024)
von: Zhou, Andy, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025) -
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2026) -
Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon
von: Koto, Fajri, et al.
Veröffentlicht: (2024) -
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024) -
A Little Leak Will Sink a Great Ship: Survey of Transparency for Large Language Models from Start to Finish
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)