Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Walkowiak, Paweł, Klonowski, Marek, Oleksy, Marcin, Janz, Arkadiusz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Personalized Large Language Models
von: Woźniak, Stanisław, et al.
Veröffentlicht: (2024)
von: Woźniak, Stanisław, et al.
Veröffentlicht: (2024)
Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language
von: Wojtasik, Konrad, et al.
Veröffentlicht: (2023)
von: Wojtasik, Konrad, et al.
Veröffentlicht: (2023)
Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
von: Kubis, Marek, et al.
Veröffentlicht: (2025)
von: Kubis, Marek, et al.
Veröffentlicht: (2025)
How to Protect Models against Adversarial Unlearning?
von: Jasiorski, Patryk, et al.
Veröffentlicht: (2025)
von: Jasiorski, Patryk, et al.
Veröffentlicht: (2025)
Developing PUGG for Polish: A Modern Approach to KBQA, MRC, and IR Dataset Construction
von: Sawczyn, Albert, et al.
Veröffentlicht: (2024)
von: Sawczyn, Albert, et al.
Veröffentlicht: (2024)
On Adversarial Robustness and Out-of-Distribution Robustness of Large Language Models
von: Yang, April, et al.
Veröffentlicht: (2024)
von: Yang, April, et al.
Veröffentlicht: (2024)
PLLuM: A Family of Polish Large Language Models
von: Kocoń, Jan, et al.
Veröffentlicht: (2025)
von: Kocoń, Jan, et al.
Veröffentlicht: (2025)
The PLLuM Instruction Corpus
von: Pęzik, Piotr, et al.
Veröffentlicht: (2025)
von: Pęzik, Piotr, et al.
Veröffentlicht: (2025)
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning
von: Yu, Simon, et al.
Veröffentlicht: (2024)
von: Yu, Simon, et al.
Veröffentlicht: (2024)
Cross-Platform Evaluation of Large Language Model Safety in Pediatric Consultations: Evolution of Adversarial Robustness and the Scale Paradox
von: Zolfaghari, Vahideh
Veröffentlicht: (2025)
von: Zolfaghari, Vahideh
Veröffentlicht: (2025)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
von: Das, Nilanjana, et al.
Veröffentlicht: (2026)
von: Das, Nilanjana, et al.
Veröffentlicht: (2026)
AI and the Law: Evaluating ChatGPT's Performance in Legal Classification
von: Weichbroth, Pawel
Veröffentlicht: (2025)
von: Weichbroth, Pawel
Veröffentlicht: (2025)
More Women, Same Stereotypes: Unpacking the Gender Bias Paradox in Large Language Models
von: Chen, Evan, et al.
Veröffentlicht: (2025)
von: Chen, Evan, et al.
Veröffentlicht: (2025)
Mechanistic Behavior Editing of Language Models
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
von: Petrova, Nora, et al.
Veröffentlicht: (2026)
von: Petrova, Nora, et al.
Veröffentlicht: (2026)
Mechanistic Indicators of Understanding in Large Language Models
von: Beckmann, Pierre, et al.
Veröffentlicht: (2025)
von: Beckmann, Pierre, et al.
Veröffentlicht: (2025)
Mechanistic Origin of Moral Indifference in Language Models
von: Li, Lingyu, et al.
Veröffentlicht: (2026)
von: Li, Lingyu, et al.
Veröffentlicht: (2026)
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
von: Goldman, Omer, et al.
Veröffentlicht: (2024)
von: Goldman, Omer, et al.
Veröffentlicht: (2024)
StylOch at PAN: Gradient-Boosted Trees with Frequency-Based Stylometric Features
von: Ochab, Jeremi K., et al.
Veröffentlicht: (2025)
von: Ochab, Jeremi K., et al.
Veröffentlicht: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
Adversarial Style Augmentation via Large Language Model for Robust Fake News Detection
von: Park, Sungwon, et al.
Veröffentlicht: (2024)
von: Park, Sungwon, et al.
Veröffentlicht: (2024)
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
Towards Quantifying Commonsense Reasoning with Mechanistic Insights
von: Joshi, Abhinav, et al.
Veröffentlicht: (2025)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Emotion Inference in Large Language Models
von: Tak, Ala N., et al.
Veröffentlicht: (2025)
von: Tak, Ala N., et al.
Veröffentlicht: (2025)
Toward Mechanistic Explanation of Deductive Reasoning in Language Models
von: Maltoni, Davide, et al.
Veröffentlicht: (2025)
von: Maltoni, Davide, et al.
Veröffentlicht: (2025)
Mechanistic Decoding of Cognitive Constructs in Large Language Models
von: Shou, Yitong, et al.
Veröffentlicht: (2026)
von: Shou, Yitong, et al.
Veröffentlicht: (2026)
Evaluating the Retrieval Robustness of Large Language Models
von: Cao, Shuyang, et al.
Veröffentlicht: (2025)
von: Cao, Shuyang, et al.
Veröffentlicht: (2025)
Self-training Language Models for Arithmetic Reasoning
von: Kadlčík, Marek, et al.
Veröffentlicht: (2024)
von: Kadlčík, Marek, et al.
Veröffentlicht: (2024)
Evaluating Robustness of Generative Search Engine on Adversarial Factual Questions
von: Hu, Xuming, et al.
Veröffentlicht: (2024)
von: Hu, Xuming, et al.
Veröffentlicht: (2024)
SlovKE: A Large-Scale Dataset and LLM Evaluation for Slovak Keyphrase Extraction
von: Števaňák, David, et al.
Veröffentlicht: (2026)
von: Števaňák, David, et al.
Veröffentlicht: (2026)
Reasoning Robustness of LLMs to Adversarial Typographical Errors
von: Gan, Esther, et al.
Veröffentlicht: (2024)
von: Gan, Esther, et al.
Veröffentlicht: (2024)
Defensive Dual Masking for Robust Adversarial Defense
von: Yang, Wangli, et al.
Veröffentlicht: (2024)
von: Yang, Wangli, et al.
Veröffentlicht: (2024)
Meta-Reasoning Improves Tool Use in Large Language Models
von: Alazraki, Lisa, et al.
Veröffentlicht: (2024)
von: Alazraki, Lisa, et al.
Veröffentlicht: (2024)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
von: Chen, Zijian, et al.
Veröffentlicht: (2025)
von: Chen, Zijian, et al.
Veröffentlicht: (2025)
Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation
von: Cantini, Riccardo, et al.
Veröffentlicht: (2024)
von: Cantini, Riccardo, et al.
Veröffentlicht: (2024)
SelfPrompt: Autonomously Evaluating LLM Robustness via Domain-Constrained Knowledge Guidelines and Refined Adversarial Prompts
von: Pei, Aihua, et al.
Veröffentlicht: (2024)
von: Pei, Aihua, et al.
Veröffentlicht: (2024)
Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice Options
von: Góral, Gracjan, et al.
Veröffentlicht: (2024)
von: Góral, Gracjan, et al.
Veröffentlicht: (2024)
Are AI-Generated Text Detectors Robust to Adversarial Perturbations?
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
von: Huang, Guanhua, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Personalized Large Language Models
von: Woźniak, Stanisław, et al.
Veröffentlicht: (2024) -
Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety
von: Janiak, Denis, et al.
Veröffentlicht: (2025) -
BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language
von: Wojtasik, Konrad, et al.
Veröffentlicht: (2023) -
Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
von: Kubis, Marek, et al.
Veröffentlicht: (2025) -
How to Protect Models against Adversarial Unlearning?
von: Jasiorski, Patryk, et al.
Veröffentlicht: (2025)