Bridging Robustness and Generalization Against Word Substitution Attacks in NLP via the Growth Bound Matrix Approach
Fuente:
arXiv
Salvato in:
| Autori principali: | Bouri, Mohammed, Saoud, Adnane |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLP
di: Ganesan, Adithya V, et al.
Pubblicazione: (2026)
di: Ganesan, Adithya V, et al.
Pubblicazione: (2026)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
NLP Verification: Towards a General Methodology for Certifying Robustness
di: Casadio, Marco, et al.
Pubblicazione: (2024)
di: Casadio, Marco, et al.
Pubblicazione: (2024)
Enhancing NLP Robustness and Generalization through LLM-Generated Contrast Sets: A Scalable Framework for Systematic Evaluation and Adversarial Training
di: Lin, Hender
Pubblicazione: (2025)
di: Lin, Hender
Pubblicazione: (2025)
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
VNLP: Turkish NLP Package
di: Turker, Meliksah, et al.
Pubblicazione: (2024)
di: Turker, Meliksah, et al.
Pubblicazione: (2024)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
di: Zhou, Andy, et al.
Pubblicazione: (2024)
di: Zhou, Andy, et al.
Pubblicazione: (2024)
Scaling BERT Models for Turkish Automatic Punctuation and Capitalization Correction
di: Saoud, Abdulkader, et al.
Pubblicazione: (2024)
di: Saoud, Abdulkader, et al.
Pubblicazione: (2024)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
di: Fu, Tingchen, et al.
Pubblicazione: (2025)
di: Fu, Tingchen, et al.
Pubblicazione: (2025)
Certified Robustness Under Bounded Levenshtein Distance
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
Deceiving Question-Answering Models: A Hybrid Word-Level Adversarial Approach
di: Li, Jiyao, et al.
Pubblicazione: (2024)
di: Li, Jiyao, et al.
Pubblicazione: (2024)
Indian Legal NLP Benchmarks : A Survey
di: Kalamkar, Prathamesh, et al.
Pubblicazione: (2021)
di: Kalamkar, Prathamesh, et al.
Pubblicazione: (2021)
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models
di: Neitemeier, Pit, et al.
Pubblicazione: (2025)
di: Neitemeier, Pit, et al.
Pubblicazione: (2025)
Interpretable AI for Time-Series: Multi-Model Heatmap Fusion with Global Attention and NLP-Generated Explanations
di: Francis, Jiztom Kavalakkatt, et al.
Pubblicazione: (2025)
di: Francis, Jiztom Kavalakkatt, et al.
Pubblicazione: (2025)
Augmenting Math Word Problems via Iterative Question Composing
di: Liu, Haoxiong, et al.
Pubblicazione: (2024)
di: Liu, Haoxiong, et al.
Pubblicazione: (2024)
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
di: Lu, Ning, et al.
Pubblicazione: (2023)
di: Lu, Ning, et al.
Pubblicazione: (2023)
Named Entity Recognition for Payment Data Using NLP
di: Nayak, Srikumar
Pubblicazione: (2026)
di: Nayak, Srikumar
Pubblicazione: (2026)
HOP to the Next Tasks and Domains for Continual Learning in NLP
di: Michieli, Umberto, et al.
Pubblicazione: (2024)
di: Michieli, Umberto, et al.
Pubblicazione: (2024)
Robust Explanations for User Trust in Enterprise NLP Systems
di: Zhang, Guilin, et al.
Pubblicazione: (2026)
di: Zhang, Guilin, et al.
Pubblicazione: (2026)
Adversarial Attacks on AI-Generated Text Detection Models: A Token Probability-Based Approach Using Embeddings
di: Kadhim, Ahmed K., et al.
Pubblicazione: (2025)
di: Kadhim, Ahmed K., et al.
Pubblicazione: (2025)
Automated Unity Game Template Generation from GDDs via NLP and Multi-Modal LLMs
di: Hassan, Amna
Pubblicazione: (2025)
di: Hassan, Amna
Pubblicazione: (2025)
Does Differential Privacy Impact Bias in Pretrained NLP Models?
di: Islam, Md. Khairul, et al.
Pubblicazione: (2024)
di: Islam, Md. Khairul, et al.
Pubblicazione: (2024)
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
di: Geng, Saibo, et al.
Pubblicazione: (2023)
di: Geng, Saibo, et al.
Pubblicazione: (2023)
Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
Group-Adaptive Adversarial Learning for Robust Fake News Detection Against Malicious Comments
di: Tong, Zhao, et al.
Pubblicazione: (2025)
di: Tong, Zhao, et al.
Pubblicazione: (2025)
Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks
di: Liu, Haoyu, et al.
Pubblicazione: (2026)
di: Liu, Haoyu, et al.
Pubblicazione: (2026)
Bridging AI Innovation and Healthcare Needs: Lessons Learned from Incorporating Modern NLP at The BC Cancer Registry
di: Gondara, Lovedeep, et al.
Pubblicazione: (2025)
di: Gondara, Lovedeep, et al.
Pubblicazione: (2025)
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
di: Yang, Shiping, et al.
Pubblicazione: (2025)
di: Yang, Shiping, et al.
Pubblicazione: (2025)
Ten Words Only Still Help: Improving Black-Box AI-Generated Text Detection via Proxy-Guided Efficient Re-Sampling
di: Shi, Yuhui, et al.
Pubblicazione: (2024)
di: Shi, Yuhui, et al.
Pubblicazione: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
di: Chu, Junjie, et al.
Pubblicazione: (2024)
di: Chu, Junjie, et al.
Pubblicazione: (2024)
Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
di: Li, Yubo, et al.
Pubblicazione: (2025)
di: Li, Yubo, et al.
Pubblicazione: (2025)
From Text to Graph: Leveraging Graph Neural Networks for Enhanced Explainability in NLP
di: Yáñez-Romero, Fabio, et al.
Pubblicazione: (2025)
di: Yáñez-Romero, Fabio, et al.
Pubblicazione: (2025)
SEMFED: Semantic-Aware Resource-Efficient Federated Learning for Heterogeneous NLP Tasks
di: Hussain, Sajid, et al.
Pubblicazione: (2025)
di: Hussain, Sajid, et al.
Pubblicazione: (2025)
Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
di: Tuck, Bryan E., et al.
Pubblicazione: (2025)
di: Tuck, Bryan E., et al.
Pubblicazione: (2025)
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
di: Kim, Minsang, et al.
Pubblicazione: (2026)
di: Kim, Minsang, et al.
Pubblicazione: (2026)
SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering
di: Li, Xiaopeng, et al.
Pubblicazione: (2024)
di: Li, Xiaopeng, et al.
Pubblicazione: (2024)
DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation
di: Li, Pingzhi, et al.
Pubblicazione: (2025)
di: Li, Pingzhi, et al.
Pubblicazione: (2025)
Word Embeddings Are Steers for Language Models
di: Han, Chi, et al.
Pubblicazione: (2023)
di: Han, Chi, et al.
Pubblicazione: (2023)
Documenti analoghi
-
From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLP
di: Ganesan, Adithya V, et al.
Pubblicazione: (2026) -
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
di: Cao, Bochuan, et al.
Pubblicazione: (2023) -
NLP Verification: Towards a General Methodology for Certifying Robustness
di: Casadio, Marco, et al.
Pubblicazione: (2024) -
Enhancing NLP Robustness and Generalization through LLM-Generated Contrast Sets: A Scalable Framework for Systematic Evaluation and Adversarial Training
di: Lin, Hender
Pubblicazione: (2025) -
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)