Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
Fuente:
arXiv
Saved in:
| Main Authors: | Tuck, Bryan E., Verma, Rakesh M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ten Words Only Still Help: Improving Black-Box AI-Generated Text Detection via Proxy-Guided Efficient Re-Sampling
by: Shi, Yuhui, et al.
Published: (2024)
by: Shi, Yuhui, et al.
Published: (2024)
Adversarial Attacks on AI-Generated Text Detection Models: A Token Probability-Based Approach Using Embeddings
by: Kadhim, Ahmed K., et al.
Published: (2025)
by: Kadhim, Ahmed K., et al.
Published: (2025)
GuideWalk: A Novel Graph-Based Word Embedding for Enhanced Text Classification
by: Mohammed, Sarmad N., et al.
Published: (2024)
by: Mohammed, Sarmad N., et al.
Published: (2024)
Word Embeddings Are Steers for Language Models
by: Han, Chi, et al.
Published: (2023)
by: Han, Chi, et al.
Published: (2023)
Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
by: Xiao, Feng, et al.
Published: (2025)
by: Xiao, Feng, et al.
Published: (2025)
Static Word Embeddings for Sentence Semantic Representation
by: Wada, Takashi, et al.
Published: (2025)
by: Wada, Takashi, et al.
Published: (2025)
Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models
by: Liang, Haoyu, et al.
Published: (2025)
by: Liang, Haoyu, et al.
Published: (2025)
SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering
by: Li, Xiaopeng, et al.
Published: (2024)
by: Li, Xiaopeng, et al.
Published: (2024)
Visualising Information Flow in Word Embeddings with Diffusion Tensor Imaging
by: Fabian, Thomas
Published: (2026)
by: Fabian, Thomas
Published: (2026)
RKadiyala at SemEval-2024 Task 8: Black-Box Word-Level Text Boundary Detection in Partially Machine Generated Texts
by: Kadiyala, Ram Mohan Rao
Published: (2024)
by: Kadiyala, Ram Mohan Rao
Published: (2024)
Algorithm Selection with Zero Domain Knowledge via Text Embeddings
by: Szeider, Stefan
Published: (2026)
by: Szeider, Stefan
Published: (2026)
Merging Text Transformer Models from Different Initializations
by: Verma, Neha, et al.
Published: (2024)
by: Verma, Neha, et al.
Published: (2024)
Exploring Gradient-Guided Masked Language Model to Detect Textual Adversarial Attacks
by: Zhang, Xiaomei, et al.
Published: (2025)
by: Zhang, Xiaomei, et al.
Published: (2025)
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022)
by: Lopardo, Gianluigi, et al.
Published: (2022)
Induced Numerical Instability: Hidden Costs in Multimodal Large Language Models
by: Wong, Wai Tuck, et al.
Published: (2026)
by: Wong, Wai Tuck, et al.
Published: (2026)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
by: Fu, Tingchen, et al.
Published: (2025)
by: Fu, Tingchen, et al.
Published: (2025)
Deceiving Question-Answering Models: A Hybrid Word-Level Adversarial Approach
by: Li, Jiyao, et al.
Published: (2024)
by: Li, Jiyao, et al.
Published: (2024)
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022)
by: Lopardo, Gianluigi, et al.
Published: (2022)
Interpretable Topic Extraction and Word Embedding Learning using row-stochastic DEDICOM
by: Hillebrand, Lars, et al.
Published: (2025)
by: Hillebrand, Lars, et al.
Published: (2025)
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
by: Wang, Dongwei, et al.
Published: (2024)
by: Wang, Dongwei, et al.
Published: (2024)
What is in a name? Mitigating Name Bias in Text Embeddings via Anonymization
by: Manchanda, Sahil, et al.
Published: (2025)
by: Manchanda, Sahil, et al.
Published: (2025)
Hakim: Farsi Text Embedding Model
by: Sarmadi, Mehran, et al.
Published: (2025)
by: Sarmadi, Mehran, et al.
Published: (2025)
AnglE-optimized Text Embeddings
by: Li, Xianming, et al.
Published: (2023)
by: Li, Xianming, et al.
Published: (2023)
RobustSentEmbed: Robust Sentence Embeddings Using Adversarial Self-Supervised Contrastive Learning
by: Asl, Javad Rafiei, et al.
Published: (2024)
by: Asl, Javad Rafiei, et al.
Published: (2024)
Unmasking the Imposters: How Censorship and Domain Adaptation Affect the Detection of Machine-Generated Tweets
by: Tuck, Bryan E., et al.
Published: (2024)
by: Tuck, Bryan E., et al.
Published: (2024)
Words as Beacons: Guiding RL Agents with High-Level Language Prompts
by: Ruiz-Gonzalez, Unai, et al.
Published: (2024)
by: Ruiz-Gonzalez, Unai, et al.
Published: (2024)
ToxiGAN: Toxic Data Augmentation via LLM-Guided Directional Adversarial Generation
by: Li, Peiran, et al.
Published: (2026)
by: Li, Peiran, et al.
Published: (2026)
Words or Vision: Do Vision-Language Models Have Blind Faith in Text?
by: Deng, Ailin, et al.
Published: (2025)
by: Deng, Ailin, et al.
Published: (2025)
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
by: Li, Chenliang, et al.
Published: (2025)
by: Li, Chenliang, et al.
Published: (2025)
Factor Augmented Supervised Learning with Text Embeddings
by: Luo, Zhanye, et al.
Published: (2025)
by: Luo, Zhanye, et al.
Published: (2025)
Empirical Characterization of Rationale Stability Under Controlled Perturbations for Explainable Pattern Recognition
by: Sakib, Abu Noman Md, et al.
Published: (2026)
by: Sakib, Abu Noman Md, et al.
Published: (2026)
Empowering Diffusion Models on the Embedding Space for Text Generation
by: Gao, Zhujin, et al.
Published: (2022)
by: Gao, Zhujin, et al.
Published: (2022)
Magic Words or Methodical Work? Challenging Conventional Wisdom in LLM-Based Political Text Annotation
by: McLaren, Lorcan, et al.
Published: (2026)
by: McLaren, Lorcan, et al.
Published: (2026)
Fine-tuning Language Models with Generative Adversarial Reward Modelling
by: Yu, Zhang Ze, et al.
Published: (2023)
by: Yu, Zhang Ze, et al.
Published: (2023)
Bolster Hallucination Detection via Prompt-Guided Data Augmentation
by: Li, Wenyun, et al.
Published: (2025)
by: Li, Wenyun, et al.
Published: (2025)
Alert-ME: An Explainability-Driven Defense Against Adversarial Examples in Transformer-Based Text Classification
by: Sabir, Bushra, et al.
Published: (2023)
by: Sabir, Bushra, et al.
Published: (2023)
A General Framework for Producing Interpretable Semantic Text Embeddings
by: Sun, Yiqun, et al.
Published: (2024)
by: Sun, Yiqun, et al.
Published: (2024)
HU at SemEval-2024 Task 8A: Can Contrastive Learning Learn Embeddings to Detect Machine-Generated Text?
by: Dipta, Shubhashis Roy, et al.
Published: (2024)
by: Dipta, Shubhashis Roy, et al.
Published: (2024)
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
by: Lu, Ning, et al.
Published: (2023)
by: Lu, Ning, et al.
Published: (2023)
Augmenting Math Word Problems via Iterative Question Composing
by: Liu, Haoxiong, et al.
Published: (2024)
by: Liu, Haoxiong, et al.
Published: (2024)
Similar Items
-
Ten Words Only Still Help: Improving Black-Box AI-Generated Text Detection via Proxy-Guided Efficient Re-Sampling
by: Shi, Yuhui, et al.
Published: (2024) -
Adversarial Attacks on AI-Generated Text Detection Models: A Token Probability-Based Approach Using Embeddings
by: Kadhim, Ahmed K., et al.
Published: (2025) -
GuideWalk: A Novel Graph-Based Word Embedding for Enhanced Text Classification
by: Mohammed, Sarmad N., et al.
Published: (2024) -
Word Embeddings Are Steers for Language Models
by: Han, Chi, et al.
Published: (2023) -
Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
by: Xiao, Feng, et al.
Published: (2025)