On Adversarial Examples for Text Classification by Perturbing Latent Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Sooksatra, Korn, Khanal, Bikram, Rivas, Pablo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
Logits of API-Protected LLMs Leak Proprietary Information
by: Finlayson, Matthew, et al.
Published: (2024)
by: Finlayson, Matthew, et al.
Published: (2024)
Shortcuts Arising from Contrast: Effective and Covert Clean-Label Attacks in Prompt-Based Learning
by: Xie, Xiaopeng, et al.
Published: (2024)
by: Xie, Xiaopeng, et al.
Published: (2024)
Dark LLMs: The Growing Threat of Unaligned AI Models
by: Fire, Michael, et al.
Published: (2025)
by: Fire, Michael, et al.
Published: (2025)
Monotonicity as an Architectural Bias for Robust Language Models
by: Cooper, Patrick, et al.
Published: (2026)
by: Cooper, Patrick, et al.
Published: (2026)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
by: Pan, Leyi, et al.
Published: (2024)
by: Pan, Leyi, et al.
Published: (2024)
A Semantic Invariant Robust Watermark for Large Language Models
by: Liu, Aiwei, et al.
Published: (2023)
by: Liu, Aiwei, et al.
Published: (2023)
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
by: Liu, Aiwei, et al.
Published: (2024)
by: Liu, Aiwei, et al.
Published: (2024)
TSDS: Data Selection for Task-Specific Model Finetuning
by: Liu, Zifan, et al.
Published: (2024)
by: Liu, Zifan, et al.
Published: (2024)
An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
by: Consoli, Sergio, et al.
Published: (2025)
by: Consoli, Sergio, et al.
Published: (2025)
Scaling Patterns in Adversarial Alignment: Evidence from Multi-LLM Jailbreak Experiments
by: Nathanson, Samuel, et al.
Published: (2025)
by: Nathanson, Samuel, et al.
Published: (2025)
ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data
by: Seneque, Gareth, et al.
Published: (2026)
by: Seneque, Gareth, et al.
Published: (2026)
From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars
by: Kornilov, Albert, et al.
Published: (2024)
by: Kornilov, Albert, et al.
Published: (2024)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
by: Wang, Yanshu, et al.
Published: (2026)
by: Wang, Yanshu, et al.
Published: (2026)
Generative AI Models: Opportunities and Risks for Industry and Authorities
by: Alt, Tobias, et al.
Published: (2024)
by: Alt, Tobias, et al.
Published: (2024)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
Data and AI governance: Promoting equity, ethics, and fairness in large language models
by: Abhishek, Alok, et al.
Published: (2025)
by: Abhishek, Alok, et al.
Published: (2025)
SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models
by: Abhishek, Alok, et al.
Published: (2026)
by: Abhishek, Alok, et al.
Published: (2026)
BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models
by: Abhishek, Alok, et al.
Published: (2025)
by: Abhishek, Alok, et al.
Published: (2025)
Reasoning Promotes Robustness in Theory of Mind Tasks
by: de Haan, Ian B., et al.
Published: (2026)
by: de Haan, Ian B., et al.
Published: (2026)
RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being
by: Ferdousi, Rahatara, et al.
Published: (2025)
by: Ferdousi, Rahatara, et al.
Published: (2025)
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs
by: Bahar, Atmane Ayoub Mansour, et al.
Published: (2024)
by: Bahar, Atmane Ayoub Mansour, et al.
Published: (2024)
Low-Resource Neural Machine Translation Using Recurrent Neural Networks and Transfer Learning: A Case Study on English-to-Igbo
by: Ekle, Ocheme Anthony, et al.
Published: (2025)
by: Ekle, Ocheme Anthony, et al.
Published: (2025)
Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation
by: Rouzegar, Hamidreza, et al.
Published: (2024)
by: Rouzegar, Hamidreza, et al.
Published: (2024)
BreakFun: Jailbreaking LLMs via Schema Exploitation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
Prompting Encoder Models for Zero-Shot Classification: A Cross-Domain Study in Italian
by: Auriemma, Serena, et al.
Published: (2024)
by: Auriemma, Serena, et al.
Published: (2024)
Large Language Models are Inconsistent and Biased Evaluators
by: Stureborg, Rickard, et al.
Published: (2024)
by: Stureborg, Rickard, et al.
Published: (2024)
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
by: Merves, Tyler H., et al.
Published: (2026)
by: Merves, Tyler H., et al.
Published: (2026)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
by: Demir, M. Mikail, et al.
Published: (2025)
by: Demir, M. Mikail, et al.
Published: (2025)
Technical Report on the Pangram AI-Generated Text Classifier
by: Emi, Bradley, et al.
Published: (2024)
by: Emi, Bradley, et al.
Published: (2024)
Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
by: Xu, Beining, et al.
Published: (2025)
by: Xu, Beining, et al.
Published: (2025)
A Study on Bias Detection and Classification in Natural Language Processing
by: Evans, Ana Sofia, et al.
Published: (2024)
by: Evans, Ana Sofia, et al.
Published: (2024)
Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
Rapid Biomedical Research Classification: The Pandemic PACT Advanced Categorisation Engine
by: Rohanian, Omid, et al.
Published: (2024)
by: Rohanian, Omid, et al.
Published: (2024)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
by: Orgad, Hadas, et al.
Published: (2024)
by: Orgad, Hadas, et al.
Published: (2024)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
by: Platt, Nolan, et al.
Published: (2025)
by: Platt, Nolan, et al.
Published: (2025)
GSDFuse: Capturing Cognitive Inconsistencies from Multi-Dimensional Weak Signals in Social Media Steganalysis
by: Huang, Kaibo, et al.
Published: (2025)
by: Huang, Kaibo, et al.
Published: (2025)
Vibe-Creation: The Epistemology of Human-AI Emergent Cognition
by: Levin, Ilya
Published: (2026)
by: Levin, Ilya
Published: (2026)
ADE: Adaptive Dictionary Embeddings -- Scaling Multi-Anchor Representations to Large Language Models
by: Demirci, Orhan, et al.
Published: (2026)
by: Demirci, Orhan, et al.
Published: (2026)
Morphological Synthesizer for Ge'ez Language: Addressing Morphological Complexity and Resource Limitations
by: Gebremariam, Gebrearegawi, et al.
Published: (2025)
by: Gebremariam, Gebrearegawi, et al.
Published: (2025)
Similar Items
-
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
by: Pan, Leyi, et al.
Published: (2025) -
Logits of API-Protected LLMs Leak Proprietary Information
by: Finlayson, Matthew, et al.
Published: (2024) -
Shortcuts Arising from Contrast: Effective and Covert Clean-Label Attacks in Prompt-Based Learning
by: Xie, Xiaopeng, et al.
Published: (2024) -
Dark LLMs: The Growing Threat of Unaligned AI Models
by: Fire, Michael, et al.
Published: (2025) -
Monotonicity as an Architectural Bias for Robust Language Models
by: Cooper, Patrick, et al.
Published: (2026)