Adversarial Lens: Exploiting Attention Layers to Generate Adversarial Examples for Evaluation
Fuente:
arXiv
Saved in:
| Main Author: | Dhole, Kaustubh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multi-Encoder Frozen-Decoder Approach for Fine-Tuning Large Language Models
by: Dhole, Kaustubh D.
Published: (2025)
by: Dhole, Kaustubh D.
Published: (2025)
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
by: Dhole, Kaustubh D.
Published: (2026)
by: Dhole, Kaustubh D.
Published: (2026)
RubricRAG: Towards Interpretable and Reliable LLM Evaluation via Domain Knowledge Retrieval for Rubric Generation
by: Dhole, Kaustubh D., et al.
Published: (2026)
by: Dhole, Kaustubh D., et al.
Published: (2026)
ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges
by: Dhole, Kaustubh D., et al.
Published: (2024)
by: Dhole, Kaustubh D., et al.
Published: (2024)
To Retrieve or Not to Retrieve? Uncertainty Detection for Dynamic Retrieval Augmented Generation
by: Dhole, Kaustubh D.
Published: (2025)
by: Dhole, Kaustubh D.
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
KAUCUS: Knowledge Augmented User Simulators for Training Language Model Assistants
by: Dhole, Kaustubh D.
Published: (2024)
by: Dhole, Kaustubh D.
Published: (2024)
A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness
by: Machlovi, Naseem, et al.
Published: (2025)
by: Machlovi, Naseem, et al.
Published: (2025)
RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
by: Goldstein, Daniel, et al.
Published: (2025)
by: Goldstein, Daniel, et al.
Published: (2025)
Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Synthius-Mem: Brain-Inspired Hallucination-Resistant Persona Memory Achieving 94.4% Memory Accuracy and 99.6% Adversarial Robustness on LoCoMo
by: Gadzhiev, Artem, et al.
Published: (2026)
by: Gadzhiev, Artem, et al.
Published: (2026)
Adversarial Attacks on Large Language Models Using Regularized Relaxation
by: Chacko, Samuel Jacob, et al.
Published: (2024)
by: Chacko, Samuel Jacob, et al.
Published: (2024)
Evaluating Long Range Dependency Handling in Code Generation LLMs
by: Assogba, Yannick, et al.
Published: (2024)
by: Assogba, Yannick, et al.
Published: (2024)
On Adversarial Examples for Text Classification by Perturbing Latent Representations
by: Sooksatra, Korn, et al.
Published: (2024)
by: Sooksatra, Korn, et al.
Published: (2024)
Consistency Evaluation of News Article Summaries Generated by Large (and Small) Language Models
by: Gilhuly, Colleen, et al.
Published: (2025)
by: Gilhuly, Colleen, et al.
Published: (2025)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
by: Vieira, Inês, et al.
Published: (2026)
by: Vieira, Inês, et al.
Published: (2026)
Forget Attention: Importance-Aware Attention Is All You Need
by: Shin, Soohyeong, et al.
Published: (2026)
by: Shin, Soohyeong, et al.
Published: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
by: Höth, Max Henning, et al.
Published: (2026)
by: Höth, Max Henning, et al.
Published: (2026)
Revisiting Intermediate-Layer Matching in Knowledge Distillation: Layer-Selection Strategy Doesn't Matter (Much)
by: Yu, Zony, et al.
Published: (2025)
by: Yu, Zony, et al.
Published: (2025)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
by: Fu, Tianyu, et al.
Published: (2024)
by: Fu, Tianyu, et al.
Published: (2024)
Bypassing DARCY Defense: Indistinguishable Universal Adversarial Triggers
by: Peng, Zuquan, et al.
Published: (2024)
by: Peng, Zuquan, et al.
Published: (2024)
No Free Swap: Protocol-Dependent Layer Redundancy in Transformers
by: Garcia, Gabriel
Published: (2026)
by: Garcia, Gabriel
Published: (2026)
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
by: Ghosh, Shubhra, et al.
Published: (2025)
by: Ghosh, Shubhra, et al.
Published: (2025)
Exploiting Novel GPT-4 APIs
by: Pelrine, Kellin, et al.
Published: (2023)
by: Pelrine, Kellin, et al.
Published: (2023)
Neural Multimodal Topic Modeling: A Comprehensive Evaluation
by: González-Pizarro, Felipe, et al.
Published: (2024)
by: González-Pizarro, Felipe, et al.
Published: (2024)
Benchmarking Cognitive Biases in Large Language Models as Evaluators
by: Koo, Ryan, et al.
Published: (2023)
by: Koo, Ryan, et al.
Published: (2023)
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
InsertRank: LLMs can reason over BM25 scores to Improve Listwise Reranking
by: Seetharaman, Rahul, et al.
Published: (2025)
by: Seetharaman, Rahul, et al.
Published: (2025)
LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation
by: Bishop, Jennifer A, et al.
Published: (2023)
by: Bishop, Jennifer A, et al.
Published: (2023)
PlanRAG: A Plan-then-Retrieval Augmented Generation for Generative Large Language Models as Decision Makers
by: Lee, Myeonghwa, et al.
Published: (2024)
by: Lee, Myeonghwa, et al.
Published: (2024)
Danoliteracy of Generative Large Language Models
by: Holm, Søren Vejlgaard, et al.
Published: (2024)
by: Holm, Søren Vejlgaard, et al.
Published: (2024)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
by: Nainani, Jatin, et al.
Published: (2024)
by: Nainani, Jatin, et al.
Published: (2024)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph
by: Wang, Fali, et al.
Published: (2025)
by: Wang, Fali, et al.
Published: (2025)
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
by: Quevedo, Ernesto, et al.
Published: (2024)
by: Quevedo, Ernesto, et al.
Published: (2024)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
by: Orgad, Hadas, et al.
Published: (2026)
by: Orgad, Hadas, et al.
Published: (2026)
Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents
by: Kim, San, et al.
Published: (2024)
by: Kim, San, et al.
Published: (2024)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
Similar Items
-
A Multi-Encoder Frozen-Decoder Approach for Fine-Tuning Large Language Models
by: Dhole, Kaustubh D.
Published: (2025) -
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
by: Dhole, Kaustubh D.
Published: (2026) -
RubricRAG: Towards Interpretable and Reliable LLM Evaluation via Domain Knowledge Retrieval for Rubric Generation
by: Dhole, Kaustubh D., et al.
Published: (2026) -
ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges
by: Dhole, Kaustubh D., et al.
Published: (2024) -
To Retrieve or Not to Retrieve? Uncertainty Detection for Dynamic Retrieval Augmented Generation
by: Dhole, Kaustubh D.
Published: (2025)