RedHerring Attack: Testing the Reliability of Attack Detection
Fuente:
arXiv
Guardado en:
| Autor principal: | Rusert, Jonathan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VertAttack: Taking advantage of Text Classifiers' horizontal vision
por: Rusert, Jonathan
Publicado: (2024)
por: Rusert, Jonathan
Publicado: (2024)
Overcoming Black-box Attack Inefficiency with Hybrid and Dynamic Select Algorithms
por: Belde, Abhinay Shankar, et al.
Publicado: (2025)
por: Belde, Abhinay Shankar, et al.
Publicado: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
por: Ashuach, Tomer, et al.
Publicado: (2025)
por: Ashuach, Tomer, et al.
Publicado: (2025)
Defending against Backdoor Attacks via Module Switching
por: Li, Weijun, et al.
Publicado: (2025)
por: Li, Weijun, et al.
Publicado: (2025)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
por: Chehbouni, Khaoula, et al.
Publicado: (2025)
por: Chehbouni, Khaoula, et al.
Publicado: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
por: Smădu, Răzvan-Alexandru, et al.
Publicado: (2025)
por: Smădu, Răzvan-Alexandru, et al.
Publicado: (2025)
UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
por: Geng, Runpeng, et al.
Publicado: (2025)
por: Geng, Runpeng, et al.
Publicado: (2025)
Test-Time Scaling of Reasoning Models for Machine Translation
por: Li, Zihao, et al.
Publicado: (2025)
por: Li, Zihao, et al.
Publicado: (2025)
Identifying Fairness Issues in Automatically Generated Testing Content
por: Stowe, Kevin, et al.
Publicado: (2024)
por: Stowe, Kevin, et al.
Publicado: (2024)
A Likelihood Ratio Test of Genetic Relationship among Languages
por: Akavarapu, V. S. D. S. Mahesh, et al.
Publicado: (2024)
por: Akavarapu, V. S. D. S. Mahesh, et al.
Publicado: (2024)
Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset
por: Galvan-Sosa, Diana, et al.
Publicado: (2025)
por: Galvan-Sosa, Diana, et al.
Publicado: (2025)
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
Distinguishing Ignorance from Error in LLM Hallucinations
por: Simhi, Adi, et al.
Publicado: (2024)
por: Simhi, Adi, et al.
Publicado: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
por: Simhi, Adi, et al.
Publicado: (2024)
por: Simhi, Adi, et al.
Publicado: (2024)
Lightweight Connective Detection Using Gradient Boosting
por: Er, Mustafa Erolcan, et al.
Publicado: (2024)
por: Er, Mustafa Erolcan, et al.
Publicado: (2024)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
por: Ge, Danying, et al.
Publicado: (2025)
por: Ge, Danying, et al.
Publicado: (2025)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
por: Toker, Michael, et al.
Publicado: (2024)
por: Toker, Michael, et al.
Publicado: (2024)
A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
por: Gaim, Fitsum, et al.
Publicado: (2025)
por: Gaim, Fitsum, et al.
Publicado: (2025)
Towards Red Teaming in Multimodal and Multilingual Translation
por: Ropers, Christophe, et al.
Publicado: (2024)
por: Ropers, Christophe, et al.
Publicado: (2024)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
por: Simhi, Adi, et al.
Publicado: (2025)
por: Simhi, Adi, et al.
Publicado: (2025)
Extracting Structured Insights from Financial News: An Augmented LLM Driven Approach
por: Dolphin, Rian, et al.
Publicado: (2024)
por: Dolphin, Rian, et al.
Publicado: (2024)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
por: Collado-Montañez, Jaime, et al.
Publicado: (2025)
por: Collado-Montañez, Jaime, et al.
Publicado: (2025)
Boosting Accuracy and Interpretability in Multilingual Hate Speech Detection Through Layer Freezing and Explainable AI
por: Bilehsavar, Meysam Shirdel, et al.
Publicado: (2026)
por: Bilehsavar, Meysam Shirdel, et al.
Publicado: (2026)
A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection
por: Kim, Songsoo, et al.
Publicado: (2025)
por: Kim, Songsoo, et al.
Publicado: (2025)
"AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
por: Rathva, Harsh, et al.
Publicado: (2025)
por: Rathva, Harsh, et al.
Publicado: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
por: Kugler, Kai
Publicado: (2025)
por: Kugler, Kai
Publicado: (2025)
TrustAI at SemEval-2024 Task 8: A Comprehensive Analysis of Multi-domain Machine Generated Text Detection Techniques
por: Urlana, Ashok, et al.
Publicado: (2024)
por: Urlana, Ashok, et al.
Publicado: (2024)
Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
por: Young, Richard
Publicado: (2025)
por: Young, Richard
Publicado: (2025)
HACK: Hallucinations Along Certainty and Knowledge Axes
por: Simhi, Adi, et al.
Publicado: (2025)
por: Simhi, Adi, et al.
Publicado: (2025)
Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
por: Pai, Aaditya
Publicado: (2026)
por: Pai, Aaditya
Publicado: (2026)
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
por: Zhang, Luyan, et al.
Publicado: (2025)
por: Zhang, Luyan, et al.
Publicado: (2025)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
por: Tu, Songjun, et al.
Publicado: (2026)
por: Tu, Songjun, et al.
Publicado: (2026)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
por: Liu, Zhongxin, et al.
Publicado: (2025)
por: Liu, Zhongxin, et al.
Publicado: (2025)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
por: Zhang, Bowen, et al.
Publicado: (2025)
por: Zhang, Bowen, et al.
Publicado: (2025)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
por: Cherif, Ahmed
Publicado: (2026)
por: Cherif, Ahmed
Publicado: (2026)
Graphemic Normalization of the Perso-Arabic Script
por: Doctor, Raiomond, et al.
Publicado: (2022)
por: Doctor, Raiomond, et al.
Publicado: (2022)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
por: Gutkin, Alexander, et al.
Publicado: (2023)
por: Gutkin, Alexander, et al.
Publicado: (2023)
KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
por: Nzeyimana, Antoine, et al.
Publicado: (2025)
por: Nzeyimana, Antoine, et al.
Publicado: (2025)
Ejemplares similares
-
VertAttack: Taking advantage of Text Classifiers' horizontal vision
por: Rusert, Jonathan
Publicado: (2024) -
Overcoming Black-box Attack Inefficiency with Hybrid and Dynamic Select Algorithms
por: Belde, Abhinay Shankar, et al.
Publicado: (2025) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
por: Ashuach, Tomer, et al.
Publicado: (2025) -
Defending against Backdoor Attacks via Module Switching
por: Li, Weijun, et al.
Publicado: (2025) -
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
por: Chehbouni, Khaoula, et al.
Publicado: (2025)