MAGPIE: A benchmark for Multi-AGent contextual PrIvacy Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Juneja, Gurusha, Pasupulati, Jayanth Naga Sai, Albalak, Alon, Hua, Wenyue, Wang, William Yang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
por: Juneja, Gurusha, et al.
Publicado: (2025)
por: Juneja, Gurusha, et al.
Publicado: (2025)
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
por: Ouyang, Yang, et al.
Publicado: (2025)
por: Ouyang, Yang, et al.
Publicado: (2025)
MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment
por: Dou, Yipu, et al.
Publicado: (2026)
por: Dou, Yipu, et al.
Publicado: (2026)
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
por: Zhou, Kaiwen, et al.
Publicado: (2025)
por: Zhou, Kaiwen, et al.
Publicado: (2025)
Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models
por: Kadali, Sri Durga Sai Sowmya, et al.
Publicado: (2026)
por: Kadali, Sri Durga Sai Sowmya, et al.
Publicado: (2026)
PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models
por: Li, Haoran, et al.
Publicado: (2023)
por: Li, Haoran, et al.
Publicado: (2023)
LeakDojo: Decoding the Leakage Threats of RAG Systems
por: Zhang, Maosen, et al.
Publicado: (2026)
por: Zhang, Maosen, et al.
Publicado: (2026)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
por: Chu, Hua-Rong, et al.
Publicado: (2026)
por: Chu, Hua-Rong, et al.
Publicado: (2026)
ContextLeak: Auditing Leakage in Private In-Context Learning Methods
por: Choi, Jacob, et al.
Publicado: (2025)
por: Choi, Jacob, et al.
Publicado: (2025)
EmojiPrompt: Generative Prompt Obfuscation for Privacy-Preserving Communication with Cloud-based LLMs
por: Lin, Sam, et al.
Publicado: (2024)
por: Lin, Sam, et al.
Publicado: (2024)
MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models
por: Gu, Tianle, et al.
Publicado: (2024)
por: Gu, Tianle, et al.
Publicado: (2024)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
por: Huang, Ruixuan, et al.
Publicado: (2025)
por: Huang, Ruixuan, et al.
Publicado: (2025)
Fully Randomized Pointers
por: Phaye, Sai Dhawal, et al.
Publicado: (2024)
por: Phaye, Sai Dhawal, et al.
Publicado: (2024)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
por: Yuan, Xiaohan, et al.
Publicado: (2024)
por: Yuan, Xiaohan, et al.
Publicado: (2024)
WorldCup Sampling for Multi-bit LLM Watermarking
por: Wang, Yidan, et al.
Publicado: (2026)
por: Wang, Yidan, et al.
Publicado: (2026)
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
por: Yan, Lecheng, et al.
Publicado: (2026)
por: Yan, Lecheng, et al.
Publicado: (2026)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
por: Wei, Zhang, et al.
Publicado: (2025)
por: Wei, Zhang, et al.
Publicado: (2025)
Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
por: Liu, Lijia, et al.
Publicado: (2025)
por: Liu, Lijia, et al.
Publicado: (2025)
Publicly Understandable Electronic Voting: A Non-Cryptographic, End-to-End Verifiable Scheme
por: Gat, Alon
Publicado: (2026)
por: Gat, Alon
Publicado: (2026)
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
por: Baumgärtner, Tim, et al.
Publicado: (2024)
por: Baumgärtner, Tim, et al.
Publicado: (2024)
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
por: Cui, Shiyao, et al.
Publicado: (2023)
por: Cui, Shiyao, et al.
Publicado: (2023)
Enhancing SQL Injection Detection and Prevention Using Generative Models
por: Dasari, Naga Sai, et al.
Publicado: (2025)
por: Dasari, Naga Sai, et al.
Publicado: (2025)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
por: Liu, Songyang, et al.
Publicado: (2026)
por: Liu, Songyang, et al.
Publicado: (2026)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
por: Shen, Guobin, et al.
Publicado: (2025)
por: Shen, Guobin, et al.
Publicado: (2025)
Toward Copyright Integrity and Verifiability via Multi-Bit Watermarking for Intelligent Transportation Systems
por: Wang, Yihao, et al.
Publicado: (2025)
por: Wang, Yihao, et al.
Publicado: (2025)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
por: Dou, Yipu, et al.
Publicado: (2026)
por: Dou, Yipu, et al.
Publicado: (2026)
LockForge: Automating Paper-to-Code for Logic Locking with Multi-Agent Reasoning LLMs
por: Saha, Akashdeep, et al.
Publicado: (2025)
por: Saha, Akashdeep, et al.
Publicado: (2025)
Rethinking Backdoor Detection Evaluation for Language Models
por: Yan, Jun, et al.
Publicado: (2024)
por: Yan, Jun, et al.
Publicado: (2024)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
por: Teja, Lekkala Sai, et al.
Publicado: (2025)
por: Teja, Lekkala Sai, et al.
Publicado: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
por: Freenor, Michael, et al.
Publicado: (2025)
por: Freenor, Michael, et al.
Publicado: (2025)
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
por: Zhou, Yukai, et al.
Publicado: (2025)
por: Zhou, Yukai, et al.
Publicado: (2025)
Data Siphoning Through Advanced Persistent Transmission Attacks At The Physical Layer
por: Hillel-Tuch, Alon
Publicado: (2026)
por: Hillel-Tuch, Alon
Publicado: (2026)
EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
por: Mirbagheri, Mohammad Reza, et al.
Publicado: (2025)
por: Mirbagheri, Mohammad Reza, et al.
Publicado: (2025)
MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors
por: Li, Yuanfan, et al.
Publicado: (2026)
por: Li, Yuanfan, et al.
Publicado: (2026)
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
por: Hastuti, Rochana Prih, et al.
Publicado: (2025)
por: Hastuti, Rochana Prih, et al.
Publicado: (2025)
TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
por: Rajore, Tanmay, et al.
Publicado: (2024)
por: Rajore, Tanmay, et al.
Publicado: (2024)
Multi-use LLM Watermarking and the False Detection Problem
por: Fu, Zihao, et al.
Publicado: (2025)
por: Fu, Zihao, et al.
Publicado: (2025)
HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
por: Ren, Xiaoxue, et al.
Publicado: (2025)
por: Ren, Xiaoxue, et al.
Publicado: (2025)
Multi-Agent Collaboration in Incident Response with Large Language Models
por: Liu, Zefang
Publicado: (2024)
por: Liu, Zefang
Publicado: (2024)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
por: Wang, Junlin, et al.
Publicado: (2024)
por: Wang, Junlin, et al.
Publicado: (2024)
Ejemplares similares
-
MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
por: Juneja, Gurusha, et al.
Publicado: (2025) -
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
por: Ouyang, Yang, et al.
Publicado: (2025) -
MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment
por: Dou, Yipu, et al.
Publicado: (2026) -
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
por: Zhou, Kaiwen, et al.
Publicado: (2025) -
Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models
por: Kadali, Sri Durga Sai Sowmya, et al.
Publicado: (2026)