Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Debenedetti, Edoardo, Rando, Javier, Paleka, Daniel, Florin, Silaghi Fineas, Albastroiu, Dragos, Cohen, Niv, Lemberg, Yuval, Ghosh, Reshmi, Wen, Rui, Salem, Ahmed, Cherubin, Giovanni, Zanella-Beguelin, Santiago, Schmid, Robin, Klemm, Victor, Miki, Takahiro, Li, Chenhao, Kraft, Stefan, Fritz, Mario, Tramèr, Florian, Abdelnabi, Sahar, Schönherr, Lea |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JuliaTrustworthyAI/CounterfactualTraining.jl: SaTML camera-ready
von: Patrick Altmeyer, et al.
Veröffentlicht: (2026)
von: Patrick Altmeyer, et al.
Veröffentlicht: (2026)
The SaTML '24 CNN Interpretability Competition: New Innovations for Concept-Level Interpretability
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
MIDST Challenge at SaTML 2025: Membership Inference over Diffusion-models-based Synthetic Tabular data
von: Shafieinejad, Masoumeh, et al.
Veröffentlicht: (2026)
von: Shafieinejad, Masoumeh, et al.
Veröffentlicht: (2026)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
Evading Black-box Classifiers Without Breaking Eggs
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2023)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2023)
Adversarial Search Engine Optimization for Large Language Models
von: Nestaas, Fredrik, et al.
Veröffentlicht: (2024)
von: Nestaas, Fredrik, et al.
Veröffentlicht: (2024)
Closed-Form Bounds for DP-SGD against Record-level Inference
von: Cherubin, Giovanni, et al.
Veröffentlicht: (2024)
von: Cherubin, Giovanni, et al.
Veröffentlicht: (2024)
Universal Jailbreak Backdoors from Poisoned Human Feedback
von: Rando, Javier, et al.
Veröffentlicht: (2023)
von: Rando, Javier, et al.
Veröffentlicht: (2023)
Pitfalls in Evaluating Language Model Forecasters
von: Paleka, Daniel, et al.
Veröffentlicht: (2025)
von: Paleka, Daniel, et al.
Veröffentlicht: (2025)
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2023)
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2023)
Large-scale online deanonymization with LLMs
von: Lermen, Simon, et al.
Veröffentlicht: (2026)
von: Lermen, Simon, et al.
Veröffentlicht: (2026)
Get my drift? Catching LLM Task Drift with Activation Deltas
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2024)
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2024)
Representations of Text and Images Align From Layer One
von: Wybitul, Evžen, et al.
Veröffentlicht: (2026)
von: Wybitul, Evžen, et al.
Veröffentlicht: (2026)
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
von: Hönig, Robert, et al.
Veröffentlicht: (2024)
von: Hönig, Robert, et al.
Veröffentlicht: (2024)
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
von: Rando, Javier, et al.
Veröffentlicht: (2025)
von: Rando, Javier, et al.
Veröffentlicht: (2025)
Separating internationalization? Episodes from the Cold War History of GAMM
von: Jason Lemberg
Veröffentlicht: (2025)
von: Jason Lemberg
Veröffentlicht: (2025)
Bromination of (AsPh2)2O: The Structure of Tribromo-Diphenylarsenic (V)
von: Luminita Silaghi Dumitrescu
Veröffentlicht: (2000)
von: Luminita Silaghi Dumitrescu
Veröffentlicht: (2000)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2024)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2024)
Beyond Membership: Limitations of Add/Remove Adjacency in Differential Privacy
von: Pradhan, Gauri, et al.
Veröffentlicht: (2025)
von: Pradhan, Gauri, et al.
Veröffentlicht: (2025)
Consistency Checks for Language Model Forecasters
von: Paleka, Daniel, et al.
Veröffentlicht: (2024)
von: Paleka, Daniel, et al.
Veröffentlicht: (2024)
Gradient-based Jailbreak Images for Multimodal Fusion Models
von: Rando, Javier, et al.
Veröffentlicht: (2024)
von: Rando, Javier, et al.
Veröffentlicht: (2024)
TAMAÑO CORPORAL Y TEMPERATURA AMBIENTAL EN POBLACIONES CAZADORAS RECOLECTORAS DEL HOLOCENO TARDIO DE PAMPA Y PATAGONIA
von: Marien Béguelin
Veröffentlicht: (2010)
von: Marien Béguelin
Veröffentlicht: (2010)
ESTIMACION DEL SEXO EN POBLACIONES DEL SUR DE SUDAMERCIA MEDIANTE FUNCIONES DISCRIMINANTES PARA EL FEMUR
von: Marien Béguelin
Veröffentlicht: (2008)
von: Marien Béguelin
Veröffentlicht: (2008)
PONIENDO BLANCO SOBRE NEGRO: ANÁLISIS QUÍMICOS Y MICROSCÓPICOS SOBRE SEDIMENTOS Y RESTOS HUMANOS DE LA LAGUNA DEL JUNCAL (VALLE DEL RÍO NEGRO, NORPATAGONIA)
von: Marien Béguelin
Veröffentlicht: (2022)
von: Marien Béguelin
Veröffentlicht: (2022)
ESTIMACION DE LA ESTATURA EN MUESTRAS DEL HOLOCENO TARDIO DEL N.O. DE SANTA CRUZ: PROBLEMAS METODOLOGICOS
von: Marien Béguelin
Veröffentlicht: (2005)
von: Marien Béguelin
Veröffentlicht: (2005)
Estimación del sexo en cazadores-recolectores de Sudamérica a partir de variables métricas del húmero
von: Marien Béguelin
Veröffentlicht: (2011)
von: Marien Béguelin
Veröffentlicht: (2011)
Variación morfométrica postcraneal en muestras tardías de restos humanos de Patagonia: una aproximación biogeográfica
von: Marien Béguelin
Veröffentlicht: (2006)
von: Marien Béguelin
Veröffentlicht: (2006)
Age determination of sediment core TML1
von: Lamy, Frank, et al.
Veröffentlicht: (2010)
von: Lamy, Frank, et al.
Veröffentlicht: (2010)
Production flexibility and trade credit under revenue uncertainty
von: Nicos Koussis, et al.
Veröffentlicht: (2024)
von: Nicos Koussis, et al.
Veröffentlicht: (2024)
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
Tropical Geometric Tools for Machine Learning: the TML package
von: Barnhill, David, et al.
Veröffentlicht: (2023)
von: Barnhill, David, et al.
Veröffentlicht: (2023)
Modeling the Feedback of AI Price Estimations on Actual Market Values
von: Silaghi, Viorel, et al.
Veröffentlicht: (2024)
von: Silaghi, Viorel, et al.
Veröffentlicht: (2024)
Vereinigungsbedingte Dimensionen regionaler Arbeitsmobilitaet
von: Schönherr, Annette
Veröffentlicht: (2019)
von: Schönherr, Annette
Veröffentlicht: (2019)
Online Hybrid-Belief POMDP with Coupled Semantic-Geometric Models
von: Lemberg, Tuvy, et al.
Veröffentlicht: (2025)
von: Lemberg, Tuvy, et al.
Veröffentlicht: (2025)
An Adversarial Perspective on Machine Unlearning for AI Safety
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
No More, No Less: Task Alignment in Terminal Agents
von: Mavali, Sina, et al.
Veröffentlicht: (2026)
von: Mavali, Sina, et al.
Veröffentlicht: (2026)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
Highlight & Summarize: RAG without the jailbreaks
von: Cherubin, Giovanni, et al.
Veröffentlicht: (2025)
von: Cherubin, Giovanni, et al.
Veröffentlicht: (2025)
Poisoning Web-Scale Training Datasets is Practical
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
JuliaTrustworthyAI/CounterfactualTraining.jl: SaTML camera-ready
von: Patrick Altmeyer, et al.
Veröffentlicht: (2026) -
The SaTML '24 CNN Interpretability Competition: New Innovations for Concept-Level Interpretability
von: Casper, Stephen, et al.
Veröffentlicht: (2024) -
MIDST Challenge at SaTML 2025: Membership Inference over Diffusion-models-based Synthetic Tabular data
von: Shafieinejad, Masoumeh, et al.
Veröffentlicht: (2026) -
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025) -
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
von: Aerni, Michael, et al.
Veröffentlicht: (2024)