ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | Hardy, Amelia F., Liu, Houjun, Griffith, Allie, Lange, Bernard, Eddy, Duncan, Kochenderfer, Mykel J. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
di: Liu, Houjun, et al.
Pubblicazione: (2026)
di: Liu, Houjun, et al.
Pubblicazione: (2026)
Brahe: A Modern Astrodynamics Library for Research and Engineering Applications
di: Eddy, Duncan, et al.
Pubblicazione: (2026)
di: Eddy, Duncan, et al.
Pubblicazione: (2026)
Optimal Ground Station Selection for Low-Earth Orbiting Satellites
di: Eddy, Duncan, et al.
Pubblicazione: (2024)
di: Eddy, Duncan, et al.
Pubblicazione: (2024)
Inferring Traffic Models in Terminal Airspace from Flight Tracks and Procedures
di: Jung, Soyeon, et al.
Pubblicazione: (2023)
di: Jung, Soyeon, et al.
Pubblicazione: (2023)
Beyond Gradient Averaging in Parallel Optimization: Improved Robustness through Gradient Agreement Filtering
di: Chaubard, Francois, et al.
Pubblicazione: (2024)
di: Chaubard, Francois, et al.
Pubblicazione: (2024)
Fault-Aware MPC for Robotic Fleet Communications Scheduling
di: Schreiber, Carlo, et al.
Pubblicazione: (2026)
di: Schreiber, Carlo, et al.
Pubblicazione: (2026)
LOPR: Latent Occupancy PRediction using Generative Models
di: Lange, Bernard, et al.
Pubblicazione: (2022)
di: Lange, Bernard, et al.
Pubblicazione: (2022)
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
di: Diao, Muxi, et al.
Pubblicazione: (2025)
di: Diao, Muxi, et al.
Pubblicazione: (2025)
Markov Decision Processes for Satellite Maneuver Planning and Collision Avoidance
di: Kuhl, William, et al.
Pubblicazione: (2025)
di: Kuhl, William, et al.
Pubblicazione: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
di: Freenor, Michael, et al.
Pubblicazione: (2025)
di: Freenor, Michael, et al.
Pubblicazione: (2025)
Scene Informer: Anchor-based Occlusion Inference and Trajectory Prediction in Partially Observable Environments
di: Lange, Bernard, et al.
Pubblicazione: (2023)
di: Lange, Bernard, et al.
Pubblicazione: (2023)
Demystifying Prompts in Language Models via Perplexity Estimation
di: Gonen, Hila, et al.
Pubblicazione: (2022)
di: Gonen, Hila, et al.
Pubblicazione: (2022)
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
di: Hardy, Amelia, et al.
Pubblicazione: (2024)
di: Hardy, Amelia, et al.
Pubblicazione: (2024)
Automated Progressive Red Teaming
di: Jiang, Bojian, et al.
Pubblicazione: (2024)
di: Jiang, Bojian, et al.
Pubblicazione: (2024)
Scalable Ground Station Selection for Large LEO Constellations
di: Kim, Grace Ra, et al.
Pubblicazione: (2025)
di: Kim, Grace Ra, et al.
Pubblicazione: (2025)
TroubleLLM: Align to Red Team Expert
di: Xu, Zhuoer, et al.
Pubblicazione: (2024)
di: Xu, Zhuoer, et al.
Pubblicazione: (2024)
One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
di: Fein, Daniel, et al.
Pubblicazione: (2026)
di: Fein, Daniel, et al.
Pubblicazione: (2026)
Anecdoctoring: Automated Red-Teaming Across Language and Place
di: Cuevas, Alejandro, et al.
Pubblicazione: (2025)
di: Cuevas, Alejandro, et al.
Pubblicazione: (2025)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
di: Reuel, Anka, et al.
Pubblicazione: (2024)
di: Reuel, Anka, et al.
Pubblicazione: (2024)
Efficient Response Generation Strategy Selection for Fine-Tuning Large Language Models Through Self-Aligned Perplexity
di: Ren, Xuan, et al.
Pubblicazione: (2025)
di: Ren, Xuan, et al.
Pubblicazione: (2025)
Plan of Thoughts: Heuristic-Guided Problem Solving with Large Language Models
di: Liu, Houjun
Pubblicazione: (2024)
di: Liu, Houjun
Pubblicazione: (2024)
Training a General Purpose Automated Red Teaming Model
di: Padmakumar, Aishwarya, et al.
Pubblicazione: (2026)
di: Padmakumar, Aishwarya, et al.
Pubblicazione: (2026)
PLPP: Prompt Learning with Perplexity Is Self-Distillation for Vision-Language Models
di: Liu, Biao, et al.
Pubblicazione: (2024)
di: Liu, Biao, et al.
Pubblicazione: (2024)
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
Safer-Instruct: Aligning Language Models with Automated Preference Data
di: Shi, Taiwei, et al.
Pubblicazione: (2023)
di: Shi, Taiwei, et al.
Pubblicazione: (2023)
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
di: Han, Vernon Toh Yan, et al.
Pubblicazione: (2024)
di: Han, Vernon Toh Yan, et al.
Pubblicazione: (2024)
Low-Perplexity LLM-Generated Sequences and Where To Find Them
di: Wuhrmann, Arthur, et al.
Pubblicazione: (2025)
di: Wuhrmann, Arthur, et al.
Pubblicazione: (2025)
Red Teaming Multimodal Language Models: Evaluating Harm Across Prompt Modalities and Models
di: Van Doren, Madison, et al.
Pubblicazione: (2025)
di: Van Doren, Madison, et al.
Pubblicazione: (2025)
Self-supervised Multi-future Occupancy Forecasting for Autonomous Driving
di: Lange, Bernard, et al.
Pubblicazione: (2024)
di: Lange, Bernard, et al.
Pubblicazione: (2024)
LeRAAT: LLM-Enabled Real-Time Aviation Advisory Tool
di: Schlichting, Marc R., et al.
Pubblicazione: (2025)
di: Schlichting, Marc R., et al.
Pubblicazione: (2025)
Gradient Co-occurrence Analysis for Detecting Unsafe Prompts in Large Language Models
di: Yang, Jingyuan, et al.
Pubblicazione: (2025)
di: Yang, Jingyuan, et al.
Pubblicazione: (2025)
Hierarchical Framework for Optimizing Wildfire Surveillance and Suppression using Human-Autonomous Teaming
di: Al-Husseini, Mahdi, et al.
Pubblicazione: (2024)
di: Al-Husseini, Mahdi, et al.
Pubblicazione: (2024)
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
di: Hu, Yujia, et al.
Pubblicazione: (2025)
di: Hu, Yujia, et al.
Pubblicazione: (2025)
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
di: Ankner, Zachary, et al.
Pubblicazione: (2024)
di: Ankner, Zachary, et al.
Pubblicazione: (2024)
When Prompt Optimization Becomes Jailbreaking: Adaptive Red-Teaming of Large Language Models
di: Shamsi, Zafir, et al.
Pubblicazione: (2026)
di: Shamsi, Zafir, et al.
Pubblicazione: (2026)
Multi-lingual Multi-turn Automated Red Teaming for LLMs
di: Singhania, Abhishek, et al.
Pubblicazione: (2025)
di: Singhania, Abhishek, et al.
Pubblicazione: (2025)
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
di: Jung, MinJae, et al.
Pubblicazione: (2026)
di: Jung, MinJae, et al.
Pubblicazione: (2026)
Red-Teaming Vision-Language-Action Models via Quality Diversity Prompt Generation for Robust Robot Policies
di: Srikanth, Siddharth, et al.
Pubblicazione: (2026)
di: Srikanth, Siddharth, et al.
Pubblicazione: (2026)
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
di: Wei, Zhang, et al.
Pubblicazione: (2025)
di: Wei, Zhang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
di: Liu, Houjun, et al.
Pubblicazione: (2026) -
Brahe: A Modern Astrodynamics Library for Research and Engineering Applications
di: Eddy, Duncan, et al.
Pubblicazione: (2026) -
Optimal Ground Station Selection for Low-Earth Orbiting Satellites
di: Eddy, Duncan, et al.
Pubblicazione: (2024) -
Inferring Traffic Models in Terminal Airspace from Flight Tracks and Procedures
di: Jung, Soyeon, et al.
Pubblicazione: (2023) -
Beyond Gradient Averaging in Parallel Optimization: Improved Robustness through Gradient Agreement Filtering
di: Chaubard, Francois, et al.
Pubblicazione: (2024)