Red-teaming Activation Probes using Prompted LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Blandfort, Phil, Graham, Robert |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs
por: Blandfort, Phil, et al.
Publicado: (2026)
por: Blandfort, Phil, et al.
Publicado: (2026)
Active Attacks: Red-teaming LLMs via Adaptive Environments
por: Yun, Taeyoung, et al.
Publicado: (2025)
por: Yun, Taeyoung, et al.
Publicado: (2025)
Detecting High-Stakes Interactions with Activation Probes
por: McKenzie, Alex, et al.
Publicado: (2025)
por: McKenzie, Alex, et al.
Publicado: (2025)
Curiosity-driven Red-teaming for Large Language Models
por: Hong, Zhang-Wei, et al.
Publicado: (2024)
por: Hong, Zhang-Wei, et al.
Publicado: (2024)
Detection of adversarial intent in Human-AI teams using LLMs
por: Musaffar, Abed K., et al.
Publicado: (2026)
por: Musaffar, Abed K., et al.
Publicado: (2026)
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models
por: Quaye, Jessica, et al.
Publicado: (2025)
por: Quaye, Jessica, et al.
Publicado: (2025)
Probe-Free Low-Rank Activation Intervention
por: Jiang, Chonghe, et al.
Publicado: (2025)
por: Jiang, Chonghe, et al.
Publicado: (2025)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
por: Nie, Yuzhou, et al.
Publicado: (2024)
por: Nie, Yuzhou, et al.
Publicado: (2024)
Probing Knowledge Holes in Unlearned LLMs
por: Ko, Myeongseob, et al.
Publicado: (2025)
por: Ko, Myeongseob, et al.
Publicado: (2025)
APEX: Probing Neural Networks via Activation Perturbation
por: Ren, Tao, et al.
Publicado: (2026)
por: Ren, Tao, et al.
Publicado: (2026)
ContextBench: Modifying Contexts for Targeted Latent Activation
por: Graham, Robert, et al.
Publicado: (2025)
por: Graham, Robert, et al.
Publicado: (2025)
Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs
por: Ashok, Arjun, et al.
Publicado: (2025)
por: Ashok, Arjun, et al.
Publicado: (2025)
Time-Prompt: Integrated Heterogeneous Prompts for Unlocking LLMs in Time Series Forecasting
por: Wang, Zesen, et al.
Publicado: (2025)
por: Wang, Zesen, et al.
Publicado: (2025)
CAP: Controllable Alignment Prompting for Unlearning in LLMs
por: Wang, Zhaokun, et al.
Publicado: (2026)
por: Wang, Zhaokun, et al.
Publicado: (2026)
Probing Graph Neural Network Activation Patterns Through Graph Topology
por: Tori, Floriano, et al.
Publicado: (2026)
por: Tori, Floriano, et al.
Publicado: (2026)
Fast and Accurate Probing of In-Training LLMs' Downstream Performances
por: Liu, Zhichen, et al.
Publicado: (2026)
por: Liu, Zhichen, et al.
Publicado: (2026)
Activated LoRA: Fine-tuned LLMs for Intrinsics
por: Greenewald, Kristjan, et al.
Publicado: (2025)
por: Greenewald, Kristjan, et al.
Publicado: (2025)
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
por: Liang, Zhiyuan, et al.
Publicado: (2025)
por: Liang, Zhiyuan, et al.
Publicado: (2025)
Evolutionary System Prompt Learning for Reinforcement Learning in LLMs
por: Zhang, Lunjun, et al.
Publicado: (2026)
por: Zhang, Lunjun, et al.
Publicado: (2026)
RAST: Reasoning Activation in LLMs via Small-model Transfer
por: Ouyang, Siru, et al.
Publicado: (2025)
por: Ouyang, Siru, et al.
Publicado: (2025)
Meta-Learning for Cold-Start Personalization in Prompt-Tuned LLMs
por: Zhao, Yushang, et al.
Publicado: (2025)
por: Zhao, Yushang, et al.
Publicado: (2025)
GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
por: Russinovich, Mark, et al.
Publicado: (2026)
por: Russinovich, Mark, et al.
Publicado: (2026)
Enhancing Instruction Following of LLMs via Activation Steering with Dynamic Rejection
por: Kang, Minjae, et al.
Publicado: (2026)
por: Kang, Minjae, et al.
Publicado: (2026)
A Fragile Number Sense: Probing the Elemental Limits of Numerical Reasoning in LLMs
por: Rahman, Roussel, et al.
Publicado: (2025)
por: Rahman, Roussel, et al.
Publicado: (2025)
The case for delegated AI autonomy for Human AI teaming in healthcare
por: Jia, Yan, et al.
Publicado: (2025)
por: Jia, Yan, et al.
Publicado: (2025)
ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
por: Zhang, Zhengyan, et al.
Publicado: (2024)
por: Zhang, Zhengyan, et al.
Publicado: (2024)
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
por: Zhou, Yifan, et al.
Publicado: (2025)
por: Zhou, Yifan, et al.
Publicado: (2025)
GraphiT: Efficient Node Classification on Text-Attributed Graphs with Prompt Optimized LLMs
por: Khoshraftar, Shima, et al.
Publicado: (2025)
por: Khoshraftar, Shima, et al.
Publicado: (2025)
Can LLMs Effectively Leverage Graph Structural Information through Prompts, and Why?
por: Huang, Jin, et al.
Publicado: (2023)
por: Huang, Jin, et al.
Publicado: (2023)
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
por: Zhang, Zeyu, et al.
Publicado: (2026)
por: Zhang, Zeyu, et al.
Publicado: (2026)
How Susceptible are LLMs to Influence in Prompts?
por: Anagnostidis, Sotiris, et al.
Publicado: (2024)
por: Anagnostidis, Sotiris, et al.
Publicado: (2024)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
por: Karvonen, Adam, et al.
Publicado: (2025)
por: Karvonen, Adam, et al.
Publicado: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
por: Heyman, Geert, et al.
Publicado: (2026)
por: Heyman, Geert, et al.
Publicado: (2026)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
por: Le, Qi, et al.
Publicado: (2025)
por: Le, Qi, et al.
Publicado: (2025)
Reinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMs
por: Zhang, Honglin, et al.
Publicado: (2025)
por: Zhang, Honglin, et al.
Publicado: (2025)
CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation
por: Liu, Ziyue, et al.
Publicado: (2025)
por: Liu, Ziyue, et al.
Publicado: (2025)
ZorBA: Zeroth-order Federated Fine-tuning of LLMs with Heterogeneous Block Activation
por: Meng, Chuiyang, et al.
Publicado: (2026)
por: Meng, Chuiyang, et al.
Publicado: (2026)
TrustGLM: Evaluating the Robustness of GraphLLMs Against Prompt, Text, and Structure Attacks
por: Zhang, Qihai, et al.
Publicado: (2025)
por: Zhang, Qihai, et al.
Publicado: (2025)
DGP: A Dual-Granularity Prompting Framework for Fraud Detection with Graph-Enhanced LLMs
por: Li, Yuan, et al.
Publicado: (2025)
por: Li, Yuan, et al.
Publicado: (2025)
Error-Driven Prompt Optimization for Arithmetic Reasoning
por: Pándy, Árpád, et al.
Publicado: (2025)
por: Pándy, Árpád, et al.
Publicado: (2025)
Ejemplares similares
-
Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs
por: Blandfort, Phil, et al.
Publicado: (2026) -
Active Attacks: Red-teaming LLMs via Adaptive Environments
por: Yun, Taeyoung, et al.
Publicado: (2025) -
Detecting High-Stakes Interactions with Activation Probes
por: McKenzie, Alex, et al.
Publicado: (2025) -
Curiosity-driven Red-teaming for Large Language Models
por: Hong, Zhang-Wei, et al.
Publicado: (2024) -
Detection of adversarial intent in Human-AI teams using LLMs
por: Musaffar, Abed K., et al.
Publicado: (2026)