How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Feldman, Shai, Romano, Yaniv |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Robust Conformal Prediction Using Privileged Information
por: Feldman, Shai, et al.
Publicado: (2024)
por: Feldman, Shai, et al.
Publicado: (2024)
Conformal Prediction with Corrupted Labels: Uncertain Imputation and Robust Re-weighting
por: Feldman, Shai, et al.
Publicado: (2025)
por: Feldman, Shai, et al.
Publicado: (2025)
Calibrated Predictive Lower Bounds on Time-to-Unsafe-Sampling in LLMs
por: Davidov, Hen, et al.
Publicado: (2025)
por: Davidov, Hen, et al.
Publicado: (2025)
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
por: Tolochinsky, Elad, et al.
Publicado: (2026)
por: Tolochinsky, Elad, et al.
Publicado: (2026)
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
por: Kumarappan, Adarsh, et al.
Publicado: (2025)
por: Kumarappan, Adarsh, et al.
Publicado: (2025)
Label Noise Robustness of Conformal Prediction
por: Einbinder, Bat-Sheva, et al.
Publicado: (2022)
por: Einbinder, Bat-Sheva, et al.
Publicado: (2022)
Semi-Supervised Hypothesis Testing by Betting on Predictions
por: Tenzer, Yaniv, et al.
Publicado: (2026)
por: Tenzer, Yaniv, et al.
Publicado: (2026)
Multi-Turn Jailbreaks Are Simpler Than They Seem
por: Yang, Xiaoxue, et al.
Publicado: (2025)
por: Yang, Xiaoxue, et al.
Publicado: (2025)
Adaptive Budget Allocation in LLM-Augmented Surveys
por: Ye, Zikun, et al.
Publicado: (2026)
por: Ye, Zikun, et al.
Publicado: (2026)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
por: Li, Nathaniel, et al.
Publicado: (2024)
por: Li, Nathaniel, et al.
Publicado: (2024)
Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking
por: He, Zhida, et al.
Publicado: (2026)
por: He, Zhida, et al.
Publicado: (2026)
Attention-Aware GNN-based Input Defense against Multi-Turn LLM Jailbreak
por: Huang, Zixuan, et al.
Publicado: (2025)
por: Huang, Zixuan, et al.
Publicado: (2025)
Multi-Task Combinatorial Bandits for Budget Allocation
por: Ge, Lin, et al.
Publicado: (2024)
por: Ge, Lin, et al.
Publicado: (2024)
ZEBRA: Zero-shot Budgeted Resource Allocation for LLM Orchestration
por: Hamri, May, et al.
Publicado: (2026)
por: Hamri, May, et al.
Publicado: (2026)
Online Learning with Improving Agents: Multiclass, Budgeted Agents and Bandit Learners
por: Ashkezari, Sajad, et al.
Publicado: (2026)
por: Ashkezari, Sajad, et al.
Publicado: (2026)
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
por: Taghibakhshi, Ali, et al.
Publicado: (2026)
por: Taghibakhshi, Ali, et al.
Publicado: (2026)
Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models
por: Li, Songze, et al.
Publicado: (2026)
por: Li, Songze, et al.
Publicado: (2026)
Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning
por: Jali, Neharika, et al.
Publicado: (2026)
por: Jali, Neharika, et al.
Publicado: (2026)
Uncertainty Quantification and Data Efficiency in AI: An Information-Theoretic Perspective
por: Simeone, Osvaldo, et al.
Publicado: (2025)
por: Simeone, Osvaldo, et al.
Publicado: (2025)
Protected Test-Time Adaptation via Online Entropy Matching: A Betting Approach
por: Bar, Yarin, et al.
Publicado: (2024)
por: Bar, Yarin, et al.
Publicado: (2024)
Efficient Budget Allocation for Large-Scale LLM-Enabled Virtual Screening
por: Li, Zaile, et al.
Publicado: (2024)
por: Li, Zaile, et al.
Publicado: (2024)
Building Math Agents with Multi-Turn Iterative Preference Learning
por: Xiong, Wei, et al.
Publicado: (2024)
por: Xiong, Wei, et al.
Publicado: (2024)
Robust Conformal Outlier Detection under Contaminated Reference Data
por: Bashari, Meshi, et al.
Publicado: (2025)
por: Bashari, Meshi, et al.
Publicado: (2025)
Mitigating Many-Shot Jailbreaking
por: Ackerman, Christopher M., et al.
Publicado: (2025)
por: Ackerman, Christopher M., et al.
Publicado: (2025)
Semi-Supervised Risk Control via Prediction-Powered Inference
por: Einbinder, Bat-Sheva, et al.
Publicado: (2024)
por: Einbinder, Bat-Sheva, et al.
Publicado: (2024)
Jailbreak Attack Initializations as Extractors of Compliance Directions
por: Levi, Amit, et al.
Publicado: (2025)
por: Levi, Amit, et al.
Publicado: (2025)
Pivotal Auto-Encoder via Self-Normalizing ReLU
por: Goldenstein, Nelson, et al.
Publicado: (2024)
por: Goldenstein, Nelson, et al.
Publicado: (2024)
Learning a Continue-Thinking Token for Enhanced Test-Time Scaling
por: Ringel, Liran, et al.
Publicado: (2025)
por: Ringel, Liran, et al.
Publicado: (2025)
SINR-Aware Deep Reinforcement Learning for Distributed Dynamic Channel Allocation in Cognitive Interference Networks
por: Cohen, Yaniv, et al.
Publicado: (2024)
por: Cohen, Yaniv, et al.
Publicado: (2024)
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
por: Shen, Yiqun, et al.
Publicado: (2025)
por: Shen, Yiqun, et al.
Publicado: (2025)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
por: Reddy, Aashray, et al.
Publicado: (2025)
por: Reddy, Aashray, et al.
Publicado: (2025)
Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative PPO
por: Jiang, Daniel R., et al.
Publicado: (2025)
por: Jiang, Daniel R., et al.
Publicado: (2025)
Testing For Distribution Shifts with Conditional Conformal Test Martingales
por: Shaer, Shalev, et al.
Publicado: (2026)
por: Shaer, Shalev, et al.
Publicado: (2026)
Jailbreaking Large Language Models in Infinitely Many Ways
por: Goldstein, Oliver, et al.
Publicado: (2025)
por: Goldstein, Oliver, et al.
Publicado: (2025)
Cascaded Transfer: Learning Many Tasks under Budget Constraints
por: Campagne, Eloi, et al.
Publicado: (2026)
por: Campagne, Eloi, et al.
Publicado: (2026)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
por: Rahman, Salman, et al.
Publicado: (2025)
por: Rahman, Salman, et al.
Publicado: (2025)
Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design
por: Wei, Quan, et al.
Publicado: (2025)
por: Wei, Quan, et al.
Publicado: (2025)
TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
por: Xiong, Xiqiao, et al.
Publicado: (2025)
por: Xiong, Xiqiao, et al.
Publicado: (2025)
Hidden Representation Clustering with Multi-Task Representation Learning towards Robust Online Budget Allocation
por: Wang, Xiaohan, et al.
Publicado: (2025)
por: Wang, Xiaohan, et al.
Publicado: (2025)
Synthetic-Powered Multiple Testing with FDR Control
por: Lee, Yonghoon, et al.
Publicado: (2026)
por: Lee, Yonghoon, et al.
Publicado: (2026)
Ejemplares similares
-
Robust Conformal Prediction Using Privileged Information
por: Feldman, Shai, et al.
Publicado: (2024) -
Conformal Prediction with Corrupted Labels: Uncertain Imputation and Robust Re-weighting
por: Feldman, Shai, et al.
Publicado: (2025) -
Calibrated Predictive Lower Bounds on Time-to-Unsafe-Sampling in LLMs
por: Davidov, Hen, et al.
Publicado: (2025) -
Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization
por: Tolochinsky, Elad, et al.
Publicado: (2026) -
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
por: Kumarappan, Adarsh, et al.
Publicado: (2025)