Toward a Dynamic Stackelberg Game-Theoretic Framework for Agentic AI Defense Against LLM Jailbreaking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Zhengye, Zhu, Quanyan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
humancompatible.detect: a Python Toolkit for Detecting Bias in AI Models
von: Matilla, German M., et al.
Veröffentlicht: (2025)
von: Matilla, German M., et al.
Veröffentlicht: (2025)
A Theoretical Framework for Adaptive Utility-Weighted Benchmarking
von: Waggoner, Philip
Veröffentlicht: (2026)
von: Waggoner, Philip
Veröffentlicht: (2026)
Creativity in the Age of AI: Rethinking the Role of Intentional Agency
von: Pearson, James S., et al.
Veröffentlicht: (2026)
von: Pearson, James S., et al.
Veröffentlicht: (2026)
Resilient Federated Chain: Transforming Blockchain Consensus into an Active Defense Layer for Federated Learning
von: García-Márquez, Mario, et al.
Veröffentlicht: (2026)
von: García-Márquez, Mario, et al.
Veröffentlicht: (2026)
Script-Based Dialog Policy Planning for LLM-Powered Conversational Agents: A Basic Architecture for an "AI Therapist"
von: Wasenmüller, Robert, et al.
Veröffentlicht: (2024)
von: Wasenmüller, Robert, et al.
Veröffentlicht: (2024)
AI Consciousness is Inevitable: A Theoretical Computer Science Perspective
von: Blum, Lenore, et al.
Veröffentlicht: (2024)
von: Blum, Lenore, et al.
Veröffentlicht: (2024)
BernGraph: Probabilistic Graph Neural Networks for EHR-based Medication Recommendations
von: Piao, Xihao, et al.
Veröffentlicht: (2024)
von: Piao, Xihao, et al.
Veröffentlicht: (2024)
A Theoretical Analysis of Soft-Label vs Hard-Label Training in Neural Networks
von: Mandal, Saptarshi, et al.
Veröffentlicht: (2024)
von: Mandal, Saptarshi, et al.
Veröffentlicht: (2024)
Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks
von: Paim, Kayua Oleques, et al.
Veröffentlicht: (2025)
von: Paim, Kayua Oleques, et al.
Veröffentlicht: (2025)
How well can a large language model explain business processes as perceived by users?
von: Fahland, Dirk, et al.
Veröffentlicht: (2024)
von: Fahland, Dirk, et al.
Veröffentlicht: (2024)
Mathematical reasoning and the computer
von: Buzzard, Kevin
Veröffentlicht: (2025)
von: Buzzard, Kevin
Veröffentlicht: (2025)
Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading
von: Sadjoli, Nicholas, et al.
Veröffentlicht: (2026)
von: Sadjoli, Nicholas, et al.
Veröffentlicht: (2026)
Modeling Clinical Concern Trajectories in Language Model Agents
von: Subaharan, Sukesh, et al.
Veröffentlicht: (2026)
von: Subaharan, Sukesh, et al.
Veröffentlicht: (2026)
Benchmarking PNW Model for MedMNIST to 100% Accuracy
von: Deng, Bo
Veröffentlicht: (2026)
von: Deng, Bo
Veröffentlicht: (2026)
Challenges and Future Directions in Agentic Reverse Engineering Systems
von: Radey, Salem, et al.
Veröffentlicht: (2026)
von: Radey, Salem, et al.
Veröffentlicht: (2026)
The Hidden Costs of AI: A Review of Energy, E-Waste, and Inequality in Model Development
von: Winsta, Jenis
Veröffentlicht: (2025)
von: Winsta, Jenis
Veröffentlicht: (2025)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
von: Aksoy, Sinan G., et al.
Veröffentlicht: (2026)
von: Aksoy, Sinan G., et al.
Veröffentlicht: (2026)
Towards a Reliable Offline Personal AI Assistant for Long Duration Spaceflight
von: Bensch, Oliver, et al.
Veröffentlicht: (2024)
von: Bensch, Oliver, et al.
Veröffentlicht: (2024)
Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
von: Turan, Berkant, et al.
Veröffentlicht: (2025)
von: Turan, Berkant, et al.
Veröffentlicht: (2025)
Beyond the Org Chart: AI and the Transformation of Invisible Work
von: Rosenthal, Stephanie, et al.
Veröffentlicht: (2026)
von: Rosenthal, Stephanie, et al.
Veröffentlicht: (2026)
LLM-Stackelberg Games: Conjectural Reasoning Equilibria and Their Applications to Spearphishing
von: Zhu, Quanyan
Veröffentlicht: (2025)
von: Zhu, Quanyan
Veröffentlicht: (2025)
AI Governance InternationaL Evaluation Index (AGILE Index) 2024
von: Zeng, Yi, et al.
Veröffentlicht: (2025)
von: Zeng, Yi, et al.
Veröffentlicht: (2025)
Authenticated Delegation and Authorized AI Agents
von: South, Tobin, et al.
Veröffentlicht: (2025)
von: South, Tobin, et al.
Veröffentlicht: (2025)
seqme: a Python library for evaluating biological sequence design
von: Møller-Larsen, Rasmus, et al.
Veröffentlicht: (2025)
von: Møller-Larsen, Rasmus, et al.
Veröffentlicht: (2025)
Adaptive Orchestration for Large-Scale Inference on Heterogeneous Accelerator Systems Balancing Cost, Performance, and Resilience
von: Biran, Yahav, et al.
Veröffentlicht: (2025)
von: Biran, Yahav, et al.
Veröffentlicht: (2025)
Tape: A Cellular Automata Benchmark for Evaluating Rule-Shift Generalization in Reinforcement Learning
von: Pan, Enze
Veröffentlicht: (2026)
von: Pan, Enze
Veröffentlicht: (2026)
RASP-Tuner: Retrieval-Augmented Soft Prompts for Context-Aware Black-Box Optimization in Non-Stationary Environments
von: Pan, Enze
Veröffentlicht: (2026)
von: Pan, Enze
Veröffentlicht: (2026)
Leveraging Diversity in Online Interactions
von: Osman, Nardine, et al.
Veröffentlicht: (2023)
von: Osman, Nardine, et al.
Veröffentlicht: (2023)
ConSensus: Multi-Agent Collaboration for Multimodal Sensing
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2026)
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2026)
Auditable Homomorphic-based Decentralized Collaborative AI with Attribute-based Differential Privacy
von: Yeh, Lo-Yao, et al.
Veröffentlicht: (2024)
von: Yeh, Lo-Yao, et al.
Veröffentlicht: (2024)
RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being
von: Ferdousi, Rahatara, et al.
Veröffentlicht: (2025)
von: Ferdousi, Rahatara, et al.
Veröffentlicht: (2025)
Towards Human Cognition Level-based Experiment Design for Counterfactual Explanations (XAI)
von: Suffian, Muhammad, et al.
Veröffentlicht: (2022)
von: Suffian, Muhammad, et al.
Veröffentlicht: (2022)
Modelling Human Values for AI Reasoning
von: Osman, Nardine, et al.
Veröffentlicht: (2024)
von: Osman, Nardine, et al.
Veröffentlicht: (2024)
Fast, close, non-singular and property-preserving approximations of entropic measures
von: Horenko, Illia, et al.
Veröffentlicht: (2025)
von: Horenko, Illia, et al.
Veröffentlicht: (2025)
Benchmarking Energy Efficiency of Large Language Models Using vLLM
von: Pronk, K., et al.
Veröffentlicht: (2025)
von: Pronk, K., et al.
Veröffentlicht: (2025)
Specification and Evaluation of Multi-Agent LLM Systems -- Prototype and Cybersecurity Applications
von: Härer, Felix
Veröffentlicht: (2025)
von: Härer, Felix
Veröffentlicht: (2025)
Charting the Future of Scholarly Knowledge with AI: A Community Perspective
von: Jiomekong, Azanzi, et al.
Veröffentlicht: (2025)
von: Jiomekong, Azanzi, et al.
Veröffentlicht: (2025)
AI-FLARES: Artificial Intelligence for the Analysis of Solar Flares Data
von: Piana, Michele, et al.
Veröffentlicht: (2024)
von: Piana, Michele, et al.
Veröffentlicht: (2024)
Feature Relevancy, Necessity and Usefulness: Complexity and Algorithms
von: Capdevielle, Tomás, et al.
Veröffentlicht: (2025)
von: Capdevielle, Tomás, et al.
Veröffentlicht: (2025)
Understanding Knowledge Transferability for Transfer Learning: A Survey
von: Wang, Haohua, et al.
Veröffentlicht: (2025)
von: Wang, Haohua, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
humancompatible.detect: a Python Toolkit for Detecting Bias in AI Models
von: Matilla, German M., et al.
Veröffentlicht: (2025) -
A Theoretical Framework for Adaptive Utility-Weighted Benchmarking
von: Waggoner, Philip
Veröffentlicht: (2026) -
Creativity in the Age of AI: Rethinking the Role of Intentional Agency
von: Pearson, James S., et al.
Veröffentlicht: (2026) -
Resilient Federated Chain: Transforming Blockchain Consensus into an Active Defense Layer for Federated Learning
von: García-Márquez, Mario, et al.
Veröffentlicht: (2026) -
Script-Based Dialog Policy Planning for LLM-Powered Conversational Agents: A Basic Architecture for an "AI Therapist"
von: Wasenmüller, Robert, et al.
Veröffentlicht: (2024)