BreakFun: Jailbreaking LLMs via Schema Exploitation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Oskooei, Amirkia Rafiei, Aktas, Mehmet S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
von: Zanbaghi, Shahin, et al.
Veröffentlicht: (2025)
von: Zanbaghi, Shahin, et al.
Veröffentlicht: (2025)
Harnessing non-adversarial robustness in large language models
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
von: Yang, Yibo
Veröffentlicht: (2025)
von: Yang, Yibo
Veröffentlicht: (2025)
Feature Selection Based on Reinforcement Learning and Hazard State Classification for Magnetic Adhesion Wall-Climbing Robots
von: Ma, Zhen, et al.
Veröffentlicht: (2025)
von: Ma, Zhen, et al.
Veröffentlicht: (2025)
Exploiting Web Search Tools of AI Agents for Data Exfiltration
von: Rall, Dennis, et al.
Veröffentlicht: (2025)
von: Rall, Dennis, et al.
Veröffentlicht: (2025)
Semantically Guided Adversarial Testing of Vision Models Using Language Models
von: Filus, Katarzyna, et al.
Veröffentlicht: (2025)
von: Filus, Katarzyna, et al.
Veröffentlicht: (2025)
"Abuse Risks are Often Inherent to Product Features": Exploring AI Vendors' Bug Bounty and Responsible Disclosure Policies
von: Piao, Yangheran, et al.
Veröffentlicht: (2025)
von: Piao, Yangheran, et al.
Veröffentlicht: (2025)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
von: Keeman, Michael
Veröffentlicht: (2026)
von: Keeman, Michael
Veröffentlicht: (2026)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
von: Henry, James
Veröffentlicht: (2026)
von: Henry, James
Veröffentlicht: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
Explainable Attention-Based LSTM Framework for Early Detection of AI-Assisted Ransomware via File System Behavioral Analysis
von: Nayak, Prabhudarshi, et al.
Veröffentlicht: (2026)
von: Nayak, Prabhudarshi, et al.
Veröffentlicht: (2026)
Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
Inference acceleration for large language models using "stairs" assisted greedy generation
von: Grigaliūnas, Domas, et al.
Veröffentlicht: (2024)
von: Grigaliūnas, Domas, et al.
Veröffentlicht: (2024)
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
von: Imanov, Olaf Yunus Laitinen, et al.
Veröffentlicht: (2026)
von: Imanov, Olaf Yunus Laitinen, et al.
Veröffentlicht: (2026)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
ProactBench: Beyond What The User Asked For
von: Harfi, Sepehr, et al.
Veröffentlicht: (2026)
von: Harfi, Sepehr, et al.
Veröffentlicht: (2026)
How much do LLMs learn from negative examples?
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
von: Fan, Jingxing, et al.
Veröffentlicht: (2025)
von: Fan, Jingxing, et al.
Veröffentlicht: (2025)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
von: Lee, Wooin, et al.
Veröffentlicht: (2026)
von: Lee, Wooin, et al.
Veröffentlicht: (2026)
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
von: Merves, Tyler H., et al.
Veröffentlicht: (2026)
von: Merves, Tyler H., et al.
Veröffentlicht: (2026)
RAR: Setting Knowledge Tripwires for Retrieval Augmented Rejection
von: Buonocore, Tommaso Mario, et al.
Veröffentlicht: (2025)
von: Buonocore, Tommaso Mario, et al.
Veröffentlicht: (2025)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
Improved ICNN-LSTM Model Classification Based on Attitude Sensor Data for Hazardous State Assessment of Magnetic Adhesion Climbing Wall Robots
von: Ma, Zhen, et al.
Veröffentlicht: (2024)
von: Ma, Zhen, et al.
Veröffentlicht: (2024)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
Efficient Fine-Tuning Methods for Portuguese Question Answering: A Comparative Study of PEFT on BERTimbau and Exploratory Evaluation of Generative LLMs
von: Nina, Mariela M., et al.
Veröffentlicht: (2026)
von: Nina, Mariela M., et al.
Veröffentlicht: (2026)
Representing LLMs in Prompt Semantic Task Space
von: Kashani, Idan, et al.
Veröffentlicht: (2025)
von: Kashani, Idan, et al.
Veröffentlicht: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
von: Dang, Kieu, et al.
Veröffentlicht: (2025) -
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
von: Hill, Brennen, et al.
Veröffentlicht: (2025) -
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025) -
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
von: Basu, Abhinaba
Veröffentlicht: (2026) -
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)