ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
Fuente:
arXiv
Saved in:
| Main Authors: | Aswal, Darpan, Hudelot, Céline |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neuro-Symbolic Frameworks: Conceptual Characterization and Empirical Comparative Analysis
by: Sinha, Sania, et al.
Published: (2025)
by: Sinha, Sania, et al.
Published: (2025)
Beyond Accuracy: Introducing a Symbolic-Mechanistic Approach to Interpretable Evaluation
by: Habibi, Reza, et al.
Published: (2026)
by: Habibi, Reza, et al.
Published: (2026)
ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark
by: Shalyt, Michael, et al.
Published: (2025)
by: Shalyt, Michael, et al.
Published: (2025)
Improving Neural-based Classification with Logical Background Knowledge
by: Ledaguenel, Arthur, et al.
Published: (2024)
by: Ledaguenel, Arthur, et al.
Published: (2024)
R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning
by: Chen, Yongchao, et al.
Published: (2025)
by: Chen, Yongchao, et al.
Published: (2025)
ONSEP: A Novel Online Neural-Symbolic Framework for Event Prediction Based on Large Language Model
by: Yu, Xuanqing, et al.
Published: (2024)
by: Yu, Xuanqing, et al.
Published: (2024)
A Complexity Map of Probabilistic Reasoning for Neurosymbolic Classification Techniques
by: Ledaguenel, Arthur, et al.
Published: (2024)
by: Ledaguenel, Arthur, et al.
Published: (2024)
Symbolic Regression with a Learned Concept Library
by: Grayeli, Arya, et al.
Published: (2024)
by: Grayeli, Arya, et al.
Published: (2024)
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
by: Aswal, Darpan, et al.
Published: (2025)
by: Aswal, Darpan, et al.
Published: (2025)
A Neuro-Symbolic Approach to Monitoring Salt Content in Food
by: Tayal, Anuja, et al.
Published: (2024)
by: Tayal, Anuja, et al.
Published: (2024)
Neural Concept Binder
by: Stammer, Wolfgang, et al.
Published: (2024)
by: Stammer, Wolfgang, et al.
Published: (2024)
Unlocking the Potential of Generative AI through Neuro-Symbolic Architectures: Benefits and Limitations
by: Bougzime, Oualid, et al.
Published: (2025)
by: Bougzime, Oualid, et al.
Published: (2025)
Ontology-Guided Neuro-Symbolic Inference: Grounding Language Models with Mathematical Domain Knowledge
by: Labre, Marcelo
Published: (2026)
by: Labre, Marcelo
Published: (2026)
Towards Learning to Reason: Comparing LLMs with Neuro-Symbolic on Arithmetic Relations in Abstract Reasoning
by: Hersche, Michael, et al.
Published: (2024)
by: Hersche, Michael, et al.
Published: (2024)
Hierarchical Neuro-Symbolic Decision Transformer
by: Baheri, Ali, et al.
Published: (2025)
by: Baheri, Ali, et al.
Published: (2025)
A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning
by: Chen, Michael K.
Published: (2025)
by: Chen, Michael K.
Published: (2025)
NePTune: A Neuro-Pythonic Framework for Tunable Compositional Reasoning on Vision-Language
by: Kamali, Danial, et al.
Published: (2025)
by: Kamali, Danial, et al.
Published: (2025)
CodeSteer: Symbolic-Augmented Language Models via Code/Text Guidance
by: Chen, Yongchao, et al.
Published: (2025)
by: Chen, Yongchao, et al.
Published: (2025)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
by: Gao, Silin, et al.
Published: (2025)
by: Gao, Silin, et al.
Published: (2025)
GOFAI meets Generative AI: Development of Expert Systems by means of Large Language Models
by: Garrido-Merchán, Eduardo C., et al.
Published: (2025)
by: Garrido-Merchán, Eduardo C., et al.
Published: (2025)
Speaking in Words, Thinking in Logic: A Dual-Process Framework in QA Systems
by: Bui, Tuan, et al.
Published: (2025)
by: Bui, Tuan, et al.
Published: (2025)
Towards Automated Functional Equation Proving: A Benchmark Dataset and A Domain-Specific In-Context Agent
by: Buali, Mahdi, et al.
Published: (2024)
by: Buali, Mahdi, et al.
Published: (2024)
Large Language Models as Mirrors of Societal Moral Standards
by: Papadopoulou, Evi, et al.
Published: (2024)
by: Papadopoulou, Evi, et al.
Published: (2024)
Quantum Knowledge Graph: Modeling Context-Dependent Triplet Validity
by: Wang, Yao, et al.
Published: (2026)
by: Wang, Yao, et al.
Published: (2026)
LLMs as mirrors of societal moral standards: reflection of cultural divergence and agreement across ethical topics
by: Meijer, Mijntje, et al.
Published: (2024)
by: Meijer, Mijntje, et al.
Published: (2024)
Cognitive LLMs: Towards Integrating Cognitive Architectures and Large Language Models for Manufacturing Decision-making
by: Wu, Siyu, et al.
Published: (2024)
by: Wu, Siyu, et al.
Published: (2024)
Synthesizing Evolving Symbolic Representations for Autonomous Systems
by: Sartor, Gabriele, et al.
Published: (2024)
by: Sartor, Gabriele, et al.
Published: (2024)
LogicMP: A Neuro-symbolic Approach for Encoding First-order Logic Constraints
by: Xu, Weidi, et al.
Published: (2023)
by: Xu, Weidi, et al.
Published: (2023)
Neuro-Symbolic ODE Discovery with Latent Grammar Flow
by: Yu, Karin, et al.
Published: (2026)
by: Yu, Karin, et al.
Published: (2026)
ANSR-DT: An Adaptive Neuro-Symbolic Learning and Reasoning Framework for Digital Twins
by: Hakim, Safayat Bin, et al.
Published: (2025)
by: Hakim, Safayat Bin, et al.
Published: (2025)
EGG-SR: Embedding Symbolic Equivalence into Symbolic Regression via Equality Graph
by: Jiang, Nan, et al.
Published: (2025)
by: Jiang, Nan, et al.
Published: (2025)
Learning Interpretable Network Dynamics via Universal Neural Symbolic Regression
by: Hu, Jiao, et al.
Published: (2024)
by: Hu, Jiao, et al.
Published: (2024)
Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning
by: Hao, Yilun, et al.
Published: (2025)
by: Hao, Yilun, et al.
Published: (2025)
Hierarchical NeuroSymbolic Approach for Comprehensive and Explainable Action Quality Assessment
by: Okamoto, Lauren, et al.
Published: (2024)
by: Okamoto, Lauren, et al.
Published: (2024)
NeSy-Edge: Neuro-Symbolic Trustworthy Self-Healing in the Computing Continuum
by: Ye, Peihan, et al.
Published: (2026)
by: Ye, Peihan, et al.
Published: (2026)
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
by: Duan, Ranjie, et al.
Published: (2025)
by: Duan, Ranjie, et al.
Published: (2025)
Incorporating Structure and Chord Constraints in Symbolic Transformer-based Melodic Harmonization
by: Kaliakatsos-Papakostas, Maximos, et al.
Published: (2025)
by: Kaliakatsos-Papakostas, Maximos, et al.
Published: (2025)
Chronosymbolic Learning: Efficient CHC Solving with Symbolic Reasoning and Inductive Learning
by: Luo, Ziyan, et al.
Published: (2023)
by: Luo, Ziyan, et al.
Published: (2023)
Auditable Unit-Aware Thresholds in Symbolic Regression via Logistic-Gated Operators
by: Deng, Ou, et al.
Published: (2025)
by: Deng, Ou, et al.
Published: (2025)
Unsupervised Symbolic Anomaly Detection
by: Hossain, Md Maruf, et al.
Published: (2026)
by: Hossain, Md Maruf, et al.
Published: (2026)
Similar Items
-
Neuro-Symbolic Frameworks: Conceptual Characterization and Empirical Comparative Analysis
by: Sinha, Sania, et al.
Published: (2025) -
Beyond Accuracy: Introducing a Symbolic-Mechanistic Approach to Interpretable Evaluation
by: Habibi, Reza, et al.
Published: (2026) -
ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark
by: Shalyt, Michael, et al.
Published: (2025) -
Improving Neural-based Classification with Logical Background Knowledge
by: Ledaguenel, Arthur, et al.
Published: (2024) -
R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning
by: Chen, Yongchao, et al.
Published: (2025)