Safer Policy Compliance with Dynamic Epistemic Fallback
Fuente:
arXiv
Saved in:
| Main Authors: | Imperial, Joseph Marvin, Madabushi, Harish Tayyar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
by: Imperial, Joseph Marvin, et al.
Published: (2025)
by: Imperial, Joseph Marvin, et al.
Published: (2025)
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance
by: Imperial, Joseph Marvin, et al.
Published: (2025)
by: Imperial, Joseph Marvin, et al.
Published: (2025)
SpeciaLex: A Benchmark for In-Context Specialized Lexicon Learning
by: Imperial, Joseph Marvin, et al.
Published: (2024)
by: Imperial, Joseph Marvin, et al.
Published: (2024)
Standardize: Aligning Language Models with Expert-Defined Standards for Content Generation
by: Imperial, Joseph Marvin, et al.
Published: (2024)
by: Imperial, Joseph Marvin, et al.
Published: (2024)
FS-RAG: A Frame Semantics Based Approach for Improved Factual Accuracy in Large Language Models
by: Madabushi, Harish Tayyar
Published: (2024)
by: Madabushi, Harish Tayyar
Published: (2024)
The Inherent Limits of Pretrained LLMs: The Unexpected Convergence of Instruction Tuning and In-Context Learning Capabilities
by: Bigoulaeva, Irina, et al.
Published: (2025)
by: Bigoulaeva, Irina, et al.
Published: (2025)
Pre-Trained Language Models Represent Some Geographic Populations Better Than Others
by: Dunn, Jonathan, et al.
Published: (2024)
by: Dunn, Jonathan, et al.
Published: (2024)
Evaluating CxG Generalisation in LLMs via Construction-Based NLI Fine Tuning
by: Mackintosh, Tom, et al.
Published: (2025)
by: Mackintosh, Tom, et al.
Published: (2025)
Dancing with Deer: A Constructional Perspective on MWEs in the Era of LLMs
by: Bonial, Claire, et al.
Published: (2025)
by: Bonial, Claire, et al.
Published: (2025)
Neither Stochastic Parroting nor AGI: LLMs Solve Tasks through Context-Directed Extrapolation from Training Data Priors
by: Madabushi, Harish Tayyar, et al.
Published: (2025)
by: Madabushi, Harish Tayyar, et al.
Published: (2025)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
by: Mendu, Sai Krishna, et al.
Published: (2025)
by: Mendu, Sai Krishna, et al.
Published: (2025)
Code-Mixed Probes Show How Pre-Trained Models Generalise On Code-Switched Text
by: De Leon, Frances A. Laureano, et al.
Published: (2024)
by: De Leon, Frances A. Laureano, et al.
Published: (2024)
Evaluating Large Language Models on Multiword Expressions in Multilingual and Code-Switched Contexts
by: De Leon, Frances Laureano, et al.
Published: (2025)
by: De Leon, Frances Laureano, et al.
Published: (2025)
Adapting Whisper for Regional Dialects: Enhancing Public Services for Vulnerable Populations in the United Kingdom
by: Torgbi, Melissa, et al.
Published: (2025)
by: Torgbi, Melissa, et al.
Published: (2025)
Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMs
by: Puerto, Haritz, et al.
Published: (2024)
by: Puerto, Haritz, et al.
Published: (2024)
Are Emergent Abilities in Large Language Models just In-Context Learning?
by: Lu, Sheng, et al.
Published: (2023)
by: Lu, Sheng, et al.
Published: (2023)
Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning
by: Niu, Jingcheng, et al.
Published: (2025)
by: Niu, Jingcheng, et al.
Published: (2025)
Word Boundary Information Isn't Useful for Encoder Language Models
by: Gow-Smith, Edward, et al.
Published: (2024)
by: Gow-Smith, Edward, et al.
Published: (2024)
Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs
by: Si, Wai Man, et al.
Published: (2026)
by: Si, Wai Man, et al.
Published: (2026)
Frictive Policy Optimization for LLMs: Epistemic Intervention, Risk-Sensitive Control, and Reflective Alignment
by: Pustejovsky, James, et al.
Published: (2026)
by: Pustejovsky, James, et al.
Published: (2026)
Beyond Memorization: Assessing Semantic Generalization in Large Language Models Using Phrasal Constructions
by: Scivetti, Wesley, et al.
Published: (2025)
by: Scivetti, Wesley, et al.
Published: (2025)
Deliberative Alignment: Reasoning Enables Safer Language Models
by: Guan, Melody Y., et al.
Published: (2024)
by: Guan, Melody Y., et al.
Published: (2024)
DynaGuard: A Dynamic Guardian Model With User-Defined Policies
by: Hoover, Monte, et al.
Published: (2025)
by: Hoover, Monte, et al.
Published: (2025)
PurpCode: Reasoning for Safer Code Generation
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
Cross-Domain Data Selection and Augmentation for Automatic Compliance Detection
by: Ikhwantri, Fariz, et al.
Published: (2026)
by: Ikhwantri, Fariz, et al.
Published: (2026)
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
by: Bakman, Yavuz, et al.
Published: (2025)
by: Bakman, Yavuz, et al.
Published: (2025)
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
by: Liu, Mickel, et al.
Published: (2025)
by: Liu, Mickel, et al.
Published: (2025)
Cross-lingual transfer of multilingual models on low resource African Languages
by: Thangaraj, Harish, et al.
Published: (2024)
by: Thangaraj, Harish, et al.
Published: (2024)
LSPO: Length-aware Dynamic Sampling for Policy Optimization in LLM Reasoning
by: Chen, Weizhe, et al.
Published: (2025)
by: Chen, Weizhe, et al.
Published: (2025)
Epistemic Traps: Rational Misalignment Driven by Model Misspecification
by: Xu, Xingcheng, et al.
Published: (2026)
by: Xu, Xingcheng, et al.
Published: (2026)
LLM-Based Robust Product Classification in Commerce and Compliance
by: Gholamian, Sina, et al.
Published: (2024)
by: Gholamian, Sina, et al.
Published: (2024)
BWArea Model: Learning World Model, Inverse Dynamics, and Policy for Controllable Language Generation
by: Jia, Chengxing, et al.
Published: (2024)
by: Jia, Chengxing, et al.
Published: (2024)
Epistemic Observability in Language Models
by: Mason, Tony, et al.
Published: (2026)
by: Mason, Tony, et al.
Published: (2026)
Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning
by: Wu, Jinyang, et al.
Published: (2026)
by: Wu, Jinyang, et al.
Published: (2026)
Universe Routing: Why Self-Evolving Agents Need Epistemic Control
by: Wang, Zhaohui Geoffrey
Published: (2026)
by: Wang, Zhaohui Geoffrey
Published: (2026)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
by: Guo, Kevin H., et al.
Published: (2026)
by: Guo, Kevin H., et al.
Published: (2026)
MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing
by: Yao, Yinsheng, et al.
Published: (2026)
by: Yao, Yinsheng, et al.
Published: (2026)
DCPO: Dynamic Clipping Policy Optimization
by: Yang, Shihui, et al.
Published: (2025)
by: Yang, Shihui, et al.
Published: (2025)
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment
by: Imperial, Joseph Marvin, et al.
Published: (2025)
by: Imperial, Joseph Marvin, et al.
Published: (2025)
Similar Items
-
Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
by: Imperial, Joseph Marvin, et al.
Published: (2025) -
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance
by: Imperial, Joseph Marvin, et al.
Published: (2025) -
SpeciaLex: A Benchmark for In-Context Specialized Lexicon Learning
by: Imperial, Joseph Marvin, et al.
Published: (2024) -
Standardize: Aligning Language Models with Expert-Defined Standards for Content Generation
by: Imperial, Joseph Marvin, et al.
Published: (2024) -
FS-RAG: A Frame Semantics Based Approach for Improved Factual Accuracy in Large Language Models
by: Madabushi, Harish Tayyar
Published: (2024)