Contextual Integrity in LLMs via Reasoning and Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Lan, Guangchen, Inan, Huseyin A., Abdelnabi, Sahar, Kulkarni, Janardhan, Wutschitz, Lukas, Shokri, Reza, Brinton, Christopher G., Sim, Robert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems
by: Lan, Guangchen
Published: (2026)
by: Lan, Guangchen
Published: (2026)
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
by: Lan, Guangchen, et al.
Published: (2026)
by: Lan, Guangchen, et al.
Published: (2026)
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
by: Lan, Guangchen, et al.
Published: (2025)
by: Lan, Guangchen, et al.
Published: (2025)
Calibrated Confidence Estimation for Tabular Question Answering
by: Voss, Lukas
Published: (2026)
by: Voss, Lukas
Published: (2026)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
by: Steele, Brady, et al.
Published: (2026)
by: Steele, Brady, et al.
Published: (2026)
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
by: Shafique, Muhammad Ali, et al.
Published: (2026)
by: Shafique, Muhammad Ali, et al.
Published: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
by: Gwak, Jiho, et al.
Published: (2025)
by: Gwak, Jiho, et al.
Published: (2025)
Planning vs Reasoning: Ablations to Test Capabilities of LoRA layers
by: Redkar, Neel
Published: (2024)
by: Redkar, Neel
Published: (2024)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
by: Easley, Eric, et al.
Published: (2026)
by: Easley, Eric, et al.
Published: (2026)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
by: Resck, Lucas, et al.
Published: (2026)
by: Resck, Lucas, et al.
Published: (2026)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
by: Heyman, Alex, et al.
Published: (2025)
by: Heyman, Alex, et al.
Published: (2025)
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
by: Zhang, Bowen, et al.
Published: (2025)
by: Zhang, Bowen, et al.
Published: (2025)
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
by: Ubukata, Shunsuke
Published: (2026)
by: Ubukata, Shunsuke
Published: (2026)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
by: Bian, Tingcheng, et al.
Published: (2026)
by: Bian, Tingcheng, et al.
Published: (2026)
Automated Bug Triaging using Instruction-Tuned Large Language Models
by: Kiashemshaki, Kiana, et al.
Published: (2025)
by: Kiashemshaki, Kiana, et al.
Published: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
by: Schesch, Benedikt, et al.
Published: (2026)
by: Schesch, Benedikt, et al.
Published: (2026)
Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
by: Eisenstadt, Roy, et al.
Published: (2025)
by: Eisenstadt, Roy, et al.
Published: (2025)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
by: Yang, Shan
Published: (2026)
by: Yang, Shan
Published: (2026)
Survey Transfer Learning: Recycling Data with Silicon Responses
by: Amini, Ali
Published: (2025)
by: Amini, Ali
Published: (2025)
BioAlchemy: Distilling Biological Literature into Reasoning-Ready Reinforcement Learning Training Data
by: Hsu, Brian, et al.
Published: (2026)
by: Hsu, Brian, et al.
Published: (2026)
HIP Network: Historical Information Passing Network for Extrapolation Reasoning on Temporal Knowledge Graph
by: He, Yongquan, et al.
Published: (2024)
by: He, Yongquan, et al.
Published: (2024)
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
by: Yuan, Xin, et al.
Published: (2025)
by: Yuan, Xin, et al.
Published: (2025)
Fuzzy, Symbolic, and Contextual: Enhancing LLM Instruction via Cognitive Scaffolding
by: Figueiredo, Vanessa
Published: (2025)
by: Figueiredo, Vanessa
Published: (2025)
Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
by: Reddy, Sandeep, et al.
Published: (2025)
by: Reddy, Sandeep, et al.
Published: (2025)
LLM-Guided Task- and Affordance-Level Exploration in Reinforcement Learning
by: Luijkx, Jelle, et al.
Published: (2025)
by: Luijkx, Jelle, et al.
Published: (2025)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
by: Asperti, Andrea, et al.
Published: (2025)
by: Asperti, Andrea, et al.
Published: (2025)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
by: Liu, Zhongxin, et al.
Published: (2025)
by: Liu, Zhongxin, et al.
Published: (2025)
On the Influence of Discourse Relations in Persuasive Texts
by: Turk, Nawar, et al.
Published: (2025)
by: Turk, Nawar, et al.
Published: (2025)
Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
by: Liao, Jianxing, et al.
Published: (2025)
by: Liao, Jianxing, et al.
Published: (2025)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
by: Saxena, Udit
Published: (2025)
by: Saxena, Udit
Published: (2025)
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
by: Chen, Xinjie, et al.
Published: (2026)
by: Chen, Xinjie, et al.
Published: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
by: McCann, Jordan F.
Published: (2026)
by: McCann, Jordan F.
Published: (2026)
Super Apriel: One Checkpoint, Many Speeds
by: Labs, SLAM, et al.
Published: (2026)
by: Labs, SLAM, et al.
Published: (2026)
OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind
by: Srishty, Sharmin Sultana, et al.
Published: (2026)
by: Srishty, Sharmin Sultana, et al.
Published: (2026)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
by: Ye, Hua, et al.
Published: (2025)
by: Ye, Hua, et al.
Published: (2025)
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
by: Han, Xudong, et al.
Published: (2025)
by: Han, Xudong, et al.
Published: (2025)
Similar Items
-
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems
by: Lan, Guangchen
Published: (2026) -
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
by: Lan, Guangchen, et al.
Published: (2026) -
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
by: Lan, Guangchen, et al.
Published: (2025) -
Calibrated Confidence Estimation for Tabular Question Answering
by: Voss, Lukas
Published: (2026) -
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
by: Steele, Brady, et al.
Published: (2026)