Reflect: Transparent Principle-Guided Reasoning for Constitutional Alignment at Scale
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bell, Henry, Zhang, Caroline, Haque, Mohammed Mobasserul, Potdar, Dhaval, Zaman, Samia, Fain, Brandon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Preferences: Learning Alignment Principles Grounded in Human Reasons and Values
von: Bell, Henry, et al.
Veröffentlicht: (2026)
von: Bell, Henry, et al.
Veröffentlicht: (2026)
Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
von: Huang, Jiameng, et al.
Veröffentlicht: (2025)
von: Huang, Jiameng, et al.
Veröffentlicht: (2025)
Transfer Q Star: Principled Decoding for LLM Alignment
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
Test-Time Scaling with Reflective Generative Model
von: Wang, Zixiao, et al.
Veröffentlicht: (2025)
von: Wang, Zixiao, et al.
Veröffentlicht: (2025)
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
von: Zhang, Yunfan, et al.
Veröffentlicht: (2025)
von: Zhang, Yunfan, et al.
Veröffentlicht: (2025)
REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning
von: Deng, Hexuan, et al.
Veröffentlicht: (2025)
von: Deng, Hexuan, et al.
Veröffentlicht: (2025)
Reasoning Boosts Opinion Alignment in LLMs
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026)
Code-Guided Reasoning for Small Language Models: Evaluating Executable MCQA Scaffolds
von: Biswas, Prateek, et al.
Veröffentlicht: (2026)
von: Biswas, Prateek, et al.
Veröffentlicht: (2026)
Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning
von: Gu, Yu, et al.
Veröffentlicht: (2026)
von: Gu, Yu, et al.
Veröffentlicht: (2026)
Resa: Transparent Reasoning Models via SAEs
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
Scaling Speculative Decoding with Lookahead Reasoning
von: Fu, Yichao, et al.
Veröffentlicht: (2025)
von: Fu, Yichao, et al.
Veröffentlicht: (2025)
GRPO++: Enhancing Dermatological Reasoning under Low Resource Settings
von: Swapnil, Ismam Nur, et al.
Veröffentlicht: (2025)
von: Swapnil, Ismam Nur, et al.
Veröffentlicht: (2025)
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
von: Shetty, Pranav, et al.
Veröffentlicht: (2025)
von: Shetty, Pranav, et al.
Veröffentlicht: (2025)
LAB: Large-Scale Alignment for ChatBots
von: Sudalairaj, Shivchander, et al.
Veröffentlicht: (2024)
von: Sudalairaj, Shivchander, et al.
Veröffentlicht: (2024)
Efficiently Scaling LLM Reasoning with Certaindex
von: Fu, Yichao, et al.
Veröffentlicht: (2024)
von: Fu, Yichao, et al.
Veröffentlicht: (2024)
Reasoning: From Reflection to Solution
von: Li, Zixi
Veröffentlicht: (2025)
von: Li, Zixi
Veröffentlicht: (2025)
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
von: Zhao, Ziqi, et al.
Veröffentlicht: (2026)
von: Zhao, Ziqi, et al.
Veröffentlicht: (2026)
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples
von: Gao, Chengqian, et al.
Veröffentlicht: (2025)
von: Gao, Chengqian, et al.
Veröffentlicht: (2025)
DSDR: Dual-Scale Diversity Regularization for Exploration in LLM Reasoning
von: Wan, Zhongwei, et al.
Veröffentlicht: (2026)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2026)
ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
von: Zhou, Ruiyang, et al.
Veröffentlicht: (2025)
von: Zhou, Ruiyang, et al.
Veröffentlicht: (2025)
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
von: Chang, Shuochen, et al.
Veröffentlicht: (2026)
von: Chang, Shuochen, et al.
Veröffentlicht: (2026)
TTSR: Test-Time Self-Reflection for Continual Reasoning Improvement
von: He, Haoyang, et al.
Veröffentlicht: (2026)
von: He, Haoyang, et al.
Veröffentlicht: (2026)
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
Beyond Neural Incompatibility: Cross-Scale Knowledge Transfer in Language Models through Latent Semantic Alignment
von: Gu, Jian, et al.
Veröffentlicht: (2025)
von: Gu, Jian, et al.
Veröffentlicht: (2025)
Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
von: Xu, Qingshu, et al.
Veröffentlicht: (2025)
von: Xu, Qingshu, et al.
Veröffentlicht: (2025)
Critique-Guided Distillation for Robust Reasoning via Refinement
von: Kapusuzoglu, Berkcan, et al.
Veröffentlicht: (2025)
von: Kapusuzoglu, Berkcan, et al.
Veröffentlicht: (2025)
ARGS: Alignment as Reward-Guided Search
von: Khanov, Maxim, et al.
Veröffentlicht: (2024)
von: Khanov, Maxim, et al.
Veröffentlicht: (2024)
Scaling Open-Ended Reasoning to Predict the Future
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
von: Chen, Hao, et al.
Veröffentlicht: (2026)
von: Chen, Hao, et al.
Veröffentlicht: (2026)
Beyond Labels: Aligning Large Language Models with Human-like Reasoning
von: Kabir, Muhammad Rafsan, et al.
Veröffentlicht: (2024)
von: Kabir, Muhammad Rafsan, et al.
Veröffentlicht: (2024)
KG-TRICK: Unifying Textual and Relational Information Completion of Knowledge for Multilingual Knowledge Graphs
von: Zhou, Zelin, et al.
Veröffentlicht: (2025)
von: Zhou, Zelin, et al.
Veröffentlicht: (2025)
Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
von: Hu, Jingcheng, et al.
Veröffentlicht: (2025)
von: Hu, Jingcheng, et al.
Veröffentlicht: (2025)
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
von: Ko, Jongwoo, et al.
Veröffentlicht: (2026)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2026)
Reinforce LLM Reasoning through Multi-Agent Reflection
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
A Hitchhiker's Guide to Scaling Law Estimation
von: Choshen, Leshem, et al.
Veröffentlicht: (2024)
von: Choshen, Leshem, et al.
Veröffentlicht: (2024)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
Transparent Neighborhood Approximation for Text Classifier Explanation
von: Cai, Yi, et al.
Veröffentlicht: (2024)
von: Cai, Yi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Preferences: Learning Alignment Principles Grounded in Human Reasons and Values
von: Bell, Henry, et al.
Veröffentlicht: (2026) -
Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression
von: Huang, Jiameng, et al.
Veröffentlicht: (2025) -
Transfer Q Star: Principled Decoding for LLM Alignment
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024) -
Test-Time Scaling with Reflective Generative Model
von: Wang, Zixiao, et al.
Veröffentlicht: (2025) -
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
von: Zhang, Yunfan, et al.
Veröffentlicht: (2025)