Large Reasoning Models Learn Better Alignment from Flawed Thinking
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, ShengYun, Smith, Eric, Evtimov, Ivan, Jiang, Song, Chen, Pin-Yu, Zhan, Hongyuan, Wang, Haozhu, Chau, Duen Horng, Pasupuleti, Mahesh, Chi, Jianfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Shape it Up! Restoring LLM Safety during Finetuning
by: Peng, ShengYun, et al.
Published: (2025)
by: Peng, ShengYun, et al.
Published: (2025)
Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
by: Peng, ShengYun, et al.
Published: (2024)
by: Peng, ShengYun, et al.
Published: (2024)
Self-Supervised Pre-Training for Table Structure Recognition Transformer
by: Peng, ShengYun, et al.
Published: (2024)
by: Peng, ShengYun, et al.
Published: (2024)
UniTable: Towards a Unified Framework for Table Recognition via Self-Supervised Pretraining
by: Peng, ShengYun, et al.
Published: (2024)
by: Peng, ShengYun, et al.
Published: (2024)
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
by: Lee, Seongmin, et al.
Published: (2025)
by: Lee, Seongmin, et al.
Published: (2025)
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
by: Phute, Mansi, et al.
Published: (2023)
by: Phute, Mansi, et al.
Published: (2023)
The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
by: Zhang, Jingyu, et al.
Published: (2025)
by: Zhang, Jingyu, et al.
Published: (2025)
Diffusion Explorer: Interactive Exploration of Diffusion Models
by: Helbling, Alec, et al.
Published: (2025)
by: Helbling, Alec, et al.
Published: (2025)
LLM Attributor: Interactive Visual Attribution for LLM Generation
by: Lee, Seongmin, et al.
Published: (2024)
by: Lee, Seongmin, et al.
Published: (2024)
UNDREAM: Bridging Differentiable Rendering and Photorealistic Simulation for End-to-end Adversarial Attacks
by: Phute, Mansi, et al.
Published: (2025)
by: Phute, Mansi, et al.
Published: (2025)
MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation
by: Wang, Zijie J., et al.
Published: (2024)
by: Wang, Zijie J., et al.
Published: (2024)
Effective Guidance for Model Attention with Simple Yes-no Annotations
by: Lee, Seongmin, et al.
Published: (2024)
by: Lee, Seongmin, et al.
Published: (2024)
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
by: Lee, Seongmin, et al.
Published: (2023)
by: Lee, Seongmin, et al.
Published: (2023)
Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
by: Chi, Jianfeng, et al.
Published: (2024)
by: Chi, Jianfeng, et al.
Published: (2024)
Nested Fusion: A Method for Learning High Resolution Latent Structure of Multi-Scale Measurement Data on Mars
by: Wright, Austin P., et al.
Published: (2024)
by: Wright, Austin P., et al.
Published: (2024)
SuperNOVA: Design Strategies and Opportunities for Interactive Visualization in Computational Notebooks
by: Wang, Zijie J., et al.
Published: (2023)
by: Wang, Zijie J., et al.
Published: (2023)
Probing LLM Hallucination from Within: Perturbation-Driven Approach via Internal Knowledge
by: Lee, Seongmin, et al.
Published: (2024)
by: Lee, Seongmin, et al.
Published: (2024)
Wordflow: Social Prompt Engineering for Large Language Models
by: Wang, Zijie J., et al.
Published: (2024)
by: Wang, Zijie J., et al.
Published: (2024)
UNIPO: Unified Interactive Visual Explanation for RL Fine-Tuning Policy Optimization
by: Cho, Aeree, et al.
Published: (2026)
by: Cho, Aeree, et al.
Published: (2026)
Dense Associative Memory Through the Lens of Random Features
by: Hoover, Benjamin, et al.
Published: (2024)
by: Hoover, Benjamin, et al.
Published: (2024)
What Time Is It? How Data Geometry Makes Time Conditioning Optional for Flow Matching
by: Helbling, Alec, et al.
Published: (2026)
by: Helbling, Alec, et al.
Published: (2026)
Persistent Pre-Training Poisoning of LLMs
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
ARCollab: Towards Multi-User Interactive Cardiovascular Surgical Planning in Mobile Augmented Reality
by: Mehta, Pratham, et al.
Published: (2024)
by: Mehta, Pratham, et al.
Published: (2024)
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
by: Helbling, Alec, et al.
Published: (2025)
by: Helbling, Alec, et al.
Published: (2025)
Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models
by: Hoover, Benjamin, et al.
Published: (2023)
by: Hoover, Benjamin, et al.
Published: (2023)
LitForager: Exploring Multimodal Literature Foraging Strategies in Immersive Sensemaking
by: Yang, Haoyang, et al.
Published: (2025)
by: Yang, Haoyang, et al.
Published: (2025)
HybridCollab: Unifying In-Person and Remote Collaboration for Cardiovascular Surgical Planning in Mobile Augmented Reality
by: Mehta, Pratham Darrpan, et al.
Published: (2025)
by: Mehta, Pratham Darrpan, et al.
Published: (2025)
Mobile Fitting Room: On-device Virtual Try-on via Diffusion Models
by: Blalock, Justin, et al.
Published: (2024)
by: Blalock, Justin, et al.
Published: (2024)
Semi-Truths: A Large-Scale Dataset of AI-Augmented Images for Evaluating Robustness of AI-Generated Image detectors
by: Pal, Anisha, et al.
Published: (2024)
by: Pal, Anisha, et al.
Published: (2024)
Can Large Reasoning Models Improve Accuracy on Mathematical Tasks Using Flawed Thinking?
by: Amjith, Saraswathy, et al.
Published: (2025)
by: Amjith, Saraswathy, et al.
Published: (2025)
FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
by: Ding, Yuyang, et al.
Published: (2025)
by: Ding, Yuyang, et al.
Published: (2025)
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
by: Recchia, Gabriel, et al.
Published: (2025)
by: Recchia, Gabriel, et al.
Published: (2025)
Multi-User Mobile Augmented Reality for Cardiovascular Surgical Planning
by: Mehta, Pratham, et al.
Published: (2024)
by: Mehta, Pratham, et al.
Published: (2024)
Safety Alignment of LMs via Non-cooperative Games
by: Paulus, Anselm, et al.
Published: (2025)
by: Paulus, Anselm, et al.
Published: (2025)
Physics-Informed Residual Learning for Safe and Adaptive Battery Charging Under Extreme Conditions
by: Pasupuleti, Ramakrishna
Published: (2026)
by: Pasupuleti, Ramakrishna
Published: (2026)
A K–R Constant–Based Framework for Predictive Stabilization, Uncertainty Regulation, and Physical Reservoir Computing
by: Pasupuleti, Ramakrishna
Published: (2026)
by: Pasupuleti, Ramakrishna
Published: (2026)
KR-Regulated Nonlinear Parabolic and Fractional Evolution Equations: Attractor Scaling Laws, Critical Thresholds, and Computational Efficiency
by: Pasupuleti, Ramakrishna
Published: (2026)
by: Pasupuleti, Ramakrishna
Published: (2026)
Pre Seismic Quiescence and Dynamical Regime Transitions in the Japan and Chile Earthquake Catalogs Evidence from KR Critical Slowing Down Indicators
by: Pasupuleti, Ramakrishna
Published: (2026)
by: Pasupuleti, Ramakrishna
Published: (2026)
Nitrosonium Ion Catalyzed Oxidative Bromination of Arenes
by: Pin‐Hsien Chen, et al.
Published: (2024)
by: Pin‐Hsien Chen, et al.
Published: (2024)
Detecting Exomoons in Free-Floating-Planet Events from Space-based Microlensing Surveys
by: Fu, Haozhu, et al.
Published: (2025)
by: Fu, Haozhu, et al.
Published: (2025)
Similar Items
-
Shape it Up! Restoring LLM Safety during Finetuning
by: Peng, ShengYun, et al.
Published: (2025) -
Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
by: Peng, ShengYun, et al.
Published: (2024) -
Self-Supervised Pre-Training for Table Structure Recognition Transformer
by: Peng, ShengYun, et al.
Published: (2024) -
UniTable: Towards a Unified Framework for Table Recognition via Self-Supervised Pretraining
by: Peng, ShengYun, et al.
Published: (2024) -
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
by: Lee, Seongmin, et al.
Published: (2025)