Large Reasoning Models Learn Better Alignment from Flawed Thinking
Fuente:
arXiv
Salvato in:
| Autori principali: | Peng, ShengYun, Smith, Eric, Evtimov, Ivan, Jiang, Song, Chen, Pin-Yu, Zhan, Hongyuan, Wang, Haozhu, Chau, Duen Horng, Pasupuleti, Mahesh, Chi, Jianfeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Shape it Up! Restoring LLM Safety during Finetuning
di: Peng, ShengYun, et al.
Pubblicazione: (2025)
di: Peng, ShengYun, et al.
Pubblicazione: (2025)
Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
di: Peng, ShengYun, et al.
Pubblicazione: (2024)
di: Peng, ShengYun, et al.
Pubblicazione: (2024)
Self-Supervised Pre-Training for Table Structure Recognition Transformer
di: Peng, ShengYun, et al.
Pubblicazione: (2024)
di: Peng, ShengYun, et al.
Pubblicazione: (2024)
UniTable: Towards a Unified Framework for Table Recognition via Self-Supervised Pretraining
di: Peng, ShengYun, et al.
Pubblicazione: (2024)
di: Peng, ShengYun, et al.
Pubblicazione: (2024)
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
di: Lee, Seongmin, et al.
Pubblicazione: (2025)
di: Lee, Seongmin, et al.
Pubblicazione: (2025)
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
di: Phute, Mansi, et al.
Pubblicazione: (2023)
di: Phute, Mansi, et al.
Pubblicazione: (2023)
The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
di: Zhang, Jingyu, et al.
Pubblicazione: (2025)
di: Zhang, Jingyu, et al.
Pubblicazione: (2025)
Diffusion Explorer: Interactive Exploration of Diffusion Models
di: Helbling, Alec, et al.
Pubblicazione: (2025)
di: Helbling, Alec, et al.
Pubblicazione: (2025)
LLM Attributor: Interactive Visual Attribution for LLM Generation
di: Lee, Seongmin, et al.
Pubblicazione: (2024)
di: Lee, Seongmin, et al.
Pubblicazione: (2024)
UNDREAM: Bridging Differentiable Rendering and Photorealistic Simulation for End-to-end Adversarial Attacks
di: Phute, Mansi, et al.
Pubblicazione: (2025)
di: Phute, Mansi, et al.
Pubblicazione: (2025)
MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation
di: Wang, Zijie J., et al.
Pubblicazione: (2024)
di: Wang, Zijie J., et al.
Pubblicazione: (2024)
Effective Guidance for Model Attention with Simple Yes-no Annotations
di: Lee, Seongmin, et al.
Pubblicazione: (2024)
di: Lee, Seongmin, et al.
Pubblicazione: (2024)
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
di: Lee, Seongmin, et al.
Pubblicazione: (2023)
di: Lee, Seongmin, et al.
Pubblicazione: (2023)
Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
di: Chi, Jianfeng, et al.
Pubblicazione: (2024)
di: Chi, Jianfeng, et al.
Pubblicazione: (2024)
Nested Fusion: A Method for Learning High Resolution Latent Structure of Multi-Scale Measurement Data on Mars
di: Wright, Austin P., et al.
Pubblicazione: (2024)
di: Wright, Austin P., et al.
Pubblicazione: (2024)
SuperNOVA: Design Strategies and Opportunities for Interactive Visualization in Computational Notebooks
di: Wang, Zijie J., et al.
Pubblicazione: (2023)
di: Wang, Zijie J., et al.
Pubblicazione: (2023)
Probing LLM Hallucination from Within: Perturbation-Driven Approach via Internal Knowledge
di: Lee, Seongmin, et al.
Pubblicazione: (2024)
di: Lee, Seongmin, et al.
Pubblicazione: (2024)
Wordflow: Social Prompt Engineering for Large Language Models
di: Wang, Zijie J., et al.
Pubblicazione: (2024)
di: Wang, Zijie J., et al.
Pubblicazione: (2024)
UNIPO: Unified Interactive Visual Explanation for RL Fine-Tuning Policy Optimization
di: Cho, Aeree, et al.
Pubblicazione: (2026)
di: Cho, Aeree, et al.
Pubblicazione: (2026)
Dense Associative Memory Through the Lens of Random Features
di: Hoover, Benjamin, et al.
Pubblicazione: (2024)
di: Hoover, Benjamin, et al.
Pubblicazione: (2024)
What Time Is It? How Data Geometry Makes Time Conditioning Optional for Flow Matching
di: Helbling, Alec, et al.
Pubblicazione: (2026)
di: Helbling, Alec, et al.
Pubblicazione: (2026)
Persistent Pre-Training Poisoning of LLMs
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
ARCollab: Towards Multi-User Interactive Cardiovascular Surgical Planning in Mobile Augmented Reality
di: Mehta, Pratham, et al.
Pubblicazione: (2024)
di: Mehta, Pratham, et al.
Pubblicazione: (2024)
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
di: Helbling, Alec, et al.
Pubblicazione: (2025)
di: Helbling, Alec, et al.
Pubblicazione: (2025)
Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models
di: Hoover, Benjamin, et al.
Pubblicazione: (2023)
di: Hoover, Benjamin, et al.
Pubblicazione: (2023)
LitForager: Exploring Multimodal Literature Foraging Strategies in Immersive Sensemaking
di: Yang, Haoyang, et al.
Pubblicazione: (2025)
di: Yang, Haoyang, et al.
Pubblicazione: (2025)
HybridCollab: Unifying In-Person and Remote Collaboration for Cardiovascular Surgical Planning in Mobile Augmented Reality
di: Mehta, Pratham Darrpan, et al.
Pubblicazione: (2025)
di: Mehta, Pratham Darrpan, et al.
Pubblicazione: (2025)
Mobile Fitting Room: On-device Virtual Try-on via Diffusion Models
di: Blalock, Justin, et al.
Pubblicazione: (2024)
di: Blalock, Justin, et al.
Pubblicazione: (2024)
Semi-Truths: A Large-Scale Dataset of AI-Augmented Images for Evaluating Robustness of AI-Generated Image detectors
di: Pal, Anisha, et al.
Pubblicazione: (2024)
di: Pal, Anisha, et al.
Pubblicazione: (2024)
Can Large Reasoning Models Improve Accuracy on Mathematical Tasks Using Flawed Thinking?
di: Amjith, Saraswathy, et al.
Pubblicazione: (2025)
di: Amjith, Saraswathy, et al.
Pubblicazione: (2025)
FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
di: Ding, Yuyang, et al.
Pubblicazione: (2025)
di: Ding, Yuyang, et al.
Pubblicazione: (2025)
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
di: Recchia, Gabriel, et al.
Pubblicazione: (2025)
di: Recchia, Gabriel, et al.
Pubblicazione: (2025)
Multi-User Mobile Augmented Reality for Cardiovascular Surgical Planning
di: Mehta, Pratham, et al.
Pubblicazione: (2024)
di: Mehta, Pratham, et al.
Pubblicazione: (2024)
Safety Alignment of LMs via Non-cooperative Games
di: Paulus, Anselm, et al.
Pubblicazione: (2025)
di: Paulus, Anselm, et al.
Pubblicazione: (2025)
Physics-Informed Residual Learning for Safe and Adaptive Battery Charging Under Extreme Conditions
di: Pasupuleti, Ramakrishna
Pubblicazione: (2026)
di: Pasupuleti, Ramakrishna
Pubblicazione: (2026)
A K–R Constant–Based Framework for Predictive Stabilization, Uncertainty Regulation, and Physical Reservoir Computing
di: Pasupuleti, Ramakrishna
Pubblicazione: (2026)
di: Pasupuleti, Ramakrishna
Pubblicazione: (2026)
KR-Regulated Nonlinear Parabolic and Fractional Evolution Equations: Attractor Scaling Laws, Critical Thresholds, and Computational Efficiency
di: Pasupuleti, Ramakrishna
Pubblicazione: (2026)
di: Pasupuleti, Ramakrishna
Pubblicazione: (2026)
Pre Seismic Quiescence and Dynamical Regime Transitions in the Japan and Chile Earthquake Catalogs Evidence from KR Critical Slowing Down Indicators
di: Pasupuleti, Ramakrishna
Pubblicazione: (2026)
di: Pasupuleti, Ramakrishna
Pubblicazione: (2026)
Nitrosonium Ion Catalyzed Oxidative Bromination of Arenes
di: Pin‐Hsien Chen, et al.
Pubblicazione: (2024)
di: Pin‐Hsien Chen, et al.
Pubblicazione: (2024)
Detecting Exomoons in Free-Floating-Planet Events from Space-based Microlensing Surveys
di: Fu, Haozhu, et al.
Pubblicazione: (2025)
di: Fu, Haozhu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Shape it Up! Restoring LLM Safety during Finetuning
di: Peng, ShengYun, et al.
Pubblicazione: (2025) -
Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
di: Peng, ShengYun, et al.
Pubblicazione: (2024) -
Self-Supervised Pre-Training for Table Structure Recognition Transformer
di: Peng, ShengYun, et al.
Pubblicazione: (2024) -
UniTable: Towards a Unified Framework for Table Recognition via Self-Supervised Pretraining
di: Peng, ShengYun, et al.
Pubblicazione: (2024) -
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
di: Lee, Seongmin, et al.
Pubblicazione: (2025)