Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | Chegini, Atoosa, Kazemi, Hamid, Souza, Garrett, Safi, Maria, Song, Yang, Bengio, Samy, Williamson, Sinead, Farajtabar, Mehrdad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
por: Kazemi, Hamid, et al.
Publicado: (2026)
por: Kazemi, Hamid, et al.
Publicado: (2026)
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
por: Chegini, Atoosa, et al.
Publicado: (2024)
por: Chegini, Atoosa, et al.
Publicado: (2024)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
por: Mirzadeh, Iman, et al.
Publicado: (2024)
por: Mirzadeh, Iman, et al.
Publicado: (2024)
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
por: Shojaee, Parshin, et al.
Publicado: (2025)
por: Shojaee, Parshin, et al.
Publicado: (2025)
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
por: Chegini, Atoosa, et al.
Publicado: (2026)
por: Chegini, Atoosa, et al.
Publicado: (2026)
RePanda: Pandas-powered Tabular Verification and Reasoning
por: Chegini, Atoosa Malemir, et al.
Publicado: (2025)
por: Chegini, Atoosa Malemir, et al.
Publicado: (2025)
RL for Reasoning by Adaptively Revealing Rationales
por: Amani, Mohammad Hossein, et al.
Publicado: (2025)
por: Amani, Mohammad Hossein, et al.
Publicado: (2025)
What do we learn from inverting CLIP models?
por: Kazemi, Hamid, et al.
Publicado: (2024)
por: Kazemi, Hamid, et al.
Publicado: (2024)
How Far Can Transformers Reason? The Globality Barrier and Inductive Scratchpad
por: Abbe, Emmanuel, et al.
Publicado: (2024)
por: Abbe, Emmanuel, et al.
Publicado: (2024)
Generalization on the Unseen, Logic Reasoning and Degree Curriculum
por: Abbe, Emmanuel, et al.
Publicado: (2023)
por: Abbe, Emmanuel, et al.
Publicado: (2023)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
por: Gao, Silin, et al.
Publicado: (2025)
por: Gao, Silin, et al.
Publicado: (2025)
Chain-of-Sketch: Enabling Global Visual Reasoning
por: Lotfi, Aryo, et al.
Publicado: (2024)
por: Lotfi, Aryo, et al.
Publicado: (2024)
Reasoning Can Hurt the Inductive Abilities of Large Language Models
por: Jin, Haibo, et al.
Publicado: (2025)
por: Jin, Haibo, et al.
Publicado: (2025)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
por: Armandpour, Mohammadreza, et al.
Publicado: (2026)
por: Armandpour, Mohammadreza, et al.
Publicado: (2026)
Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks
por: Saberi, Mehrdad, et al.
Publicado: (2023)
por: Saberi, Mehrdad, et al.
Publicado: (2023)
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
por: Liu, Ruikang, et al.
Publicado: (2025)
por: Liu, Ruikang, et al.
Publicado: (2025)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
por: Li, Jeffrey, et al.
Publicado: (2025)
por: Li, Jeffrey, et al.
Publicado: (2025)
Perfect‐Recall and Bootstrapping Reasoning
por: Michael Cohen
Publicado: (2025)
por: Michael Cohen
Publicado: (2025)
Trace Length is a Simple Uncertainty Signal in Reasoning Models
por: Devic, Siddartha, et al.
Publicado: (2025)
por: Devic, Siddartha, et al.
Publicado: (2025)
AI Safety for Everyone
por: Gyevnar, Balint, et al.
Publicado: (2025)
por: Gyevnar, Balint, et al.
Publicado: (2025)
Self-Supervised Learning with Gaussian Processes
por: Duan, Yunshan, et al.
Publicado: (2025)
por: Duan, Yunshan, et al.
Publicado: (2025)
Posterior Uncertainty Quantification in Neural Networks using Data Augmentation
por: Wu, Luhuan, et al.
Publicado: (2024)
por: Wu, Luhuan, et al.
Publicado: (2024)
Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning
por: Sun, Yunxin, et al.
Publicado: (2025)
por: Sun, Yunxin, et al.
Publicado: (2025)
Conformal Thinking: Risk Control for Reasoning on a Compute Budget
por: Wang, Xi, et al.
Publicado: (2026)
por: Wang, Xi, et al.
Publicado: (2026)
CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning
por: He, Jie, et al.
Publicado: (2025)
por: He, Jie, et al.
Publicado: (2025)
Can Large Reasoning Models Improve Accuracy on Mathematical Tasks Using Flawed Thinking?
por: Amjith, Saraswathy, et al.
Publicado: (2025)
por: Amjith, Saraswathy, et al.
Publicado: (2025)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
por: Alizadeh, Keivan, et al.
Publicado: (2026)
por: Alizadeh, Keivan, et al.
Publicado: (2026)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
por: Wang, Zhecan, et al.
Publicado: (2024)
por: Wang, Zhecan, et al.
Publicado: (2024)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
por: Maheswaran, Monishwaran, et al.
Publicado: (2025)
por: Maheswaran, Monishwaran, et al.
Publicado: (2025)
Fast Adversarial Attacks on Language Models In One GPU Minute
por: Sadasivan, Vinu Sankar, et al.
Publicado: (2024)
por: Sadasivan, Vinu Sankar, et al.
Publicado: (2024)
Causal Temporal Reasoning for Markov Decision Processes
por: Kazemi, Milad, et al.
Publicado: (2022)
por: Kazemi, Milad, et al.
Publicado: (2022)
To Copy or Not to Copy: Copying Is Easier to Induce Than Recall
por: Farahani, Mehrdad, et al.
Publicado: (2026)
por: Farahani, Mehrdad, et al.
Publicado: (2026)
Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
por: Xu, Jillian, et al.
Publicado: (2025)
por: Xu, Jillian, et al.
Publicado: (2025)
Estimation of genetic and environmental relationships between milk yield and different measures of mastitis and hyperkeratosis in Holstein cows
por: Arash Chegini
Publicado: (2016)
por: Arash Chegini
Publicado: (2016)
CAMRI Loss: Improving Recall of a Specific Class without Sacrificing Accuracy
por: Nishiyama, Daiki, et al.
Publicado: (2022)
por: Nishiyama, Daiki, et al.
Publicado: (2022)
Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning
por: Shi, Jian, et al.
Publicado: (2026)
por: Shi, Jian, et al.
Publicado: (2026)
Psychological Safety of Critical Care Nurses in Finland
por: Anna Garrett, et al.
Publicado: (2026)
por: Anna Garrett, et al.
Publicado: (2026)
Hurt, Baby, Hurt
por: Scott III, William Walter
Publicado: (2025)
por: Scott III, William Walter
Publicado: (2025)
STAIR: Improving Safety Alignment with Introspective Reasoning
por: Zhang, Yichi, et al.
Publicado: (2025)
por: Zhang, Yichi, et al.
Publicado: (2025)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
por: Mehta, Sachin, et al.
Publicado: (2024)
por: Mehta, Sachin, et al.
Publicado: (2024)
Ejemplares similares
-
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
por: Kazemi, Hamid, et al.
Publicado: (2026) -
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
por: Chegini, Atoosa, et al.
Publicado: (2024) -
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
por: Mirzadeh, Iman, et al.
Publicado: (2024) -
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
por: Shojaee, Parshin, et al.
Publicado: (2025) -
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
por: Chegini, Atoosa, et al.
Publicado: (2026)