Differential Smoothing Mitigates Sharpening and Improves LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Gai, Jingchu, Zeng, Guanning, Zhang, Huaqing, Raghunathan, Aditi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
by: Gai, Jingchu, et al.
Published: (2026)
by: Gai, Jingchu, et al.
Published: (2026)
Taming the Curses of Multiagency in Robust Markov Games with Large State Space through Linear Function Approximation
by: Gai, Jingchu, et al.
Published: (2026)
by: Gai, Jingchu, et al.
Published: (2026)
Weight Ensembling Improves Reasoning in Language Models
by: Dang, Xingyu, et al.
Published: (2025)
by: Dang, Xingyu, et al.
Published: (2025)
Mitigating Bias in RAG: Controlling the Embedder
by: Kim, Taeyoun, et al.
Published: (2025)
by: Kim, Taeyoun, et al.
Published: (2025)
Momentum Streams for Optimizer-Inspired Transformers
by: Gai, Jingchu, et al.
Published: (2026)
by: Gai, Jingchu, et al.
Published: (2026)
Homomorphism Expressivity of Spectral Invariant Graph Neural Networks
by: Gai, Jingchu, et al.
Published: (2025)
by: Gai, Jingchu, et al.
Published: (2025)
Reasoning as an Adaptive Defense for Safety
by: Kim, Taeyoun, et al.
Published: (2025)
by: Kim, Taeyoun, et al.
Published: (2025)
Memorization Sinks: Isolating Memorization during LLM Training
by: Ghosal, Gaurav R., et al.
Published: (2025)
by: Ghosal, Gaurav R., et al.
Published: (2025)
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
by: Zhong, Ziqian, et al.
Published: (2025)
by: Zhong, Ziqian, et al.
Published: (2025)
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
by: Watts, Ishaan, et al.
Published: (2026)
by: Watts, Ishaan, et al.
Published: (2026)
Why is SAM Robust to Label Noise?
by: Baek, Christina, et al.
Published: (2024)
by: Baek, Christina, et al.
Published: (2024)
Lossless Anti-Distillation Sampling
by: Diao, Zibo, et al.
Published: (2026)
by: Diao, Zibo, et al.
Published: (2026)
Sharpen Your Flow: Sharpness-Aware Sampling for Flow Matching
by: Gupta, Aditi, et al.
Published: (2026)
by: Gupta, Aditi, et al.
Published: (2026)
Multitask Learning Can Improve Worst-Group Outcomes
by: Kulkarni, Atharva, et al.
Published: (2023)
by: Kulkarni, Atharva, et al.
Published: (2023)
Self-Trained Verification for Training- and Test-Time Self-Improvement
by: Wu, Chen Henry, et al.
Published: (2026)
by: Wu, Chen Henry, et al.
Published: (2026)
Sharpness-Aware Minimization Enhances Feature Quality via Balanced Learning
by: Springer, Jacob Mitchell, et al.
Published: (2024)
by: Springer, Jacob Mitchell, et al.
Published: (2024)
Early Data Exposure Improves Robustness to Subsequent Fine-Tuning
by: Feng, Lawrence, et al.
Published: (2026)
by: Feng, Lawrence, et al.
Published: (2026)
Repetition Improves Language Model Embeddings
by: Springer, Jacob Mitchell, et al.
Published: (2024)
by: Springer, Jacob Mitchell, et al.
Published: (2024)
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
by: Zhong, Ziqian, et al.
Published: (2025)
by: Zhong, Ziqian, et al.
Published: (2025)
Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable Rewards
by: Zeng, Guanning, et al.
Published: (2025)
by: Zeng, Guanning, et al.
Published: (2025)
Understanding Finetuning for Factual Knowledge Extraction
by: Ghosal, Gaurav, et al.
Published: (2024)
by: Ghosal, Gaurav, et al.
Published: (2024)
Breaking the Curse of Multiagency in Robust Multi-Agent Reinforcement Learning
by: Shi, Laixi, et al.
Published: (2024)
by: Shi, Laixi, et al.
Published: (2024)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
by: Kim, Taeyoun, et al.
Published: (2024)
by: Kim, Taeyoun, et al.
Published: (2024)
Understanding Catastrophic Forgetting in Language Models via Implicit Inference
by: Kotha, Suhas, et al.
Published: (2023)
by: Kotha, Suhas, et al.
Published: (2023)
R$^2$PO: Decoupling Training Trajectories from Inference Responses for LLM Reasoning
by: Wang, Jingchu, et al.
Published: (2026)
by: Wang, Jingchu, et al.
Published: (2026)
Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation
by: Nguyen, Hieu, et al.
Published: (2025)
by: Nguyen, Hieu, et al.
Published: (2025)
Beyond Weisfeiler-Lehman: A Quantitative Framework for GNN Expressiveness
by: Zhang, Bohang, et al.
Published: (2024)
by: Zhang, Bohang, et al.
Published: (2024)
Laplacian Score Sharpening for Mitigating Hallucination in Diffusion Models
by: C, Barath Chandran., et al.
Published: (2025)
by: C, Barath Chandran., et al.
Published: (2025)
Mode-Conditioning Unlocks Superior Test-Time Scaling
by: Wu, Chen Henry, et al.
Published: (2025)
by: Wu, Chen Henry, et al.
Published: (2025)
Exact Unlearning of Finetuning Data via Model Merging at Scale
by: Kuo, Kevin, et al.
Published: (2025)
by: Kuo, Kevin, et al.
Published: (2025)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
A Minimalist Example of Edge-of-Stability and Progressive Sharpening
by: Liu, Liming, et al.
Published: (2025)
by: Liu, Liming, et al.
Published: (2025)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
by: Maini, Pratyush, et al.
Published: (2023)
by: Maini, Pratyush, et al.
Published: (2023)
Vulnerability of Text-Matching in ML/AI Conference Reviewer Assignments to Collusions
by: Hsieh, Jhih-Yi, et al.
Published: (2024)
by: Hsieh, Jhih-Yi, et al.
Published: (2024)
Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?
by: Zhong, Ziqian, et al.
Published: (2026)
by: Zhong, Ziqian, et al.
Published: (2026)
Label Smoothing Improves Gradient Ascent in LLM Unlearning
by: Pang, Zirui, et al.
Published: (2025)
by: Pang, Zirui, et al.
Published: (2025)
Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
Predicting the Performance of Foundation Models via Agreement-on-the-Line
by: Saxena, Rahul, et al.
Published: (2024)
by: Saxena, Rahul, et al.
Published: (2024)
Differentiable Conformal Training for LLM Reasoning Factuality
by: Hittesdorf, Nathan, et al.
Published: (2026)
by: Hittesdorf, Nathan, et al.
Published: (2026)
An Improved Privacy and Utility Analysis of Differentially Private SGD with Bounded Domain and Smooth Losses
by: Liang, Hao, et al.
Published: (2025)
by: Liang, Hao, et al.
Published: (2025)
Similar Items
-
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
by: Gai, Jingchu, et al.
Published: (2026) -
Taming the Curses of Multiagency in Robust Markov Games with Large State Space through Linear Function Approximation
by: Gai, Jingchu, et al.
Published: (2026) -
Weight Ensembling Improves Reasoning in Language Models
by: Dang, Xingyu, et al.
Published: (2025) -
Mitigating Bias in RAG: Controlling the Embedder
by: Kim, Taeyoun, et al.
Published: (2025) -
Momentum Streams for Optimizer-Inspired Transformers
by: Gai, Jingchu, et al.
Published: (2026)