What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Ming, Li, Yanhong, Zhou, Tianyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Agents Thinking Fast and Slow: A Talker-Reasoner Architecture
von: Christakopoulou, Konstantina, et al.
Veröffentlicht: (2024)
von: Christakopoulou, Konstantina, et al.
Veröffentlicht: (2024)
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
von: Wang, Shouren, et al.
Veröffentlicht: (2025)
von: Wang, Shouren, et al.
Veröffentlicht: (2025)
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition
von: Ye, Guanghao, et al.
Veröffentlicht: (2025)
von: Ye, Guanghao, et al.
Veröffentlicht: (2025)
Stable Adaptive Thinking via Advantage Shaping and Length-Aware Gradient Regulation
von: Xu, Zihang, et al.
Veröffentlicht: (2026)
von: Xu, Zihang, et al.
Veröffentlicht: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
Think Globally, Group Locally: Evaluating LLMs Using Multi-Lingual Word Grouping Games
von: Guerra-Solano, César, et al.
Veröffentlicht: (2025)
von: Guerra-Solano, César, et al.
Veröffentlicht: (2025)
Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
von: Wang, Ziyan, et al.
Veröffentlicht: (2025)
Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective
von: Xu, Chengyin, et al.
Veröffentlicht: (2025)
von: Xu, Chengyin, et al.
Veröffentlicht: (2025)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
von: Hong, Yining, et al.
Veröffentlicht: (2024)
von: Hong, Yining, et al.
Veröffentlicht: (2024)
Do Multilingual LLMs Think In English?
von: Schut, Lisa, et al.
Veröffentlicht: (2025)
von: Schut, Lisa, et al.
Veröffentlicht: (2025)
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
System-1.x: Learning to Balance Fast and Slow Planning with Language Models
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2024)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2024)
FADE: Why Bad Descriptions Happen to Good Features
von: Puri, Bruno, et al.
Veröffentlicht: (2025)
von: Puri, Bruno, et al.
Veröffentlicht: (2025)
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2025)
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2025)
RuleR: Improving LLM Controllability by Rule-based Data Recycling
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Thinker: Learning to Think Fast and Slow
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
Continuous Approximations for Improving Quantization Aware Training of LLMs
von: Li, He, et al.
Veröffentlicht: (2024)
von: Li, He, et al.
Veröffentlicht: (2024)
MeTHanol: Modularized Thinking Language Models with Intermediate Layer Thinking, Decoding and Bootstrapping Reasoning
von: Xi, Ningyuan, et al.
Veröffentlicht: (2024)
von: Xi, Ningyuan, et al.
Veröffentlicht: (2024)
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
von: Kang, Yipeng, et al.
Veröffentlicht: (2024)
von: Kang, Yipeng, et al.
Veröffentlicht: (2024)
Reverse Thinking Makes LLMs Stronger Reasoners
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden Principles
von: Wei, Qingyan, et al.
Veröffentlicht: (2025)
von: Wei, Qingyan, et al.
Veröffentlicht: (2025)
Not All Layers of LLMs Are Necessary During Inference
von: Fan, Siqi, et al.
Veröffentlicht: (2024)
von: Fan, Siqi, et al.
Veröffentlicht: (2024)
Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents
von: Li, Xirui, et al.
Veröffentlicht: (2026)
von: Li, Xirui, et al.
Veröffentlicht: (2026)
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Reasoning Models Don't Always Say What They Think
von: Chen, Yanda, et al.
Veröffentlicht: (2025)
von: Chen, Yanda, et al.
Veröffentlicht: (2025)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
von: Lv, Keyu, et al.
Veröffentlicht: (2026)
von: Lv, Keyu, et al.
Veröffentlicht: (2026)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
A Decomposition Perspective to Long-context Reasoning for LLMs
von: Xiao, Yanling, et al.
Veröffentlicht: (2026)
von: Xiao, Yanling, et al.
Veröffentlicht: (2026)
Federated Learning with Layer Skipping: Efficient Training of Large Language Models for Healthcare NLP
von: Zhang, Lihong, et al.
Veröffentlicht: (2025)
von: Zhang, Lihong, et al.
Veröffentlicht: (2025)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
ATLaS: Agent Tuning via Learning Critical Steps
von: Chen, Zhixun, et al.
Veröffentlicht: (2025)
von: Chen, Zhixun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
von: Li, Ming, et al.
Veröffentlicht: (2025) -
Agents Thinking Fast and Slow: A Talker-Reasoner Architecture
von: Christakopoulou, Konstantina, et al.
Veröffentlicht: (2024) -
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
von: Fan, Chenrui, et al.
Veröffentlicht: (2025) -
Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory
von: Li, Ming, et al.
Veröffentlicht: (2025) -
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
von: Li, Ming, et al.
Veröffentlicht: (2024)