Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Yein, Jeong, Minbyul, Kang, Jaewoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information
by: Park, Yein, et al.
Published: (2025)
by: Park, Yein, et al.
Published: (2025)
ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains
by: Park, Yein, et al.
Published: (2024)
by: Park, Yein, et al.
Published: (2024)
Improving Medical Reasoning through Retrieval and Self-Reflection with Retrieval-Augmented Large Language Models
by: Jeong, Minbyul, et al.
Published: (2024)
by: Jeong, Minbyul, et al.
Published: (2024)
ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
by: Park, Yein, et al.
Published: (2025)
by: Park, Yein, et al.
Published: (2025)
ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway
by: Park, Jueon, et al.
Published: (2026)
by: Park, Jueon, et al.
Published: (2026)
MolDeTox: Evaluating Language Model's Stepwise Fragment Editing for Molecular Detoxification
by: Park, Jueon, et al.
Published: (2026)
by: Park, Jueon, et al.
Published: (2026)
OLAPH: Improving Factuality in Biomedical Long-form Question Answering
by: Jeong, Minbyul, et al.
Published: (2024)
by: Jeong, Minbyul, et al.
Published: (2024)
Healthcare AI GYM for Medical Agents
by: Jeong, Minbyul
Published: (2026)
by: Jeong, Minbyul
Published: (2026)
CoTox: Chain-of-Thought-Based Molecular Toxicity Reasoning and Prediction
by: Park, Jueon, et al.
Published: (2025)
by: Park, Jueon, et al.
Published: (2025)
Assessing LLM Reasoning Steps via Principal Knowledge Grounding
by: Hwang, Hyeon, et al.
Published: (2025)
by: Hwang, Hyeon, et al.
Published: (2025)
Trustworthy Agents for Electronic Health Records through Confidence Estimation
by: Song, Yongwoo, et al.
Published: (2025)
by: Song, Yongwoo, et al.
Published: (2025)
Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
by: Singla, Pratham, et al.
Published: (2025)
by: Singla, Pratham, et al.
Published: (2025)
System Message Generation for User Preferences using Open-Source Models
by: Jeong, Minbyul, et al.
Published: (2025)
by: Jeong, Minbyul, et al.
Published: (2025)
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost
by: Xuan, Richmond Sin Jing, et al.
Published: (2026)
by: Xuan, Richmond Sin Jing, et al.
Published: (2026)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
by: Park, Jungwoo, et al.
Published: (2025)
by: Park, Jungwoo, et al.
Published: (2025)
Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
InvThink: Premortem Reasoning for Safer Language Models
by: Kim, Yubin, et al.
Published: (2025)
by: Kim, Yubin, et al.
Published: (2025)
DialogueReason: Rule-Based RL Sparks Dialogue Reasoning in LLMs
by: Shu, Yubo, et al.
Published: (2025)
by: Shu, Yubo, et al.
Published: (2025)
Multi-agent Long-term 3D Human Pose Forecasting via Interaction-aware Trajectory Conditioning
by: Jeong, Jaewoo, et al.
Published: (2024)
by: Jeong, Jaewoo, et al.
Published: (2024)
Tracing Mathematical Proficiency Through Problem-Solving Processes
by: Park, Jungyang, et al.
Published: (2025)
by: Park, Jungyang, et al.
Published: (2025)
Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
CRADLE-VAE: Enhancing Single-Cell Gene Perturbation Modeling with Counterfactual Reasoning-based Artifact Disentanglement
by: Baek, Seungheun, et al.
Published: (2024)
by: Baek, Seungheun, et al.
Published: (2024)
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning
by: Gan, Siyuan, et al.
Published: (2026)
by: Gan, Siyuan, et al.
Published: (2026)
Monet: Mixture of Monosemantic Experts for Transformers
by: Park, Jungwoo, et al.
Published: (2024)
by: Park, Jungwoo, et al.
Published: (2024)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models
by: An, Sohyun, et al.
Published: (2025)
by: An, Sohyun, et al.
Published: (2025)
GPO-VAE: Modeling Explainable Gene Perturbation Responses utilizing GRN-Aligned Parameter Optimization
by: Baek, Seungheun, et al.
Published: (2025)
by: Baek, Seungheun, et al.
Published: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
Guided Trajectory Generation with Diffusion Models for Offline Model-based Optimization
by: Yun, Taeyoung, et al.
Published: (2024)
by: Yun, Taeyoung, et al.
Published: (2024)
Intrinsically Interpretable Attention via Sparse Post-Training
by: Draye, Florent, et al.
Published: (2025)
by: Draye, Florent, et al.
Published: (2025)
Sparks of Tabular Reasoning via Text2SQL Reinforcement Learning
by: Stoisser, Josefa Lia, et al.
Published: (2025)
by: Stoisser, Josefa Lia, et al.
Published: (2025)
Sparks of Rationality: Do Reasoning LLMs Align with Human Judgment and Choice?
by: Tak, Ala N., et al.
Published: (2026)
by: Tak, Ala N., et al.
Published: (2026)
Efficient Post-Training Refinement of Latent Reasoning in Large Language Models
by: Wang, Xinyuan, et al.
Published: (2025)
by: Wang, Xinyuan, et al.
Published: (2025)
Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
by: Kang, Hyeongyu, et al.
Published: (2025)
by: Kang, Hyeongyu, et al.
Published: (2025)
Think How to Think: Mitigating Overthinking with Autonomous Difficulty Cognition in Large Reasoning Models
by: Liu, Yongjiang, et al.
Published: (2025)
by: Liu, Yongjiang, et al.
Published: (2025)
Adaptive Dual Reasoner: Large Reasoning Models Can Think Efficiently by Hybrid Reasoning
by: Zhang, Yujian, et al.
Published: (2025)
by: Zhang, Yujian, et al.
Published: (2025)
Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
by: Xu, Jillian, et al.
Published: (2025)
by: Xu, Jillian, et al.
Published: (2025)
Emergent Search and Backtracking in Latent Reasoning Models
by: Cui, Jasmine, et al.
Published: (2026)
by: Cui, Jasmine, et al.
Published: (2026)
EP-HDC: Hyperdimensional Computing with Encrypted Parameters for High-Throughput Privacy-Preserving Inference
by: Park, Jaewoo, et al.
Published: (2025)
by: Park, Jaewoo, et al.
Published: (2025)
Similar Items
-
Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information
by: Park, Yein, et al.
Published: (2025) -
ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains
by: Park, Yein, et al.
Published: (2024) -
Improving Medical Reasoning through Retrieval and Self-Reflection with Retrieval-Augmented Large Language Models
by: Jeong, Minbyul, et al.
Published: (2024) -
ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
by: Park, Yein, et al.
Published: (2025) -
ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway
by: Park, Jueon, et al.
Published: (2026)