Understanding Linear Probing then Fine-tuning Language Models from NTK Perspective
Fuente:
arXiv
Guardado en:
| Autores principales: | Tomihari, Akiyoshi, Sato, Issei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Understanding Transformer Optimization via Gradient Heterogeneity
por: Tomihari, Akiyoshi, et al.
Publicado: (2025)
por: Tomihari, Akiyoshi, et al.
Publicado: (2025)
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation
por: Tomihari, Akiyoshi, et al.
Publicado: (2026)
por: Tomihari, Akiyoshi, et al.
Publicado: (2026)
Learning Dynamics in RL Post-Training for Language Models
por: Tomihari, Akiyoshi
Publicado: (2026)
por: Tomihari, Akiyoshi
Publicado: (2026)
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
por: Tomihari, Akiyoshi, et al.
Publicado: (2025)
por: Tomihari, Akiyoshi, et al.
Publicado: (2025)
Understanding Generalization in Physics Informed Models through Affine Variety Dimensions
por: Koshizuka, Takeshi, et al.
Publicado: (2025)
por: Koshizuka, Takeshi, et al.
Publicado: (2025)
Understanding the Expressivity and Trainability of Fourier Neural Operator: A Mean-Field Perspective
por: Koshizuka, Takeshi, et al.
Publicado: (2023)
por: Koshizuka, Takeshi, et al.
Publicado: (2023)
Can Test-time Computation Mitigate Reproduction Bias in Neural Symbolic Regression?
por: Sato, Shun, et al.
Publicado: (2025)
por: Sato, Shun, et al.
Publicado: (2025)
Understanding NTK Variance in Implicit Neural Representations
por: Ou, Chengguang, et al.
Publicado: (2025)
por: Ou, Chengguang, et al.
Publicado: (2025)
Understanding Fine-tuning in Approximate Unlearning: A Theoretical Perspective
por: Ding, Meng, et al.
Publicado: (2024)
por: Ding, Meng, et al.
Publicado: (2024)
End-to-End Training Induces Information Bottleneck through Layer-Role Differentiation: A Comparative Analysis with Layer-wise Training
por: Sakamoto, Keitaro, et al.
Publicado: (2024)
por: Sakamoto, Keitaro, et al.
Publicado: (2024)
Benign Overfitting in Token Selection of Attention Mechanism
por: Sakamoto, Keitaro, et al.
Publicado: (2024)
por: Sakamoto, Keitaro, et al.
Publicado: (2024)
Multiplicative Logit Adjustment Approximates Neural-Collapse-Aware Decision Boundary Adjustment
por: Hasegawa, Naoya, et al.
Publicado: (2024)
por: Hasegawa, Naoya, et al.
Publicado: (2024)
On Expressive Power of Looped Transformers: Theoretical Analysis and Enhancement via Timestep Encoding
por: Xu, Kevin, et al.
Publicado: (2024)
por: Xu, Kevin, et al.
Publicado: (2024)
Top-Down Bayesian Posterior Sampling for Sum-Product Networks
por: Yokoi, Soma, et al.
Publicado: (2024)
por: Yokoi, Soma, et al.
Publicado: (2024)
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
por: Sakamoto, Keitaro, et al.
Publicado: (2025)
por: Sakamoto, Keitaro, et al.
Publicado: (2025)
Exploring Weight Balancing on Long-Tailed Recognition Problem
por: Hasegawa, Naoya, et al.
Publicado: (2023)
por: Hasegawa, Naoya, et al.
Publicado: (2023)
Label-NTK Alignments and A Tighter Convergence Bound in the NTK Regime
por: Marreddy, Ruchirinkil, et al.
Publicado: (2026)
por: Marreddy, Ruchirinkil, et al.
Publicado: (2026)
Probe-based Fine-tuning for Reducing Toxicity
por: Wehner, Jan, et al.
Publicado: (2025)
por: Wehner, Jan, et al.
Publicado: (2025)
Takeuchi's Information Criteria as Generalization Measures for DNNs Close to NTK Regime
por: Naganuma, Hiroki, et al.
Publicado: (2026)
por: Naganuma, Hiroki, et al.
Publicado: (2026)
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
por: Wenger, Jonathan, et al.
Publicado: (2023)
por: Wenger, Jonathan, et al.
Publicado: (2023)
Rethinking Associative Memory Mechanism in Induction Head
por: Wang, Shuo, et al.
Publicado: (2024)
por: Wang, Shuo, et al.
Publicado: (2024)
On the Optimal Memorization Capacity of Transformers
por: Kajitsuka, Tokio, et al.
Publicado: (2024)
por: Kajitsuka, Tokio, et al.
Publicado: (2024)
Max-pooling Network Revisited: Analyzing the Role of Semantic Probability in Multiple Instance Learning for Hallucination Detection
por: Fujikawa, Shota, et al.
Publicado: (2026)
por: Fujikawa, Shota, et al.
Publicado: (2026)
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
por: Xu, Kevin, et al.
Publicado: (2025)
por: Xu, Kevin, et al.
Publicado: (2025)
Fix Initial Codes and Iteratively Refine Textual Directions Toward Safe Multi-Turn Code Correction
por: Tanaka, Yuto, et al.
Publicado: (2026)
por: Tanaka, Yuto, et al.
Publicado: (2026)
Training NTK to Generalize with KARE
por: Schwab, Johannes, et al.
Publicado: (2025)
por: Schwab, Johannes, et al.
Publicado: (2025)
Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU?
por: Hościłowicz, Jakub, et al.
Publicado: (2023)
por: Hościłowicz, Jakub, et al.
Publicado: (2023)
A Formal Comparison Between Chain of Thought and Latent Thought
por: Xu, Kevin, et al.
Publicado: (2025)
por: Xu, Kevin, et al.
Publicado: (2025)
Adversarial Robustness of NTK Neural Networks
por: Hou, Yuxuan
Publicado: (2026)
por: Hou, Yuxuan
Publicado: (2026)
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
por: Kajitsuka, Tokio, et al.
Publicado: (2023)
por: Kajitsuka, Tokio, et al.
Publicado: (2023)
Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective
por: Ghosh, Bishwamittra, et al.
Publicado: (2026)
por: Ghosh, Bishwamittra, et al.
Publicado: (2026)
MLPs at the EOC: Spectrum of the NTK
por: Terjék, Dávid, et al.
Publicado: (2025)
por: Terjék, Dávid, et al.
Publicado: (2025)
MLPs at the EOC: Concentration of the NTK
por: Terjék, Dávid, et al.
Publicado: (2025)
por: Terjék, Dávid, et al.
Publicado: (2025)
On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm Perspective
por: Xie, Zeke, et al.
Publicado: (2020)
por: Xie, Zeke, et al.
Publicado: (2020)
A Study of Optimizations for Fine-tuning Large Language Models
por: Singh, Arjun, et al.
Publicado: (2024)
por: Singh, Arjun, et al.
Publicado: (2024)
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
por: Je, Gyung Hyun, et al.
Publicado: (2025)
por: Je, Gyung Hyun, et al.
Publicado: (2025)
Shortcut Features as Top Eigenfunctions of NTK: A Linear Neural Network Case and More
por: Lim, Jinwoo, et al.
Publicado: (2026)
por: Lim, Jinwoo, et al.
Publicado: (2026)
A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding
por: Shen, Yiqing, et al.
Publicado: (2024)
por: Shen, Yiqing, et al.
Publicado: (2024)
Feature Identification via the Empirical NTK
por: Lin, Jennifer
Publicado: (2025)
por: Lin, Jennifer
Publicado: (2025)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
por: Kim, Minseon, et al.
Publicado: (2025)
por: Kim, Minseon, et al.
Publicado: (2025)
Ejemplares similares
-
Understanding Transformer Optimization via Gradient Heterogeneity
por: Tomihari, Akiyoshi, et al.
Publicado: (2025) -
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation
por: Tomihari, Akiyoshi, et al.
Publicado: (2026) -
Learning Dynamics in RL Post-Training for Language Models
por: Tomihari, Akiyoshi
Publicado: (2026) -
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
por: Tomihari, Akiyoshi, et al.
Publicado: (2025) -
Understanding Generalization in Physics Informed Models through Affine Variety Dimensions
por: Koshizuka, Takeshi, et al.
Publicado: (2025)