Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Hongru, Kailkhura, Bhavya, Wang, Zhangyang, Liang, Yingbin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
por: Huang, Ruiquan, et al.
Publicado: (2025)
por: Huang, Ruiquan, et al.
Publicado: (2025)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
por: Yang, Tong, et al.
Publicado: (2024)
por: Yang, Tong, et al.
Publicado: (2024)
Neural Networks with Sparse Activation Induced by Large Bias: Tighter Analysis with Bias-Generalized NTK
por: Yang, Hongru, et al.
Publicado: (2023)
por: Yang, Hongru, et al.
Publicado: (2023)
Constrained Discrete Diffusion
por: Cardei, Michael, et al.
Publicado: (2025)
por: Cardei, Michael, et al.
Publicado: (2025)
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
por: Pal, Soumyadeep, et al.
Publicado: (2025)
por: Pal, Soumyadeep, et al.
Publicado: (2025)
Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence
por: Nava, Andres, et al.
Publicado: (2026)
por: Nava, Andres, et al.
Publicado: (2026)
LoCoCo: Dropping In Convolutions for Long Context Compression
por: Cai, Ruisi, et al.
Publicado: (2024)
por: Cai, Ruisi, et al.
Publicado: (2024)
Constraint-Rectified Training for Efficient Chain-of-Thought
por: Wu, Qinhang, et al.
Publicado: (2026)
por: Wu, Qinhang, et al.
Publicado: (2026)
Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion
por: Christopher, Jacob K, et al.
Publicado: (2024)
por: Christopher, Jacob K, et al.
Publicado: (2024)
Low-rank finetuning for LLMs: A fairness perspective
por: Das, Saswat, et al.
Publicado: (2024)
por: Das, Saswat, et al.
Publicado: (2024)
UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making
por: Duan, Jinhao, et al.
Publicado: (2025)
por: Duan, Jinhao, et al.
Publicado: (2025)
Training Neural Networks as Recognizers of Formal Languages
por: Butoi, Alexandra, et al.
Publicado: (2024)
por: Butoi, Alexandra, et al.
Publicado: (2024)
SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning
por: Jia, Jinghan, et al.
Publicado: (2024)
por: Jia, Jinghan, et al.
Publicado: (2024)
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
por: Duan, Jinhao, et al.
Publicado: (2023)
por: Duan, Jinhao, et al.
Publicado: (2023)
Meta ControlNet: Enhancing Task Adaptation via Meta Learning
por: Yang, Junjie, et al.
Publicado: (2023)
por: Yang, Junjie, et al.
Publicado: (2023)
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
por: Duan, Jinhao, et al.
Publicado: (2024)
por: Duan, Jinhao, et al.
Publicado: (2024)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
por: Geiping, Jonas, et al.
Publicado: (2025)
por: Geiping, Jonas, et al.
Publicado: (2025)
Reciprocal Co-Training (RCT): Coupling Gradient-Based and Non-Differentiable Models via Reinforcement Learning
por: Tian, Yunshuo, et al.
Publicado: (2026)
por: Tian, Yunshuo, et al.
Publicado: (2026)
World Properties without World Models: Recovering Spatial and Temporal Structure from Co-occurrence Statistics in Static Word Embeddings
por: Barenholtz, Elan
Publicado: (2026)
por: Barenholtz, Elan
Publicado: (2026)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
por: Vasudeva, Bhavya, et al.
Publicado: (2026)
por: Vasudeva, Bhavya, et al.
Publicado: (2026)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
por: Cai, Ruisi, et al.
Publicado: (2024)
por: Cai, Ruisi, et al.
Publicado: (2024)
Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models
por: Karkada, Dhruva, et al.
Publicado: (2025)
por: Karkada, Dhruva, et al.
Publicado: (2025)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
por: Deora, Puneesh, et al.
Publicado: (2025)
por: Deora, Puneesh, et al.
Publicado: (2025)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
por: Li, Hongkang, et al.
Publicado: (2024)
por: Li, Hongkang, et al.
Publicado: (2024)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
por: Nguyen, Dang, et al.
Publicado: (2025)
por: Nguyen, Dang, et al.
Publicado: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
por: Yang, Songlin, et al.
Publicado: (2023)
por: Yang, Songlin, et al.
Publicado: (2023)
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
por: Zhao, Siyan, et al.
Publicado: (2025)
por: Zhao, Siyan, et al.
Publicado: (2025)
SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
por: Huang, Tianjin, et al.
Publicado: (2025)
por: Huang, Tianjin, et al.
Publicado: (2025)
Tracking the Feature Dynamics in LLM Training: A Mechanistic Study
por: Xu, Yang, et al.
Publicado: (2024)
por: Xu, Yang, et al.
Publicado: (2024)
On the Duality between Gradient Transformations and Adapters
por: Torroba-Hennigen, Lucas, et al.
Publicado: (2025)
por: Torroba-Hennigen, Lucas, et al.
Publicado: (2025)
HIPO: Instruction Hierarchy via Constrained Reinforcement Learning
por: Chen, Keru, et al.
Publicado: (2026)
por: Chen, Keru, et al.
Publicado: (2026)
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages
por: Yang, Andy, et al.
Publicado: (2023)
por: Yang, Andy, et al.
Publicado: (2023)
A Gradient Analysis Framework for Rewarding Good and Penalizing Bad Examples in Language Models
por: Tuan, Yi-Lin, et al.
Publicado: (2024)
por: Tuan, Yi-Lin, et al.
Publicado: (2024)
DFPO: Scaling Value Modeling via Distributional Flow towards Robust and Generalizable LLM Post-Training
por: Zhu, Dingwei, et al.
Publicado: (2026)
por: Zhu, Dingwei, et al.
Publicado: (2026)
ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
por: Meshram, Pragati Shuddhodhan, et al.
Publicado: (2024)
por: Meshram, Pragati Shuddhodhan, et al.
Publicado: (2024)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
por: Huang, Zeyi, et al.
Publicado: (2026)
por: Huang, Zeyi, et al.
Publicado: (2026)
A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
por: Hoang, Nhat M., et al.
Publicado: (2025)
por: Hoang, Nhat M., et al.
Publicado: (2025)
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
por: Bartoldson, Brian R., et al.
Publicado: (2024)
por: Bartoldson, Brian R., et al.
Publicado: (2024)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
por: Jaiswal, Ajay, et al.
Publicado: (2023)
por: Jaiswal, Ajay, et al.
Publicado: (2023)
Non-asymptotic Convergence of Training Transformers for Next-token Prediction
por: Huang, Ruiquan, et al.
Publicado: (2024)
por: Huang, Ruiquan, et al.
Publicado: (2024)
Ejemplares similares
-
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
por: Huang, Ruiquan, et al.
Publicado: (2025) -
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
por: Yang, Tong, et al.
Publicado: (2024) -
Neural Networks with Sparse Activation Induced by Large Bias: Tighter Analysis with Bias-Generalized NTK
por: Yang, Hongru, et al.
Publicado: (2023) -
Constrained Discrete Diffusion
por: Cardei, Michael, et al.
Publicado: (2025) -
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
por: Pal, Soumyadeep, et al.
Publicado: (2025)