When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
Fuente:
arXiv
Guardado en:
| Autores principales: | Higuchi, Rei, Kawata, Ryotaro, Nishikawa, Naoki, Oko, Kazusato, Yamaguchi, Shoichiro, Kobayashi, Sosuke, Tokui, Seiya, Hayashi, Kohei, Okanohara, Daisuke, Suzuki, Taiji |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Direct Distributional Optimization for Provable Alignment of Diffusion Models
por: Kawata, Ryotaro, et al.
Publicado: (2025)
por: Kawata, Ryotaro, et al.
Publicado: (2025)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
por: Nishikawa, Naoki, et al.
Publicado: (2025)
por: Nishikawa, Naoki, et al.
Publicado: (2025)
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
por: Higuchi, Rei, et al.
Publicado: (2026)
por: Higuchi, Rei, et al.
Publicado: (2026)
Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations
por: Oko, Kazusato, et al.
Publicado: (2024)
por: Oko, Kazusato, et al.
Publicado: (2024)
Pretrained transformer efficiently learns low-dimensional target functions in-context
por: Oko, Kazusato, et al.
Publicado: (2024)
por: Oko, Kazusato, et al.
Publicado: (2024)
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
por: Kawata, Ryotaro, et al.
Publicado: (2025)
por: Kawata, Ryotaro, et al.
Publicado: (2025)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
por: Kawata, Ryotaro, et al.
Publicado: (2026)
por: Kawata, Ryotaro, et al.
Publicado: (2026)
Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
por: Lee, Jason D., et al.
Publicado: (2024)
por: Lee, Jason D., et al.
Publicado: (2024)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
por: Higuchi, Rei, et al.
Publicado: (2025)
por: Higuchi, Rei, et al.
Publicado: (2025)
Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
por: Kim, Juno, et al.
Publicado: (2023)
por: Kim, Juno, et al.
Publicado: (2023)
Flow matching achieves almost minimax optimal convergence
por: Fukumizu, Kenji, et al.
Publicado: (2024)
por: Fukumizu, Kenji, et al.
Publicado: (2024)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
por: Nishikawa, Naoki, et al.
Publicado: (2024)
por: Nishikawa, Naoki, et al.
Publicado: (2024)
From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
por: Kawata, Ryotaro, et al.
Publicado: (2025)
por: Kawata, Ryotaro, et al.
Publicado: (2025)
A Statistical Theory of Contrastive Pre-training and Multimodal Generative AI
por: Oko, Kazusato, et al.
Publicado: (2025)
por: Oko, Kazusato, et al.
Publicado: (2025)
Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO
por: Takakura, Shokichi, et al.
Publicado: (2026)
por: Takakura, Shokichi, et al.
Publicado: (2026)
Speed-accuracy relations for diffusion models: Wisdom from nonequilibrium thermodynamics and optimal transport
por: Ikeda, Kotaro, et al.
Publicado: (2024)
por: Ikeda, Kotaro, et al.
Publicado: (2024)
A Thermodynamic Theory of Learning Part II: Critical Period Closure and Continual Learning Failure
por: Okanohara, Daisuke
Publicado: (2026)
por: Okanohara, Daisuke
Publicado: (2026)
A Thermodynamic Theory of Learning I: Irreversible Ensemble Transport and Epistemic Costs
por: Okanohara, Daisuke
Publicado: (2026)
por: Okanohara, Daisuke
Publicado: (2026)
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
por: Wachi, Akifumi, et al.
Publicado: (2026)
por: Wachi, Akifumi, et al.
Publicado: (2026)
ROFBS$α$: Real Time Backup System Decoupled from ML Based Ransomware Detection
por: Higuchi, Kosuke, et al.
Publicado: (2025)
por: Higuchi, Kosuke, et al.
Publicado: (2025)
Impact of File-Open Hook Points on Backup Ratio in ROFBS on XFS
por: Higuchi, Kosuke, et al.
Publicado: (2026)
por: Higuchi, Kosuke, et al.
Publicado: (2026)
Spike No More: Stabilizing the Pre-training of Large Language Models
por: Takase, Sho, et al.
Publicado: (2023)
por: Takase, Sho, et al.
Publicado: (2023)
Generative Model for Constructing Reaction Path from Initial to Final States
por: Hayashi, Akihide, et al.
Publicado: (2024)
por: Hayashi, Akihide, et al.
Publicado: (2024)
Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning
por: Yano, Kazuki, et al.
Publicado: (2026)
por: Yano, Kazuki, et al.
Publicado: (2026)
Multi-Robot Patrol Algorithm with Distributed Coordination and Consciousness of the Base Station's Situation Awareness
por: Kobayashi, Kazuho, et al.
Publicado: (2023)
por: Kobayashi, Kazuho, et al.
Publicado: (2023)
Distributed Algorithm with Emergent Area Partitioning and Base Station's Situation Awareness for Multi-Robot Patrolling
por: Kobayashi, Kazuho, et al.
Publicado: (2026)
por: Kobayashi, Kazuho, et al.
Publicado: (2026)
ExSampling: a system for the real-time ensemble performance of field-recorded environmental sounds
por: Kobayashi, Atsuya, et al.
Publicado: (2020)
por: Kobayashi, Atsuya, et al.
Publicado: (2020)
Two applications of stochastic thermodynamics to hydrodynamics
por: Yoshimura, Kohei, et al.
Publicado: (2023)
por: Yoshimura, Kohei, et al.
Publicado: (2023)
Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models
por: Suzuki, Satoshi, et al.
Publicado: (2025)
por: Suzuki, Satoshi, et al.
Publicado: (2025)
The Infinite Dyson Brownian Motion with $β=2$ Does Not Have a Spectral Gap
por: Suzuki, Kohei
Publicado: (2024)
por: Suzuki, Kohei
Publicado: (2024)
Walking inhibition effect of different surfaces in apple snail, Pomacea canaliculata
por: Satoru Tachibana, et al.
Publicado: (2026)
por: Satoru Tachibana, et al.
Publicado: (2026)
Unveiling multilayered barriers to agile methodologies: an exploratory study on relationships among barriers
por: Karen Kawata Kobayashi
Publicado: (2025)
por: Karen Kawata Kobayashi
Publicado: (2025)
Infinite variety of thermodynamic speed limits with general activities
por: Nagayama, Ryuna, et al.
Publicado: (2024)
por: Nagayama, Ryuna, et al.
Publicado: (2024)
Efficient Construction of Model Family through Progressive Training Using Model Expansion
por: Yano, Kazuki, et al.
Publicado: (2025)
por: Yano, Kazuki, et al.
Publicado: (2025)
When Do We Feel Sorry for Others? : An Externality of Lake Use as an Example
por: Kawata, Y.
Publicado: (2015)
por: Kawata, Y.
Publicado: (2015)
Does High Unemployment Rate Result in a High Divorce Rate?: A Test for Japan
por: Yukichika Kawata
Publicado: (2008)
por: Yukichika Kawata
Publicado: (2008)
Latent Granular Resynthesis using Neural Audio Codecs
por: Tokui, Nao, et al.
Publicado: (2025)
por: Tokui, Nao, et al.
Publicado: (2025)
An energy landscape-based theoretical framework for understanding the emergence of functions in a living system under the dynamical component interaction
por: Suzuki, Ryunosuke, et al.
Publicado: (2025)
por: Suzuki, Ryunosuke, et al.
Publicado: (2025)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
por: Kim, Juno, et al.
Publicado: (2024)
por: Kim, Juno, et al.
Publicado: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
por: Watanabe, Chihiro, et al.
Publicado: (2021)
por: Watanabe, Chihiro, et al.
Publicado: (2021)
Ejemplares similares
-
Direct Distributional Optimization for Provable Alignment of Diffusion Models
por: Kawata, Ryotaro, et al.
Publicado: (2025) -
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
por: Nishikawa, Naoki, et al.
Publicado: (2025) -
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
por: Higuchi, Rei, et al.
Publicado: (2026) -
Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations
por: Oko, Kazusato, et al.
Publicado: (2024) -
Pretrained transformer efficiently learns low-dimensional target functions in-context
por: Oko, Kazusato, et al.
Publicado: (2024)