Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Michael Y., Petty, Jackson, Shi, Chuan, Merrill, William, Linzen, Tal |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction
por: Petty, Jackson, et al.
Publicado: (2026)
por: Petty, Jackson, et al.
Publicado: (2026)
How Does Code Pretraining Affect Language Model Task Performance?
por: Petty, Jackson, et al.
Publicado: (2024)
por: Petty, Jackson, et al.
Publicado: (2024)
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
por: Petty, Jackson, et al.
Publicado: (2025)
por: Petty, Jackson, et al.
Publicado: (2025)
Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time
por: Hu, Michael Y., et al.
Publicado: (2026)
por: Hu, Michael Y., et al.
Publicado: (2026)
Rapid Word Learning Through Meta In-Context Learning
por: Wang, Wentao, et al.
Publicado: (2025)
por: Wang, Wentao, et al.
Publicado: (2025)
Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy
por: Dong, Yihong, et al.
Publicado: (2026)
por: Dong, Yihong, et al.
Publicado: (2026)
A Systematic Comparison of Syllogistic Reasoning in Humans and Language Models
por: Eisape, Tiwalayo, et al.
Publicado: (2023)
por: Eisape, Tiwalayo, et al.
Publicado: (2023)
Language Models Struggle to Use Representations Learned In-Context
por: Lepori, Michael A., et al.
Publicado: (2026)
por: Lepori, Michael A., et al.
Publicado: (2026)
Entailment Semantics Can Be Extracted from an Ideal Language Model
por: Merrill, William, et al.
Publicado: (2022)
por: Merrill, William, et al.
Publicado: (2022)
The Illusion of State in State-Space Models
por: Merrill, William, et al.
Publicado: (2024)
por: Merrill, William, et al.
Publicado: (2024)
Heterogeneity in Formal Linguistic Competence of Language Models: Is Data the Real Bottleneck?
por: Renduchintala, H S V N S Kowndinya, et al.
Publicado: (2026)
por: Renduchintala, H S V N S Kowndinya, et al.
Publicado: (2026)
Synthetic continued pretraining
por: Yang, Zitong, et al.
Publicado: (2024)
por: Yang, Zitong, et al.
Publicado: (2024)
In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax
por: Mueller, Aaron, et al.
Publicado: (2023)
por: Mueller, Aaron, et al.
Publicado: (2023)
Relative Value Biases in Large Language Models
por: Hayes, William M., et al.
Publicado: (2024)
por: Hayes, William M., et al.
Publicado: (2024)
Large Language Models are Biased Reinforcement Learners
por: Hayes, William M., et al.
Publicado: (2024)
por: Hayes, William M., et al.
Publicado: (2024)
Do pretrained Transformers Learn In-Context by Gradient Descent?
por: Shen, Lingfeng, et al.
Publicado: (2023)
por: Shen, Lingfeng, et al.
Publicado: (2023)
Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
por: Qiu, Linlu, et al.
Publicado: (2025)
por: Qiu, Linlu, et al.
Publicado: (2025)
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
por: Huang, Xinting, et al.
Publicado: (2026)
por: Huang, Xinting, et al.
Publicado: (2026)
Uncovering Biases with Reflective Large Language Models
por: Chang, Edward Y.
Publicado: (2024)
por: Chang, Edward Y.
Publicado: (2024)
Self-Improving Pretraining: using post-trained models to pretrain better models
por: Tan, Ellen Xiaoqing, et al.
Publicado: (2026)
por: Tan, Ellen Xiaoqing, et al.
Publicado: (2026)
A Formal Comparison Between Chain of Thought and Latent Thought
por: Xu, Kevin, et al.
Publicado: (2025)
por: Xu, Kevin, et al.
Publicado: (2025)
An Investigation of Linguistic Biases in LLM-Based Recommendations
por: Venkateswaran, Nitin, et al.
Publicado: (2026)
por: Venkateswaran, Nitin, et al.
Publicado: (2026)
Perceptions of Linguistic Uncertainty by Language Models and Humans
por: Belem, Catarina G, et al.
Publicado: (2024)
por: Belem, Catarina G, et al.
Publicado: (2024)
Linguistic Blind Spots of Large Language Models
por: Cheng, Jiali, et al.
Publicado: (2025)
por: Cheng, Jiali, et al.
Publicado: (2025)
MedicalBERT: enhancing biomedical natural language processing using pretrained BERT-based model
por: Reddy, K. Sahit, et al.
Publicado: (2025)
por: Reddy, K. Sahit, et al.
Publicado: (2025)
The Impact of Depth on Compositional Generalization in Transformer Language Models
por: Petty, Jackson, et al.
Publicado: (2023)
por: Petty, Jackson, et al.
Publicado: (2023)
Synthetic bootstrapped pretraining
por: Yang, Zitong, et al.
Publicado: (2025)
por: Yang, Zitong, et al.
Publicado: (2025)
Do Large Language Models Show Biases in Causal Learning?
por: Carro, Maria Victoria, et al.
Publicado: (2024)
por: Carro, Maria Victoria, et al.
Publicado: (2024)
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
por: Yang, Xikang, et al.
Publicado: (2025)
por: Yang, Xikang, et al.
Publicado: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
por: Lan, Michael, et al.
Publicado: (2023)
por: Lan, Michael, et al.
Publicado: (2023)
Evaluating $n$-Gram Novelty of Language Models Using Rusty-DAWG
por: Merrill, William, et al.
Publicado: (2024)
por: Merrill, William, et al.
Publicado: (2024)
Formal-LLM: Integrating Formal Language and Natural Language for Controllable LLM-based Agents
por: Li, Zelong, et al.
Publicado: (2024)
por: Li, Zelong, et al.
Publicado: (2024)
Transformer Circuit Faithfulness Metrics are not Robust
por: Miller, Joseph, et al.
Publicado: (2024)
por: Miller, Joseph, et al.
Publicado: (2024)
Large Language Models are Geographically Biased
por: Manvi, Rohin, et al.
Publicado: (2024)
por: Manvi, Rohin, et al.
Publicado: (2024)
Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning
por: Dijujin, Negin Hashemi, et al.
Publicado: (2025)
por: Dijujin, Negin Hashemi, et al.
Publicado: (2025)
Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
por: Yang, Nakyeong, et al.
Publicado: (2023)
por: Yang, Nakyeong, et al.
Publicado: (2023)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
por: Lee, Isack, et al.
Publicado: (2024)
por: Lee, Isack, et al.
Publicado: (2024)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
por: Svete, Anej, et al.
Publicado: (2026)
por: Svete, Anej, et al.
Publicado: (2026)
BPDec: Unveiling the Potential of Masked Language Modeling Decoder in BERT pretraining
por: Liang, Wen, et al.
Publicado: (2024)
por: Liang, Wen, et al.
Publicado: (2024)
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning
por: Merrill, Scott, et al.
Publicado: (2026)
por: Merrill, Scott, et al.
Publicado: (2026)
Ejemplares similares
-
Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction
por: Petty, Jackson, et al.
Publicado: (2026) -
How Does Code Pretraining Affect Language Model Task Performance?
por: Petty, Jackson, et al.
Publicado: (2024) -
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
por: Petty, Jackson, et al.
Publicado: (2025) -
Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time
por: Hu, Michael Y., et al.
Publicado: (2026) -
Rapid Word Learning Through Meta In-Context Learning
por: Wang, Wentao, et al.
Publicado: (2025)