The Impact of Depth on Compositional Generalization in Transformer Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Petty, Jackson, van Steenkiste, Sjoerd, Dasgupta, Ishita, Sha, Fei, Garrette, Dan, Linzen, Tal |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Does Code Pretraining Affect Language Model Task Performance?
by: Petty, Jackson, et al.
Published: (2024)
by: Petty, Jackson, et al.
Published: (2024)
A Systematic Comparison of Syllogistic Reasoning in Humans and Language Models
by: Eisape, Tiwalayo, et al.
Published: (2023)
by: Eisape, Tiwalayo, et al.
Published: (2023)
Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
by: Qiu, Linlu, et al.
Published: (2025)
by: Qiu, Linlu, et al.
Published: (2025)
Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction
by: Petty, Jackson, et al.
Published: (2026)
by: Petty, Jackson, et al.
Published: (2026)
In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax
by: Mueller, Aaron, et al.
Published: (2023)
by: Mueller, Aaron, et al.
Published: (2023)
Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases
by: Hu, Michael Y., et al.
Published: (2025)
by: Hu, Michael Y., et al.
Published: (2025)
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
by: Petty, Jackson, et al.
Published: (2025)
by: Petty, Jackson, et al.
Published: (2025)
Do Language Models' Words Refer?
by: Mandelkern, Matthew, et al.
Published: (2023)
by: Mandelkern, Matthew, et al.
Published: (2023)
Entailment Semantics Can Be Extracted from an Ideal Language Model
by: Merrill, William, et al.
Published: (2022)
by: Merrill, William, et al.
Published: (2022)
SPAWNing Structural Priming Predictions from a Cognitively Motivated Parser
by: Prasad, Grusha, et al.
Published: (2024)
by: Prasad, Grusha, et al.
Published: (2024)
What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length
by: Tjuatja, Lindia, et al.
Published: (2024)
by: Tjuatja, Lindia, et al.
Published: (2024)
Emergence of Linear Truth Encodings in Language Models
by: Ravfogel, Shauli, et al.
Published: (2025)
by: Ravfogel, Shauli, et al.
Published: (2025)
Multilingual Prompting for Improving LLM Generation Diversity
by: Wang, Qihan, et al.
Published: (2025)
by: Wang, Qihan, et al.
Published: (2025)
Deconstructing sentence disambiguation by joint latent modeling of reading paradigms: LLM surprisal is not enough
by: Paape, Dario, et al.
Published: (2026)
by: Paape, Dario, et al.
Published: (2026)
Manipulating language models' training data to study syntactic constraint learning: the case of English passivization
by: Leong, Cara Su-Yi, et al.
Published: (2024)
by: Leong, Cara Su-Yi, et al.
Published: (2024)
Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis
by: Timkey, William, et al.
Published: (2026)
by: Timkey, William, et al.
Published: (2026)
Language Models Struggle to Use Representations Learned In-Context
by: Lepori, Michael A., et al.
Published: (2026)
by: Lepori, Michael A., et al.
Published: (2026)
How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuning
by: Choenni, Rochelle, et al.
Published: (2023)
by: Choenni, Rochelle, et al.
Published: (2023)
The Illusion of State in State-Space Models
by: Merrill, William, et al.
Published: (2024)
by: Merrill, William, et al.
Published: (2024)
Benchmarking Vision Language Models for Cultural Understanding
by: Nayak, Shravan, et al.
Published: (2024)
by: Nayak, Shravan, et al.
Published: (2024)
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models
by: Zheng, Lin, et al.
Published: (2026)
by: Zheng, Lin, et al.
Published: (2026)
Can You Learn Semantics Through Next-Word Prediction? The Case of Entailment
by: Merrill, William, et al.
Published: (2024)
by: Merrill, William, et al.
Published: (2024)
How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?
by: Liu, Ryan, et al.
Published: (2024)
by: Liu, Ryan, et al.
Published: (2024)
Rapid Word Learning Through Meta In-Context Learning
by: Wang, Wentao, et al.
Published: (2025)
by: Wang, Wentao, et al.
Published: (2025)
Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility
by: Lepori, Michael A., et al.
Published: (2025)
by: Lepori, Michael A., et al.
Published: (2025)
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
by: Fleisig, Eve, et al.
Published: (2024)
by: Fleisig, Eve, et al.
Published: (2024)
Standard Language Ideology in AI-Generated Language
by: Smith, Genevieve, et al.
Published: (2024)
by: Smith, Genevieve, et al.
Published: (2024)
Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time
by: Hu, Michael Y., et al.
Published: (2026)
by: Hu, Michael Y., et al.
Published: (2026)
DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
by: Jiao, Cathy, et al.
Published: (2025)
by: Jiao, Cathy, et al.
Published: (2025)
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
by: Chen, Hung-Hsuan
Published: (2026)
by: Chen, Hung-Hsuan
Published: (2026)
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations
by: Deshpande, Vijeta, et al.
Published: (2025)
by: Deshpande, Vijeta, et al.
Published: (2025)
Linearity of Relation Decoding in Transformer Language Models
by: Hernandez, Evan, et al.
Published: (2023)
by: Hernandez, Evan, et al.
Published: (2023)
Private Text Generation by Seeding Large Language Model Prompts
by: Nagesh, Supriya, et al.
Published: (2025)
by: Nagesh, Supriya, et al.
Published: (2025)
Rewiring the Transformer with Depth-Wise LSTMs
by: Xu, Hongfei, et al.
Published: (2020)
by: Xu, Hongfei, et al.
Published: (2020)
Sentiment Analysis of Code-Mixed Languages leveraging Resource Rich Languages
by: Choudhary, Nurendra, et al.
Published: (2018)
by: Choudhary, Nurendra, et al.
Published: (2018)
Compositional Generalization with Grounded Language Models
by: Wold, Sondre, et al.
Published: (2024)
by: Wold, Sondre, et al.
Published: (2024)
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
by: Blum, Carter, et al.
Published: (2025)
by: Blum, Carter, et al.
Published: (2025)
'Since Lawyers are Males..': Examining Implicit Gender Bias in Hindi Language Generation by LLMs
by: Joshi, Ishika, et al.
Published: (2024)
by: Joshi, Ishika, et al.
Published: (2024)
Similar Items
-
How Does Code Pretraining Affect Language Model Task Performance?
by: Petty, Jackson, et al.
Published: (2024) -
A Systematic Comparison of Syllogistic Reasoning in Humans and Language Models
by: Eisape, Tiwalayo, et al.
Published: (2023) -
Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
by: Qiu, Linlu, et al.
Published: (2025) -
Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction
by: Petty, Jackson, et al.
Published: (2026) -
In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax
by: Mueller, Aaron, et al.
Published: (2023)