Tree-Planted Transformers: Unidirectional Transformer Language Models with Implicit Syntactic Supervision
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yoshida, Ryo, Someya, Taiga, Oseki, Yohei |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Derivational Probing: Unveiling the Layer-wise Derivation of Syntactic Structures in Neural Language Models
par: Someya, Taiga, et autres
Publié: (2025)
par: Someya, Taiga, et autres
Publié: (2025)
Language Acquisition Device in Large Language Models
par: Mita, Masato, et autres
Publié: (2026)
par: Mita, Masato, et autres
Publié: (2026)
An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal
par: Yoshida, Ryo, et autres
Publié: (2026)
par: Yoshida, Ryo, et autres
Publié: (2026)
If Attention Serves as a Cognitive Model of Human Memory Retrieval, What is the Plausible Memory Representation?
par: Yoshida, Ryo, et autres
Publié: (2025)
par: Yoshida, Ryo, et autres
Publié: (2025)
Composition, Attention, or Both?
par: Yoshida, Ryo, et autres
Publié: (2022)
par: Yoshida, Ryo, et autres
Publié: (2022)
Developmentally-plausible Working Memory Shapes a Critical Period for Language Acquisition
par: Mita, Masato, et autres
Publié: (2025)
par: Mita, Masato, et autres
Publié: (2025)
Modeling Human Sentence Processing with Left-Corner Recurrent Neural Network Grammars
par: Yoshida, Ryo, et autres
Publié: (2021)
par: Yoshida, Ryo, et autres
Publié: (2021)
Emergent Word Order Universals from Cognitively-Motivated Language Models
par: Kuribayashi, Tatsuki, et autres
Publié: (2024)
par: Kuribayashi, Tatsuki, et autres
Publié: (2024)
Rethinking the Relationship between the Power Law and Hierarchical Structures
par: Nakaishi, Kai, et autres
Publié: (2025)
par: Nakaishi, Kai, et autres
Publié: (2025)
Psychometric Predictive Power of Large Language Models
par: Kuribayashi, Tatsuki, et autres
Publié: (2023)
par: Kuribayashi, Tatsuki, et autres
Publié: (2023)
Tree Transformers are an Ineffective Model of Syntactic Constituency
par: Ginn, Michael
Publié: (2024)
par: Ginn, Michael
Publié: (2024)
Dual Alignment Between Language Model Layers and Human Sentence Processing
par: Kuribayashi, Tatsuki, et autres
Publié: (2026)
par: Kuribayashi, Tatsuki, et autres
Publié: (2026)
Is Structure Dependence Shaped for Efficient Communication?: A Case Study on Coordination
par: Kajikawa, Kohei, et autres
Publié: (2024)
par: Kajikawa, Kohei, et autres
Publié: (2024)
Can Language Models Learn Typologically Implausible Languages?
par: Xu, Tianyang, et autres
Publié: (2025)
par: Xu, Tianyang, et autres
Publié: (2025)
Large Language Models Are Human-Like Internally
par: Kuribayashi, Tatsuki, et autres
Publié: (2025)
par: Kuribayashi, Tatsuki, et autres
Publié: (2025)
BabyLM Challenge: Exploring the Effect of Variation Sets on Language Model Training Efficiency
par: Haga, Akari, et autres
Publié: (2024)
par: Haga, Akari, et autres
Publié: (2024)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
par: Lee, Tzu-Yun, et autres
Publié: (2025)
par: Lee, Tzu-Yun, et autres
Publié: (2025)
A Systematic Study of Compositional Syntactic Transformer Language Models
par: Zhao, Yida, et autres
Publié: (2025)
par: Zhao, Yida, et autres
Publié: (2025)
Syntactic Learnability of Echo State Neural Language Models at Scale
par: Ueda, Ryo, et autres
Publié: (2025)
par: Ueda, Ryo, et autres
Publié: (2025)
Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality
par: Harada, Yuto, et autres
Publié: (2025)
par: Harada, Yuto, et autres
Publié: (2025)
Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at Scale
par: Hu, Xiang, et autres
Publié: (2024)
par: Hu, Xiang, et autres
Publié: (2024)
Information Locality as an Inductive Bias for Neural Language Models
par: Someya, Taiga, et autres
Publié: (2025)
par: Someya, Taiga, et autres
Publié: (2025)
Can Language Models Induce Grammatical Knowledge from Indirect Evidence?
par: Oba, Miyu, et autres
Publié: (2024)
par: Oba, Miyu, et autres
Publié: (2024)
The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models
par: Graichen, Nora, et autres
Publié: (2026)
par: Graichen, Nora, et autres
Publié: (2026)
Distantly Supervised Morpho-Syntactic Model for Relation Extraction
par: Gutehrlé, Nicolas, et autres
Publié: (2024)
par: Gutehrlé, Nicolas, et autres
Publié: (2024)
Exclusive Unlearning
par: Sasaki, Mutsumi, et autres
Publié: (2026)
par: Sasaki, Mutsumi, et autres
Publié: (2026)
Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic Transformations
par: Lindemann, Matthias, et autres
Publié: (2024)
par: Lindemann, Matthias, et autres
Publié: (2024)
Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs
par: Boguraev, Sasha, et autres
Publié: (2026)
par: Boguraev, Sasha, et autres
Publié: (2026)
Making the Most of your Model: Methods for Finetuning and Applying Pretrained Transformers
par: Yoshida, Davis
Publié: (2024)
par: Yoshida, Davis
Publié: (2024)
Investigating Syntactic Biases in Multilingual Transformers with RC Attachment Ambiguities in Italian and English
par: Kamerath, Michael, et autres
Publié: (2025)
par: Kamerath, Michael, et autres
Publié: (2025)
Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models
par: Kryvosheieva, Daria, et autres
Publié: (2024)
par: Kryvosheieva, Daria, et autres
Publié: (2024)
Understanding Syntactic Generalization in Structure-inducing Language Models
par: Arps, David, et autres
Publié: (2025)
par: Arps, David, et autres
Publié: (2025)
Towards Harnessing Large Language Models for Comprehension of Conversational Grounding
par: Jokinen, Kristiina, et autres
Publié: (2024)
par: Jokinen, Kristiina, et autres
Publié: (2024)
Implicit Reasoning in Transformers is Reasoning through Shortcuts
par: Lin, Tianhe, et autres
Publié: (2025)
par: Lin, Tianhe, et autres
Publié: (2025)
Targeted Syntactic Evaluation of Language Models on Georgian Case Alignment
par: Gallagher, Daniel, et autres
Publié: (2026)
par: Gallagher, Daniel, et autres
Publié: (2026)
Universal Syntactic Structures: Modeling Syntax for Various Natural Languages
par: Kim, Min K., et autres
Publié: (2023)
par: Kim, Min K., et autres
Publié: (2023)
Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models
par: Kumon, Ryoma, et autres
Publié: (2026)
par: Kumon, Ryoma, et autres
Publié: (2026)
Syntactic Evolution in Language Usage
par: Kumar, Surbhit
Publié: (2025)
par: Kumar, Surbhit
Publié: (2025)
Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance
par: Ozaki, Shintaro, et autres
Publié: (2025)
par: Ozaki, Shintaro, et autres
Publié: (2025)
Sneaking Syntax into Transformer Language Models with Tree Regularization
par: Nandi, Ananjan, et autres
Publié: (2024)
par: Nandi, Ananjan, et autres
Publié: (2024)
Documents similaires
-
Derivational Probing: Unveiling the Layer-wise Derivation of Syntactic Structures in Neural Language Models
par: Someya, Taiga, et autres
Publié: (2025) -
Language Acquisition Device in Large Language Models
par: Mita, Masato, et autres
Publié: (2026) -
An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal
par: Yoshida, Ryo, et autres
Publié: (2026) -
If Attention Serves as a Cognitive Model of Human Memory Retrieval, What is the Plausible Memory Representation?
par: Yoshida, Ryo, et autres
Publié: (2025) -
Composition, Attention, or Both?
par: Yoshida, Ryo, et autres
Publié: (2022)