Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic Transformations
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lindemann, Matthias, Koller, Alexander, Titov, Ivan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by Simulation
par: Lindemann, Matthias, et autres
Publié: (2023)
par: Lindemann, Matthias, et autres
Publié: (2023)
Cache & Distil: Optimising API Calls to Large Language Models
par: Ramírez, Guillem, et autres
Publié: (2023)
par: Ramírez, Guillem, et autres
Publié: (2023)
Investigating Syntactic Biases in Multilingual Transformers with RC Attachment Ambiguities in Italian and English
par: Kamerath, Michael, et autres
Publié: (2025)
par: Kamerath, Michael, et autres
Publié: (2025)
Positional Biases Shift as Inputs Approach Context Window Limits
par: Veseli, Blerta, et autres
Publié: (2025)
par: Veseli, Blerta, et autres
Publié: (2025)
Language Models Need Inductive Biases to Count Inductively
par: Chang, Yingshan, et autres
Publié: (2024)
par: Chang, Yingshan, et autres
Publié: (2024)
The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models
par: Liu, Yan, et autres
Publié: (2024)
par: Liu, Yan, et autres
Publié: (2024)
Generalisation First, Memorisation Second? Memorisation Localisation for Natural Language Classification Tasks
par: Dankers, Verna, et autres
Publié: (2024)
par: Dankers, Verna, et autres
Publié: (2024)
On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
par: Naous, Tarek, et autres
Publié: (2025)
par: Naous, Tarek, et autres
Publié: (2025)
Unveiling Divergent Inductive Biases of LLMs on Temporal Data
par: Kishore, Sindhu, et autres
Publié: (2024)
par: Kishore, Sindhu, et autres
Publié: (2024)
Tree Transformers are an Ineffective Model of Syntactic Constituency
par: Ginn, Michael
Publié: (2024)
par: Ginn, Michael
Publié: (2024)
Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at Scale
par: Hu, Xiang, et autres
Publié: (2024)
par: Hu, Xiang, et autres
Publié: (2024)
Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention
par: Zeris, Athanasios
Publié: (2026)
par: Zeris, Athanasios
Publié: (2026)
PART: Pre-trained Authorship Representation Transformer
par: Huertas-Tato, Javier, et autres
Publié: (2022)
par: Huertas-Tato, Javier, et autres
Publié: (2022)
Simple and effective data augmentation for compositional generalization
par: Yao, Yuekun, et autres
Publié: (2024)
par: Yao, Yuekun, et autres
Publié: (2024)
Fine-grained Controllable Text Generation through In-context Learning with Feedback
par: Thillainathan, Sarubi, et autres
Publié: (2024)
par: Thillainathan, Sarubi, et autres
Publié: (2024)
A Survey on Complex Tasks for Goal-Directed Interactive Agents
par: Hartmann, Mareike, et autres
Publié: (2024)
par: Hartmann, Mareike, et autres
Publié: (2024)
Predicting generalization performance with correctness discriminators
par: Yao, Yuekun, et autres
Publié: (2023)
par: Yao, Yuekun, et autres
Publié: (2023)
Unlearning Traces the Influential Training Data of Language Models
par: Isonuma, Masaru, et autres
Publié: (2024)
par: Isonuma, Masaru, et autres
Publié: (2024)
M-Wanda: Improving One-Shot Pruning for Multilingual LLMs
par: Choenni, Rochelle, et autres
Publié: (2025)
par: Choenni, Rochelle, et autres
Publié: (2025)
Grammar Assistance Using Syntactic Structures (GAUSS)
par: Zamaraeva, Olga, et autres
Publié: (2024)
par: Zamaraeva, Olga, et autres
Publié: (2024)
Language models can learn implicit multi-hop reasoning, but only if they have lots of training data
par: Yao, Yuekun, et autres
Publié: (2025)
par: Yao, Yuekun, et autres
Publié: (2025)
Menzerath-Altmann Law for Syntactic Structures in Ukrainian
par: Buk, Solomija, et autres
Publié: (2007)
par: Buk, Solomija, et autres
Publié: (2007)
Joint Localization and Activation Editing for Low-Resource Fine-Tuning
par: Lai, Wen, et autres
Publié: (2025)
par: Lai, Wen, et autres
Publié: (2025)
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
par: Martinez, Richard Diehl, et autres
Publié: (2024)
par: Martinez, Richard Diehl, et autres
Publié: (2024)
Shaping Shared Languages: Human and Large Language Models' Inductive Biases in Emergent Communication
par: Kouwenhoven, Tom, et autres
Publié: (2025)
par: Kouwenhoven, Tom, et autres
Publié: (2025)
Tree-Planted Transformers: Unidirectional Transformer Language Models with Implicit Syntactic Supervision
par: Yoshida, Ryo, et autres
Publié: (2024)
par: Yoshida, Ryo, et autres
Publié: (2024)
Frege in the Flesh: Biolinguistics and the Neural Enforcement of Syntactic Structures
par: Murphy, Elliot
Publié: (2026)
par: Murphy, Elliot
Publié: (2026)
Understanding Syntactic Generalization in Structure-inducing Language Models
par: Arps, David, et autres
Publié: (2025)
par: Arps, David, et autres
Publié: (2025)
Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection
par: Ramírez, Guillem, et autres
Publié: (2024)
par: Ramírez, Guillem, et autres
Publié: (2024)
Explanation Regularisation through the Lens of Attributions
par: Ferreira, Pedro, et autres
Publié: (2024)
par: Ferreira, Pedro, et autres
Publié: (2024)
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
par: Ferreira, Pedro, et autres
Publié: (2025)
par: Ferreira, Pedro, et autres
Publié: (2025)
On Initializing Transformers with Pre-trained Embeddings
par: Kim, Ha Young, et autres
Publié: (2024)
par: Kim, Ha Young, et autres
Publié: (2024)
An Empirical Investigation of Matrix Factorization Methods for Pre-trained Transformers
par: Gupta, Ashim, et autres
Publié: (2024)
par: Gupta, Ashim, et autres
Publié: (2024)
What's New in My Data? Novelty Exploration via Contrastive Generation
par: Isonuma, Masaru, et autres
Publié: (2024)
par: Isonuma, Masaru, et autres
Publié: (2024)
Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs
par: Boguraev, Sasha, et autres
Publié: (2026)
par: Boguraev, Sasha, et autres
Publié: (2026)
Universal Syntactic Structures: Modeling Syntax for Various Natural Languages
par: Kim, Min K., et autres
Publié: (2023)
par: Kim, Min K., et autres
Publié: (2023)
Catalytic Role Of Noise And Necessity Of Inductive Biases In The Emergence Of Compositional Communication
par: Kuciński, Łukasz, et autres
Publié: (2021)
par: Kuciński, Łukasz, et autres
Publié: (2021)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
par: Kraus, Oliver, et autres
Publié: (2026)
par: Kraus, Oliver, et autres
Publié: (2026)
A Dialogue Game for Eliciting Balanced Collaboration
par: Jeknić, Isidora, et autres
Publié: (2024)
par: Jeknić, Isidora, et autres
Publié: (2024)
LLMs syntactically adapt their language use to their conversational partner
par: Kandra, Florian, et autres
Publié: (2025)
par: Kandra, Florian, et autres
Publié: (2025)
Documents similaires
-
SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by Simulation
par: Lindemann, Matthias, et autres
Publié: (2023) -
Cache & Distil: Optimising API Calls to Large Language Models
par: Ramírez, Guillem, et autres
Publié: (2023) -
Investigating Syntactic Biases in Multilingual Transformers with RC Attachment Ambiguities in Italian and English
par: Kamerath, Michael, et autres
Publié: (2025) -
Positional Biases Shift as Inputs Approach Context Window Limits
par: Veseli, Blerta, et autres
Publié: (2025) -
Language Models Need Inductive Biases to Count Inductively
par: Chang, Yingshan, et autres
Publié: (2024)