Disaggregation Reveals Hidden Training Dynamics: The Case of Agreement Attraction
Fuente:
arXiv
Saved in:
| Main Authors: | Michaelov, James A., Arnett, Catherine |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revenge of the Fallen? Recurrent Models Match Transformers at Predicting Human Language Comprehension Metrics
by: Michaelov, James A., et al.
Published: (2024)
by: Michaelov, James A., et al.
Published: (2024)
On the Acquisition of Shared Grammatical Representations in Bilingual Language Models
by: Arnett, Catherine, et al.
Published: (2025)
by: Arnett, Catherine, et al.
Published: (2025)
Different Tokenization Schemes Lead to Comparable Performance in Spanish Number Agreement
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
by: Michaelov, James A., et al.
Published: (2025)
by: Michaelov, James A., et al.
Published: (2025)
Emergent inabilities? Inverse scaling over the course of pretraining
by: Michaelov, James A., et al.
Published: (2023)
by: Michaelov, James A., et al.
Published: (2023)
N-gram-like Language Models Predict Reading Time Best
by: Michaelov, James A., et al.
Published: (2026)
by: Michaelov, James A., et al.
Published: (2026)
BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization
by: Land, Sander, et al.
Published: (2025)
by: Land, Sander, et al.
Published: (2025)
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
by: Chizhov, Pavel, et al.
Published: (2024)
by: Chizhov, Pavel, et al.
Published: (2024)
Diverging Transformer Predictions for Human Sentence Processing: A Comprehensive Analysis of Agreement Attraction Effects
by: von der Malsburg, Titus, et al.
Published: (2026)
by: von der Malsburg, Titus, et al.
Published: (2026)
Why do language models perform worse for morphologically complex languages?
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
Toxicity of the Commons: Curating Open-Source Pre-Training Data
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events
by: Michaelov, James A., et al.
Published: (2025)
by: Michaelov, James A., et al.
Published: (2025)
How Open Must Language Models be to Enable Reliable Scientific Inference?
by: Michaelov, James A., et al.
Published: (2026)
by: Michaelov, James A., et al.
Published: (2026)
A Bit of a Problem: Measurement Disparities in Dataset Sizes Across Languages
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
Evaluating Morphological Alignment of Tokenizers in 70 Languages
by: Arnett, Catherine, et al.
Published: (2025)
by: Arnett, Catherine, et al.
Published: (2025)
Weight Tying Biases Token Embeddings Towards the Output Space
by: Lopardo, Antonio, et al.
Published: (2026)
by: Lopardo, Antonio, et al.
Published: (2026)
Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMs
by: Trott, Sean, et al.
Published: (2026)
by: Trott, Sean, et al.
Published: (2026)
Explaining and Mitigating Crosslingual Tokenizer Inequities
by: Arnett, Catherine, et al.
Published: (2025)
by: Arnett, Catherine, et al.
Published: (2025)
Goldfish: Monolingual Language Models for 350 Languages
by: Chang, Tyler A., et al.
Published: (2024)
by: Chang, Tyler A., et al.
Published: (2024)
Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
by: Langlais, Pierre-Carl, et al.
Published: (2025)
by: Langlais, Pierre-Carl, et al.
Published: (2025)
Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Re-defining Humor Data Objects for AI Humor Research
by: Arnett, Anna, et al.
Published: (2026)
by: Arnett, Anna, et al.
Published: (2026)
Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation
by: James, Joseph
Published: (2026)
by: James, Joseph
Published: (2026)
Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models
by: Karkada, Dhruva, et al.
Published: (2025)
by: Karkada, Dhruva, et al.
Published: (2025)
On the Similarity of Circuits across Languages: a Case Study on the Subject-verb Agreement Task
by: Ferrando, Javier, et al.
Published: (2024)
by: Ferrando, Javier, et al.
Published: (2024)
Jailbreaking Large Language Diffusion Models: Revealing Hidden Safety Flaws in Diffusion-Based Text Generation
by: Zhang, Yuanhe, et al.
Published: (2025)
by: Zhang, Yuanhe, et al.
Published: (2025)
RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
Saying the Unsaid: Revealing the Hidden Language of Multimodal Systems Through Telephone Games
by: Zhao, Juntu, et al.
Published: (2025)
by: Zhao, Juntu, et al.
Published: (2025)
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning
by: Ye, Xiaotian, et al.
Published: (2026)
by: Ye, Xiaotian, et al.
Published: (2026)
Estimating Agreement by Chance for Sequence Annotation
by: Li, Diya, et al.
Published: (2024)
by: Li, Diya, et al.
Published: (2024)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
by: Turk, Matt
Published: (2026)
by: Turk, Matt
Published: (2026)
The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?
by: Zhao, Qinyu, et al.
Published: (2024)
by: Zhao, Qinyu, et al.
Published: (2024)
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
by: Li, Mingjie, et al.
Published: (2026)
by: Li, Mingjie, et al.
Published: (2026)
Can Large Language Models Detect Verbal Indicators of Romantic Attraction?
by: Matz, Sandra C., et al.
Published: (2024)
by: Matz, Sandra C., et al.
Published: (2024)
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
by: Potter, Yujin, et al.
Published: (2024)
by: Potter, Yujin, et al.
Published: (2024)
Explanation Bias is a Product: Revealing the Hidden Lexical and Position Preferences in Post-Hoc Feature Attribution
by: Kamp, Jonathan, et al.
Published: (2025)
by: Kamp, Jonathan, et al.
Published: (2025)
Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs
by: Blandfort, Phil, et al.
Published: (2026)
by: Blandfort, Phil, et al.
Published: (2026)
Revealing the Inherent Instructability of Pre-Trained Language Models
by: An, Seokhyun, et al.
Published: (2024)
by: An, Seokhyun, et al.
Published: (2024)
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs
by: Berezin, Sergey, et al.
Published: (2025)
by: Berezin, Sergey, et al.
Published: (2025)
Similar Items
-
Revenge of the Fallen? Recurrent Models Match Transformers at Predicting Human Language Comprehension Metrics
by: Michaelov, James A., et al.
Published: (2024) -
On the Acquisition of Shared Grammatical Representations in Bilingual Language Models
by: Arnett, Catherine, et al.
Published: (2025) -
Different Tokenization Schemes Lead to Comparable Performance in Spanish Number Agreement
by: Arnett, Catherine, et al.
Published: (2024) -
Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
by: Michaelov, James A., et al.
Published: (2025) -
Emergent inabilities? Inverse scaling over the course of pretraining
by: Michaelov, James A., et al.
Published: (2023)