Frequency Explains the Inverse Correlation of Large Language Models' Size, Training Data Amount, and Surprisal's Fit to Reading Times
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Oh, Byung-Doh, Yue, Shisen, Schuler, William |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage
von: Oh, Byung-Doh, et al.
Veröffentlicht: (2025)
von: Oh, Byung-Doh, et al.
Veröffentlicht: (2025)
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal
von: Oh, Byung-Doh, et al.
Veröffentlicht: (2024)
von: Oh, Byung-Doh, et al.
Veröffentlicht: (2024)
Linear Recency Bias During Training Improves Transformers' Fit to Reading Times
von: Clark, Christian, et al.
Veröffentlicht: (2024)
von: Clark, Christian, et al.
Veröffentlicht: (2024)
Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word Probabilities
von: Oh, Byung-Doh, et al.
Veröffentlicht: (2024)
von: Oh, Byung-Doh, et al.
Veröffentlicht: (2024)
Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal
von: Nair, Sathvik, et al.
Veröffentlicht: (2026)
von: Nair, Sathvik, et al.
Veröffentlicht: (2026)
How Well Does First-Token Entropy Approximate Word Entropy as a Psycholinguistic Predictor?
von: Clark, Christian, et al.
Veröffentlicht: (2025)
von: Clark, Christian, et al.
Veröffentlicht: (2025)
Adaptive Pre-training Data Detection for Large Language Models via Surprising Tokens
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
SR-TTT: Surprisal-Aware Residual Test-Time Training
von: P, Swamynathan V
Veröffentlicht: (2026)
von: P, Swamynathan V
Veröffentlicht: (2026)
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
von: Akyürek, Ekin, et al.
Veröffentlicht: (2024)
von: Akyürek, Ekin, et al.
Veröffentlicht: (2024)
Training Language Models to Explain Their Own Computations
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Test-Time Training on Nearest Neighbors for Large Language Models
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
von: Hardt, Moritz, et al.
Veröffentlicht: (2023)
Explaining Large Language Models with gSMILE
von: Dehghani, Zeinab, et al.
Veröffentlicht: (2025)
von: Dehghani, Zeinab, et al.
Veröffentlicht: (2025)
Scaling Law for Language Models Training Considering Batch Size
von: Shuai, Xian, et al.
Veröffentlicht: (2024)
von: Shuai, Xian, et al.
Veröffentlicht: (2024)
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
On The Truthfulness of 'Surprisingly Likely' Responses of Large Language Models
von: Goel, Naman
Veröffentlicht: (2023)
von: Goel, Naman
Veröffentlicht: (2023)
No One Size Fits All: QueryBandits for Hallucination Mitigation
von: Cho, Nicole, et al.
Veröffentlicht: (2026)
von: Cho, Nicole, et al.
Veröffentlicht: (2026)
Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Chien, et al.
Veröffentlicht: (2025)
Sequence-Level Leakage Risk of Training Data in Large Language Models
von: Tiwari, Trishita, et al.
Veröffentlicht: (2024)
von: Tiwari, Trishita, et al.
Veröffentlicht: (2024)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
von: Sun, Hao, et al.
Veröffentlicht: (2025)
von: Sun, Hao, et al.
Veröffentlicht: (2025)
JoPA:Explaining Large Language Model's Generation via Joint Prompt Attribution
von: Chang, Yurui, et al.
Veröffentlicht: (2024)
von: Chang, Yurui, et al.
Veröffentlicht: (2024)
Regurgitative Training: The Value of Real Data in Training Large Language Models
von: Zhang, Jinghui, et al.
Veröffentlicht: (2024)
von: Zhang, Jinghui, et al.
Veröffentlicht: (2024)
Explingo: Explaining AI Predictions using Large Language Models
von: Zytek, Alexandra, et al.
Veröffentlicht: (2024)
von: Zytek, Alexandra, et al.
Veröffentlicht: (2024)
Explaining Large Language Models Decisions Using Shapley Values
von: Mohammadi, Behnam
Veröffentlicht: (2024)
von: Mohammadi, Behnam
Veröffentlicht: (2024)
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
von: Liang, Hao, et al.
Veröffentlicht: (2026)
von: Liang, Hao, et al.
Veröffentlicht: (2026)
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
von: Ye, Jiayuan, et al.
Veröffentlicht: (2026)
von: Ye, Jiayuan, et al.
Veröffentlicht: (2026)
Detecting Training Data of Large Language Models via Expectation Maximization
von: Kim, Gyuwan, et al.
Veröffentlicht: (2024)
von: Kim, Gyuwan, et al.
Veröffentlicht: (2024)
Exploring Cross-Client Memorization of Training Data in Large Language Models for Federated Learning
von: Udsa, Tinnakit, et al.
Veröffentlicht: (2025)
von: Udsa, Tinnakit, et al.
Veröffentlicht: (2025)
Explaining the role of Intrinsic Dimensionality in Adversarial Training
von: Altinisik, Enes, et al.
Veröffentlicht: (2024)
von: Altinisik, Enes, et al.
Veröffentlicht: (2024)
Federated Learning with Layer Skipping: Efficient Training of Large Language Models for Healthcare NLP
von: Zhang, Lihong, et al.
Veröffentlicht: (2025)
von: Zhang, Lihong, et al.
Veröffentlicht: (2025)
One Size Fits None: Heuristic Collapse in LLM Investment Advice
von: Ross, Jillian, et al.
Veröffentlicht: (2026)
von: Ross, Jillian, et al.
Veröffentlicht: (2026)
Query-Conditioned Test-Time Self-Training for Large Language Models
von: Song, Chaehee, et al.
Veröffentlicht: (2026)
von: Song, Chaehee, et al.
Veröffentlicht: (2026)
Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models
von: Zhang, Jingyang, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyang, et al.
Veröffentlicht: (2024)
Comparative Performance Evaluation of Large Language Models for Extracting Molecular Interactions and Pathway Knowledge
von: Park, Gilchan, et al.
Veröffentlicht: (2023)
von: Park, Gilchan, et al.
Veröffentlicht: (2023)
Escaping Collapse: The Strength of Weak Data for Large Language Model Training
von: Amin, Kareem, et al.
Veröffentlicht: (2025)
von: Amin, Kareem, et al.
Veröffentlicht: (2025)
Fairness Definitions in Language Models Explained
von: Yin, Zhipeng, et al.
Veröffentlicht: (2024)
von: Yin, Zhipeng, et al.
Veröffentlicht: (2024)
Self-Training Large Language Models with Confident Reasoning
von: Jang, Hyosoon, et al.
Veröffentlicht: (2025)
von: Jang, Hyosoon, et al.
Veröffentlicht: (2025)
Linear Dynamics in the RLVR Training of Large Language Models
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
Large Language Models as Interpolated and Extrapolated Event Predictors
von: Zhang, Libo, et al.
Veröffentlicht: (2024)
von: Zhang, Libo, et al.
Veröffentlicht: (2024)
Explaining Graph Neural Networks with Large Language Models: A Counterfactual Perspective for Molecular Property Prediction
von: He, Yinhan, et al.
Veröffentlicht: (2024)
von: He, Yinhan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage
von: Oh, Byung-Doh, et al.
Veröffentlicht: (2025) -
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal
von: Oh, Byung-Doh, et al.
Veröffentlicht: (2024) -
Linear Recency Bias During Training Improves Transformers' Fit to Reading Times
von: Clark, Christian, et al.
Veröffentlicht: (2024) -
Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word Probabilities
von: Oh, Byung-Doh, et al.
Veröffentlicht: (2024) -
Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal
von: Nair, Sathvik, et al.
Veröffentlicht: (2026)