Linear Recency Bias During Training Improves Transformers' Fit to Reading Times
Fuente:
arXiv
Salvato in:
| Autori principali: | Clark, Christian, Oh, Byung-Doh, Schuler, William |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Frequency Explains the Inverse Correlation of Large Language Models' Size, Training Data Amount, and Surprisal's Fit to Reading Times
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024)
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024)
How Well Does First-Token Entropy Approximate Word Entropy as a Psycholinguistic Predictor?
di: Clark, Christian, et al.
Pubblicazione: (2025)
di: Clark, Christian, et al.
Pubblicazione: (2025)
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024)
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024)
Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word Probabilities
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024)
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024)
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage
di: Oh, Byung-Doh, et al.
Pubblicazione: (2025)
di: Oh, Byung-Doh, et al.
Pubblicazione: (2025)
Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal
di: Nair, Sathvik, et al.
Pubblicazione: (2026)
di: Nair, Sathvik, et al.
Pubblicazione: (2026)
LayerNorm Induces Recency Bias in Transformer Decoders
di: Kim, Junu, et al.
Pubblicazione: (2025)
di: Kim, Junu, et al.
Pubblicazione: (2025)
Citation Amnesia: On The Recency Bias of NLP and Other Academic Fields
di: Wahle, Jan Philip, et al.
Pubblicazione: (2024)
di: Wahle, Jan Philip, et al.
Pubblicazione: (2024)
Vectors from Larger Language Models Predict Human Reading Time and fMRI Data More Poorly when Dimensionality Expansion is Controlled
di: Lin, Yi-Chien, et al.
Pubblicazione: (2025)
di: Lin, Yi-Chien, et al.
Pubblicazione: (2025)
Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly
di: Lin, Yi-Chien, et al.
Pubblicazione: (2025)
di: Lin, Yi-Chien, et al.
Pubblicazione: (2025)
Emergence of Primacy and Recency Effect in Mamba: A Mechanistic Point of View
di: Airlangga, Muhammad Cendekia, et al.
Pubblicazione: (2025)
di: Airlangga, Muhammad Cendekia, et al.
Pubblicazione: (2025)
Big Tech-Funded AI Papers Have Higher Citation Impact, Greater Insularity, and Larger Recency Bias
di: Gnewuch, Max Martin, et al.
Pubblicazione: (2025)
di: Gnewuch, Max Martin, et al.
Pubblicazione: (2025)
How often do Answers Change? Estimating Recency Requirements in Question Answering
di: Piryani, Bhawna, et al.
Pubblicazione: (2026)
di: Piryani, Bhawna, et al.
Pubblicazione: (2026)
One-Topic-Doesn't-Fit-All: Transcreating Reading Comprehension Test for Personalized Learning
di: Han, Jieun, et al.
Pubblicazione: (2025)
di: Han, Jieun, et al.
Pubblicazione: (2025)
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
di: Oh, Gyutaek, et al.
Pubblicazione: (2025)
di: Oh, Gyutaek, et al.
Pubblicazione: (2025)
RoBIn: A Transformer-Based Model For Risk Of Bias Inference With Machine Reading Comprehension
di: Dias, Abel Corrêa, et al.
Pubblicazione: (2024)
di: Dias, Abel Corrêa, et al.
Pubblicazione: (2024)
Measuring Recency Bias In Sequential Recommendation Systems
di: Oh, Jeonglyul, et al.
Pubblicazione: (2024)
di: Oh, Jeonglyul, et al.
Pubblicazione: (2024)
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
di: Ye, Jiayuan, et al.
Pubblicazione: (2026)
di: Ye, Jiayuan, et al.
Pubblicazione: (2026)
Experiments in News Bias Detection with Pre-Trained Neural Transformers
di: Menzner, Tim, et al.
Pubblicazione: (2024)
di: Menzner, Tim, et al.
Pubblicazione: (2024)
To Words and Beyond: Probing Large Language Models for Sentence-Level Psycholinguistic Norms of Memorability and Reading Times
di: Clark, Thomas Hikaru, et al.
Pubblicazione: (2026)
di: Clark, Thomas Hikaru, et al.
Pubblicazione: (2026)
Gated Linear Attention Transformers with Hardware-Efficient Training
di: Yang, Songlin, et al.
Pubblicazione: (2023)
di: Yang, Songlin, et al.
Pubblicazione: (2023)
ModRWKV: Transformer Multimodality in Linear Time
di: Kang, Jiale, et al.
Pubblicazione: (2025)
di: Kang, Jiale, et al.
Pubblicazione: (2025)
Probing for Reading Times
di: Tsipidi, Eleftheria, et al.
Pubblicazione: (2026)
di: Tsipidi, Eleftheria, et al.
Pubblicazione: (2026)
Fairness Dynamics During Training
di: Patel, Krishna, et al.
Pubblicazione: (2025)
di: Patel, Krishna, et al.
Pubblicazione: (2025)
No Training Wheels: Steering Vectors for Bias Correction at Inference Time
di: Gupta, Aviral, et al.
Pubblicazione: (2025)
di: Gupta, Aviral, et al.
Pubblicazione: (2025)
Improving LLM Abilities in Idiomatic Translation
di: Donthi, Sundesh, et al.
Pubblicazione: (2024)
di: Donthi, Sundesh, et al.
Pubblicazione: (2024)
Transformer-VQ: Linear-Time Transformers via Vector Quantization
di: Lingle, Lucas D.
Pubblicazione: (2023)
di: Lingle, Lucas D.
Pubblicazione: (2023)
The Effect of Surprisal on Reading Times in Information Seeking and Repeated Reading
di: Klein, Keren Gruteke, et al.
Pubblicazione: (2024)
di: Klein, Keren Gruteke, et al.
Pubblicazione: (2024)
Analyzing Bias in Swiss Federal Supreme Court Judgments Using Facebook's Holistic Bias Dataset: Implications for Language Model Training
di: Wehnert, Sabine, et al.
Pubblicazione: (2025)
di: Wehnert, Sabine, et al.
Pubblicazione: (2025)
AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
di: Oh, Gyutaek, et al.
Pubblicazione: (2025)
di: Oh, Gyutaek, et al.
Pubblicazione: (2025)
Improved Models for Media Bias Detection and Subcategorization
di: Menzner, Tim, et al.
Pubblicazione: (2024)
di: Menzner, Tim, et al.
Pubblicazione: (2024)
Enhancing Training Data Attribution for Large Language Models with Fitting Error Consideration
di: Wu, Kangxi, et al.
Pubblicazione: (2024)
di: Wu, Kangxi, et al.
Pubblicazione: (2024)
Bias A-head? Analyzing Bias in Transformer-Based Language Model Attention Heads
di: Yang, Yi, et al.
Pubblicazione: (2023)
di: Yang, Yi, et al.
Pubblicazione: (2023)
Non-Linear Inference Time Intervention: Improving LLM Truthfulness
di: Hoscilowicz, Jakub, et al.
Pubblicazione: (2024)
di: Hoscilowicz, Jakub, et al.
Pubblicazione: (2024)
Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles
di: Lia, Nusrat Jahan, et al.
Pubblicazione: (2025)
di: Lia, Nusrat Jahan, et al.
Pubblicazione: (2025)
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
di: Huang, Ruiquan, et al.
Pubblicazione: (2025)
di: Huang, Ruiquan, et al.
Pubblicazione: (2025)
Re-Reading Improves Reasoning in Large Language Models
di: Xu, Xiaohan, et al.
Pubblicazione: (2023)
di: Xu, Xiaohan, et al.
Pubblicazione: (2023)
Mitigating Length Bias in RLHF through a Causal Lens
di: Kim, Hyeonji, et al.
Pubblicazione: (2025)
di: Kim, Hyeonji, et al.
Pubblicazione: (2025)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
di: Din, Alexander Yom, et al.
Pubblicazione: (2023)
di: Din, Alexander Yom, et al.
Pubblicazione: (2023)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
di: Oh, Juhyun, et al.
Pubblicazione: (2024)
di: Oh, Juhyun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Frequency Explains the Inverse Correlation of Large Language Models' Size, Training Data Amount, and Surprisal's Fit to Reading Times
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024) -
How Well Does First-Token Entropy Approximate Word Entropy as a Psycholinguistic Predictor?
di: Clark, Christian, et al.
Pubblicazione: (2025) -
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024) -
Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word Probabilities
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024) -
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage
di: Oh, Byung-Doh, et al.
Pubblicazione: (2025)