Just read twice: closing the recall gap for recurrent language models
Fuente:
arXiv
Salvato in:
| Autori principali: | Arora, Simran, Timalsina, Aman, Singhal, Aaryan, Spector, Benjamin, Eyuboglu, Sabri, Zhao, Xinyi, Rao, Ashish, Rudra, Atri, Ré, Christopher |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Simple linear attention language models balance the recall-throughput tradeoff
di: Arora, Simran, et al.
Pubblicazione: (2024)
di: Arora, Simran, et al.
Pubblicazione: (2024)
ThunderKittens: Simple, Fast, and Adorable AI Kernels
di: Spector, Benjamin F., et al.
Pubblicazione: (2024)
di: Spector, Benjamin F., et al.
Pubblicazione: (2024)
Towards Learning High-Precision Least Squares Algorithms with Sequence Models
di: Liu, Jerry, et al.
Pubblicazione: (2025)
di: Liu, Jerry, et al.
Pubblicazione: (2025)
LoLCATs: On Low-Rank Linearizing of Large Language Models
di: Zhang, Michael, et al.
Pubblicazione: (2024)
di: Zhang, Michael, et al.
Pubblicazione: (2024)
Cartridges: Lightweight and general-purpose long context representations via self-study
di: Eyuboglu, Sabri, et al.
Pubblicazione: (2025)
di: Eyuboglu, Sabri, et al.
Pubblicazione: (2025)
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
di: Arora, Simran, et al.
Pubblicazione: (2023)
di: Arora, Simran, et al.
Pubblicazione: (2023)
ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels
di: Sul, Stuart H., et al.
Pubblicazione: (2025)
di: Sul, Stuart H., et al.
Pubblicazione: (2025)
Constructing Efficient Fact-Storing MLPs for Transformers
di: Dugan, Owen, et al.
Pubblicazione: (2025)
di: Dugan, Owen, et al.
Pubblicazione: (2025)
Minions: Cost-efficient Collaboration Between On-device and Cloud Language Models
di: Narayan, Avanika, et al.
Pubblicazione: (2025)
di: Narayan, Avanika, et al.
Pubblicazione: (2025)
Explaining vague language
di: Égré, Paul, et al.
Pubblicazione: (2024)
di: Égré, Paul, et al.
Pubblicazione: (2024)
Late Time Acceleration with Observational Constraints in Modified Theories of Gravity
di: Arora, Simran
Pubblicazione: (2023)
di: Arora, Simran
Pubblicazione: (2023)
Counting Clinical Trials: New Evidence on Pharmaceutical Sector Productivity
di: Durvasula, Maya M., et al.
Pubblicazione: (2024)
di: Durvasula, Maya M., et al.
Pubblicazione: (2024)
Adaptive Rank Allocation: Speeding Up Modern Transformers with RaNA Adapters
di: Garcia, Roberto, et al.
Pubblicazione: (2025)
di: Garcia, Roberto, et al.
Pubblicazione: (2025)
Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT
di: Saad-Falcon, Jon, et al.
Pubblicazione: (2024)
di: Saad-Falcon, Jon, et al.
Pubblicazione: (2024)
Revisiting associative recall in modern recurrent models
di: Okpekpe, Destiny, et al.
Pubblicazione: (2025)
di: Okpekpe, Destiny, et al.
Pubblicazione: (2025)
The unregulated plant‐based ‘milk’ industry: A threat to nutrition, health and safety?
di: Simran Kaur Arora
Pubblicazione: (2024)
di: Simran Kaur Arora
Pubblicazione: (2024)
BWLer: Barycentric Weight Layer Elucidates a Precision-Conditioning Tradeoff for PINNs
di: Liu, Jerry, et al.
Pubblicazione: (2025)
di: Liu, Jerry, et al.
Pubblicazione: (2025)
Bayesian and Machine-Learning Analyses of Nonminimal $f(Q)$ Gravity and $H_0$ Tension
di: Arora, Simran, et al.
Pubblicazione: (2025)
di: Arora, Simran, et al.
Pubblicazione: (2025)
Towards Testable Type-III Leptogenesis in Non-Standard Early Universe Scenarios
di: Arora, Simran, et al.
Pubblicazione: (2026)
di: Arora, Simran, et al.
Pubblicazione: (2026)
KernelBench: Can LLMs Write Efficient GPU Kernels?
di: Ouyang, Anne, et al.
Pubblicazione: (2025)
di: Ouyang, Anne, et al.
Pubblicazione: (2025)
Interacting bosonic dark energy and fermionic dark matter in Einstein scalar Gauss-Bonnet gravity
di: Arora, Simran, et al.
Pubblicazione: (2025)
di: Arora, Simran, et al.
Pubblicazione: (2025)
Revisiting kink-like parametrization and constraints using OHD/Pantheon+/BAO samples
di: Arora, Simran, et al.
Pubblicazione: (2023)
di: Arora, Simran, et al.
Pubblicazione: (2023)
A sparse resolution of the DiPerna-Majda gap problem for $2$D Euler equations
di: Domínguez, Oscar, et al.
Pubblicazione: (2024)
di: Domínguez, Oscar, et al.
Pubblicazione: (2024)
Just rephrase it! Uncertainty estimation in closed-source language models via multiple rephrased queries
di: Yang, Adam, et al.
Pubblicazione: (2024)
di: Yang, Adam, et al.
Pubblicazione: (2024)
Time Slip as a Perceptual Construct: A Scientific and Probabilistic Analysis Rejecting Temporal Reversal
di: maan, Aaryan
Pubblicazione: (2026)
di: maan, Aaryan
Pubblicazione: (2026)
Teleportation Limits: A Unified Quantum-Classical and Consciousness-Based Framework for Physical and Identity Constraints
di: maan, Aaryan
Pubblicazione: (2026)
di: maan, Aaryan
Pubblicazione: (2026)
Temporal Recurrence of Pandemics: A 101-Year Cycle Hypothesis
di: Aaryan khan
Pubblicazione: (2025)
di: Aaryan khan
Pubblicazione: (2025)
KinetiDiff: Docking-Guided Diffusion for De Novo ACVR1 Inhibitor Design in Fibrodysplasia Ossificans Progressiva
di: Patel, Aaryan
Pubblicazione: (2026)
di: Patel, Aaryan
Pubblicazione: (2026)
How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
di: Park, Simon, et al.
Pubblicazione: (2025)
di: Park, Simon, et al.
Pubblicazione: (2025)
Embedding Generalized CP Symmetry in One Zero Texture Neutrino Mass Models
di: Priya, et al.
Pubblicazione: (2025)
di: Priya, et al.
Pubblicazione: (2025)
Constraining Spatial Curvature with Priors from Swampland Conjectures
di: Arora, Simran, et al.
Pubblicazione: (2026)
di: Arora, Simran, et al.
Pubblicazione: (2026)
Evaluation of mechanical, permeation, and degradation properties of poly(hydroxybutyrate) blends for sustainable packaging
di: Simran Ahuja, et al.
Pubblicazione: (2024)
di: Simran Ahuja, et al.
Pubblicazione: (2024)
The eV-Scale Sterile Neutrino and Neutrinoless Double Beta Decay
di: Priya, et al.
Pubblicazione: (2026)
di: Priya, et al.
Pubblicazione: (2026)
Optimistic Verifiable Training by Controlling Hardware Nondeterminism
di: Srivastava, Megha, et al.
Pubblicazione: (2024)
di: Srivastava, Megha, et al.
Pubblicazione: (2024)
Breaking and making genes: the genesis of novel traits in plants
di: Fiza Hamid, et al.
Pubblicazione: (2025)
di: Fiza Hamid, et al.
Pubblicazione: (2025)
Gifts that give twice are not twice as nice: Consumers avoid socially responsible gifts for picky recipients
di: Shih‐Chun (Daniel) Chin, et al.
Pubblicazione: (2024)
di: Shih‐Chun (Daniel) Chin, et al.
Pubblicazione: (2024)
What is Liberation? – Ācārya Sthaneshwar Timalsina
di: Timalsina, Staneshwar
Pubblicazione: (2023)
di: Timalsina, Staneshwar
Pubblicazione: (2023)
Shiva Nirvana Stotram (with lyrics) -- Devotional Song recited by Acharya Sthaneshwar
di: Timalsina, Staneshwar
Pubblicazione: (2023)
di: Timalsina, Staneshwar
Pubblicazione: (2023)
Teachings on the Truth of Mother Kali -- short answers from Acharya Sthaneshwar
di: Timalsina, Staneshwar
Pubblicazione: (2023)
di: Timalsina, Staneshwar
Pubblicazione: (2023)
What is anu(ta)? – Ācārya dr. Sthaneshwar Timalsina
di: Timalsina, Staneshwar
Pubblicazione: (2022)
di: Timalsina, Staneshwar
Pubblicazione: (2022)
Documenti analoghi
-
Simple linear attention language models balance the recall-throughput tradeoff
di: Arora, Simran, et al.
Pubblicazione: (2024) -
ThunderKittens: Simple, Fast, and Adorable AI Kernels
di: Spector, Benjamin F., et al.
Pubblicazione: (2024) -
Towards Learning High-Precision Least Squares Algorithms with Sequence Models
di: Liu, Jerry, et al.
Pubblicazione: (2025) -
LoLCATs: On Low-Rank Linearizing of Large Language Models
di: Zhang, Michael, et al.
Pubblicazione: (2024) -
Cartridges: Lightweight and general-purpose long context representations via self-study
di: Eyuboglu, Sabri, et al.
Pubblicazione: (2025)