Gespeichert in:
| Hauptverfasser: | Raposo, David, Ritter, Sam, Richards, Blake, Lillicrap, Timothy, Humphreys, Peter Conway, Santoro, Adam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2404.02258 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
von: Bae, Sangmin, et al.
Veröffentlicht: (2025)
von: Bae, Sangmin, et al.
Veröffentlicht: (2025)
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
von: Li, Melody Zixuan, et al.
Veröffentlicht: (2025)
von: Li, Melody Zixuan, et al.
Veröffentlicht: (2025)
A path to natural language through tokenisation and transformers
von: Berman, David S., et al.
Veröffentlicht: (2026)
von: Berman, David S., et al.
Veröffentlicht: (2026)
Detecting out-of-distribution text using topological features of transformer-based language models
von: Pollano, Andres, et al.
Veröffentlicht: (2023)
von: Pollano, Andres, et al.
Veröffentlicht: (2023)
MoDification: Mixture of Depths Made Easy
von: Zhang, Chen, et al.
Veröffentlicht: (2024)
von: Zhang, Chen, et al.
Veröffentlicht: (2024)
Physical models realizing the transformer architecture of large language models
von: Chen, Zeqian
Veröffentlicht: (2025)
von: Chen, Zeqian
Veröffentlicht: (2025)
Zero-shot data citation function classification using transformer-based large language models (LLMs)
von: Byers, Neil, et al.
Veröffentlicht: (2025)
von: Byers, Neil, et al.
Veröffentlicht: (2025)
Comparison of different Unique hard attention transformer models by the formal languages they can recognize
von: Ryvkin, Leonid
Veröffentlicht: (2025)
von: Ryvkin, Leonid
Veröffentlicht: (2025)
Training Agents Inside of Scalable World Models
von: Hafner, Danijar, et al.
Veröffentlicht: (2025)
von: Hafner, Danijar, et al.
Veröffentlicht: (2025)
Alignment faking in large language models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
How do language models learn facts? Dynamics, curricula and hallucinations
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2025)
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2025)
Question answering system of bridge design specification based on large language model
von: Zhang, Leye, et al.
Veröffentlicht: (2024)
von: Zhang, Leye, et al.
Veröffentlicht: (2024)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
von: Knupp, Jonas, et al.
Veröffentlicht: (2026)
von: Knupp, Jonas, et al.
Veröffentlicht: (2026)
Auditing language models for hidden objectives
von: Marks, Samuel, et al.
Veröffentlicht: (2025)
von: Marks, Samuel, et al.
Veröffentlicht: (2025)
Aligning language models with human preferences
von: Korbak, Tomasz
Veröffentlicht: (2024)
von: Korbak, Tomasz
Veröffentlicht: (2024)
Evaluating language models as risk scores
von: Cruz, André F., et al.
Veröffentlicht: (2024)
von: Cruz, André F., et al.
Veröffentlicht: (2024)
Exploring prompts to elicit memorization in masked language model-based named entity recognition
von: Xia, Yuxi, et al.
Veröffentlicht: (2024)
von: Xia, Yuxi, et al.
Veröffentlicht: (2024)
Attention based Bidirectional GRU hybrid model for inappropriate content detection in Urdu language
von: Shoukat, Ezzah, et al.
Veröffentlicht: (2025)
von: Shoukat, Ezzah, et al.
Veröffentlicht: (2025)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
Dynamic layer selection in decoder-only transformers
von: Glavas, Theodore, et al.
Veröffentlicht: (2024)
von: Glavas, Theodore, et al.
Veröffentlicht: (2024)
Mastering Diverse Domains through World Models
von: Hafner, Danijar, et al.
Veröffentlicht: (2023)
von: Hafner, Danijar, et al.
Veröffentlicht: (2023)
A comparison of pipelines for the translation of a low resource language based on transformers
von: Bonfanti, Chiara, et al.
Veröffentlicht: (2025)
von: Bonfanti, Chiara, et al.
Veröffentlicht: (2025)
The language of time: a language model perspective on time-series foundation models
von: Xie, Yi, et al.
Veröffentlicht: (2025)
von: Xie, Yi, et al.
Veröffentlicht: (2025)
Amortizing intractable inference in large language models
von: Hu, Edward J., et al.
Veröffentlicht: (2023)
von: Hu, Edward J., et al.
Veröffentlicht: (2023)
Continuous-Depth Transformers with Learned Control Dynamics
von: Jemley, Peter
Veröffentlicht: (2026)
von: Jemley, Peter
Veröffentlicht: (2026)
A meta-analysis on the performance of machine-learning based language models for sentiment analysis
von: Rohde, Elena, et al.
Veröffentlicht: (2025)
von: Rohde, Elena, et al.
Veröffentlicht: (2025)
Perturbed examples reveal invariances shared by language models
von: Rawal, Ruchit, et al.
Veröffentlicht: (2023)
von: Rawal, Ruchit, et al.
Veröffentlicht: (2023)
A mean teacher algorithm for unlearning of language models
von: Klochkov, Yegor
Veröffentlicht: (2025)
von: Klochkov, Yegor
Veröffentlicht: (2025)
Do language models plan ahead for future tokens?
von: Wu, Wilson, et al.
Veröffentlicht: (2024)
von: Wu, Wilson, et al.
Veröffentlicht: (2024)
Visualizing token importance for black-box language models
von: Rauba, Paulius, et al.
Veröffentlicht: (2025)
von: Rauba, Paulius, et al.
Veröffentlicht: (2025)
Representation in large language models
von: Yetman, Cameron
Veröffentlicht: (2025)
von: Yetman, Cameron
Veröffentlicht: (2025)
Investigating and Alleviating Harm Amplification in LLM Interactions
von: Guo, Ruohao, et al.
Veröffentlicht: (2026)
von: Guo, Ruohao, et al.
Veröffentlicht: (2026)
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
von: Guo, Ruohao, et al.
Veröffentlicht: (2023)
von: Guo, Ruohao, et al.
Veröffentlicht: (2023)
Prompt reinforcing for long-term planning of large language models
von: Lin, Hsien-Chin, et al.
Veröffentlicht: (2025)
von: Lin, Hsien-Chin, et al.
Veröffentlicht: (2025)
Machine-generated text detection prevents language model collapse
von: Drayson, George, et al.
Veröffentlicht: (2025)
von: Drayson, George, et al.
Veröffentlicht: (2025)
Fresh in memory: Training-order recency is linearly encoded in language model activations
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2025)
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2025)
Language Models can Self-Improve at State-Value Estimation for Better Search
von: Mendes, Ethan, et al.
Veröffentlicht: (2025)
von: Mendes, Ethan, et al.
Veröffentlicht: (2025)
No Need to Talk: Asynchronous Mixture of Language Models
von: Filippova, Anastasiia, et al.
Veröffentlicht: (2024)
von: Filippova, Anastasiia, et al.
Veröffentlicht: (2024)
Lightweight reranking for language model generations
von: Jain, Siddhartha, et al.
Veröffentlicht: (2023)
von: Jain, Siddhartha, et al.
Veröffentlicht: (2023)
Boosting classification reliability of NLP transformer models in the long run
von: Kmetty, Zoltán, et al.
Veröffentlicht: (2023)
von: Kmetty, Zoltán, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
von: Bae, Sangmin, et al.
Veröffentlicht: (2025) -
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
von: Li, Melody Zixuan, et al.
Veröffentlicht: (2025) -
A path to natural language through tokenisation and transformers
von: Berman, David S., et al.
Veröffentlicht: (2026) -
Detecting out-of-distribution text using topological features of transformer-based language models
von: Pollano, Andres, et al.
Veröffentlicht: (2023) -
MoDification: Mixture of Depths Made Easy
von: Zhang, Chen, et al.
Veröffentlicht: (2024)