Enregistré dans:
| Auteurs principaux: | Zucchet, Nicolas, Bornschein, Jörg, Chan, Stephanie, Lampinen, Andrew, Pascanu, Razvan, De, Soham |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2503.21676 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
On the generalization of language models from in-context learning and finetuning: a controlled study
par: Lampinen, Andrew K., et autres
Publié: (2025)
par: Lampinen, Andrew K., et autres
Publié: (2025)
Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models
par: Rannen-Triki, Amal, et autres
Publié: (2024)
par: Rannen-Triki, Amal, et autres
Publié: (2024)
The emergence of sparse attention: impact of data distribution and benefits of repetition
par: Zucchet, Nicolas, et autres
Publié: (2025)
par: Zucchet, Nicolas, et autres
Publié: (2025)
The Illusion of Stochasticity in LLMs
par: Gu, Xiangming, et autres
Publié: (2026)
par: Gu, Xiangming, et autres
Publié: (2026)
The broader spectrum of in-context learning
par: Lampinen, Andrew Kyle, et autres
Publié: (2024)
par: Lampinen, Andrew Kyle, et autres
Publié: (2024)
Linear representations in language models can change dramatically over a conversation
par: Lampinen, Andrew Kyle, et autres
Publié: (2026)
par: Lampinen, Andrew Kyle, et autres
Publié: (2026)
The in-context inductive biases of vision-language models differ across modalities
par: Allen, Kelsey, et autres
Publié: (2025)
par: Allen, Kelsey, et autres
Publié: (2025)
Transformers need glasses! Information over-squashing in language tasks
par: Barbero, Federico, et autres
Publié: (2024)
par: Barbero, Federico, et autres
Publié: (2024)
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
par: Schmied, Thomas, et autres
Publié: (2025)
par: Schmied, Thomas, et autres
Publié: (2025)
Round and Round We Go! What makes Rotary Positional Encodings useful?
par: Barbero, Federico, et autres
Publié: (2024)
par: Barbero, Federico, et autres
Publié: (2024)
Perplexity Cannot Always Tell Right from Wrong
par: Veličković, Petar, et autres
Publié: (2026)
par: Veličković, Petar, et autres
Publié: (2026)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
par: Orvieto, Antonio, et autres
Publié: (2023)
par: Orvieto, Antonio, et autres
Publié: (2023)
Latent learning: episodic memory complements parametric learning by enabling flexible reuse of experiences
par: Lampinen, Andrew Kyle, et autres
Publié: (2025)
par: Lampinen, Andrew Kyle, et autres
Publié: (2025)
Learned feature representations are biased by complexity, learning order, position, and more
par: Lampinen, Andrew Kyle, et autres
Publié: (2024)
par: Lampinen, Andrew Kyle, et autres
Publié: (2024)
Transformers meet Neural Algorithmic Reasoners
par: Bounsi, Wilfried, et autres
Publié: (2024)
par: Bounsi, Wilfried, et autres
Publié: (2024)
Just-in-time and distributed task representations in language models
par: Li, Yuxuan, et autres
Publié: (2025)
par: Li, Yuxuan, et autres
Publié: (2025)
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
par: Gu, Xiangming, et autres
Publié: (2026)
par: Gu, Xiangming, et autres
Publié: (2026)
Fine-Tuned In-Context Learners for Efficient Adaptation
par: Bornschein, Jorg, et autres
Publié: (2025)
par: Bornschein, Jorg, et autres
Publié: (2025)
Language models show human-like content effects on reasoning tasks
par: Dasgupta, Ishita, et autres
Publié: (2022)
par: Dasgupta, Ishita, et autres
Publié: (2022)
Filter Equivariant Functions: A symmetric account of length-general extrapolation on lists
par: Lewis, Owen, et autres
Publié: (2025)
par: Lewis, Owen, et autres
Publié: (2025)
Interpretability Illusions in the Generalization of Simplified Models
par: Friedman, Dan, et autres
Publié: (2023)
par: Friedman, Dan, et autres
Publié: (2023)
Kalman Filter for Online Classification of Non-Stationary Data
par: Titsias, Michalis K., et autres
Publié: (2023)
par: Titsias, Michalis K., et autres
Publié: (2023)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
par: De, Soham, et autres
Publié: (2024)
par: De, Soham, et autres
Publié: (2024)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
par: Lee, Jin Hwa, et autres
Publié: (2025)
par: Lee, Jin Hwa, et autres
Publié: (2025)
Survey on reinforcement learning for language processing
par: Uc-Cetina, Victor, et autres
Publié: (2021)
par: Uc-Cetina, Victor, et autres
Publié: (2021)
Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency
par: Goel, Aman, et autres
Publié: (2025)
par: Goel, Aman, et autres
Publié: (2025)
The representation landscape of few-shot learning and fine-tuning in large language models
par: Doimo, Diego, et autres
Publié: (2024)
par: Doimo, Diego, et autres
Publié: (2024)
Meta-learning how to Share Credit among Macro-Actions
par: Hosu, Ionel-Alexandru, et autres
Publié: (2025)
par: Hosu, Ionel-Alexandru, et autres
Publié: (2025)
An evolutionary perspective on modes of learning in Transformers
par: Ku, Alexander Y., et autres
Publié: (2025)
par: Ku, Alexander Y., et autres
Publié: (2025)
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
par: Raposo, David, et autres
Publié: (2024)
par: Raposo, David, et autres
Publié: (2024)
Enhancing ASD detection accuracy: a combined approach of machine learning and deep learning models with natural language processing
par: Rubio-Martín, Sergio, et autres
Publié: (2024)
par: Rubio-Martín, Sergio, et autres
Publié: (2024)
Reducing hallucination in structured outputs via Retrieval-Augmented Generation
par: Béchard, Patrice, et autres
Publié: (2024)
par: Béchard, Patrice, et autres
Publié: (2024)
Large language models reorganize representational geometry during in-context learning
par: Xiong, Hua-Dong, et autres
Publié: (2026)
par: Xiong, Hua-Dong, et autres
Publié: (2026)
A meta-analysis on the performance of machine-learning based language models for sentiment analysis
par: Rohde, Elena, et autres
Publié: (2025)
par: Rohde, Elena, et autres
Publié: (2025)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
par: Zucchet, Nicolas, et autres
Publié: (2024)
par: Zucchet, Nicolas, et autres
Publié: (2024)
Evaluating Representations with Readout Model Switching
par: Li, Yazhe, et autres
Publié: (2023)
par: Li, Yazhe, et autres
Publié: (2023)
Denoising Autoregressive Representation Learning
par: Li, Yazhe, et autres
Publié: (2024)
par: Li, Yazhe, et autres
Publié: (2024)
DevBench: A multimodal developmental benchmark for language learning
par: Tan, Alvin Wei Ming, et autres
Publié: (2024)
par: Tan, Alvin Wei Ming, et autres
Publié: (2024)
Perturbation: A simple and efficient adversarial tracer for representation learning in language models
par: Rozner, Joshua, et autres
Publié: (2026)
par: Rozner, Joshua, et autres
Publié: (2026)
Aligning language models with human preferences
par: Korbak, Tomasz
Publié: (2024)
par: Korbak, Tomasz
Publié: (2024)
Documents similaires
-
On the generalization of language models from in-context learning and finetuning: a controlled study
par: Lampinen, Andrew K., et autres
Publié: (2025) -
Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models
par: Rannen-Triki, Amal, et autres
Publié: (2024) -
The emergence of sparse attention: impact of data distribution and benefits of repetition
par: Zucchet, Nicolas, et autres
Publié: (2025) -
The Illusion of Stochasticity in LLMs
par: Gu, Xiangming, et autres
Publié: (2026) -
The broader spectrum of in-context learning
par: Lampinen, Andrew Kyle, et autres
Publié: (2024)