Salvato in:
| Autori principali: | Zucchet, Nicolas, Bornschein, Jörg, Chan, Stephanie, Lampinen, Andrew, Pascanu, Razvan, De, Soham |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2503.21676 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On the generalization of language models from in-context learning and finetuning: a controlled study
di: Lampinen, Andrew K., et al.
Pubblicazione: (2025)
di: Lampinen, Andrew K., et al.
Pubblicazione: (2025)
Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models
di: Rannen-Triki, Amal, et al.
Pubblicazione: (2024)
di: Rannen-Triki, Amal, et al.
Pubblicazione: (2024)
The emergence of sparse attention: impact of data distribution and benefits of repetition
di: Zucchet, Nicolas, et al.
Pubblicazione: (2025)
di: Zucchet, Nicolas, et al.
Pubblicazione: (2025)
The Illusion of Stochasticity in LLMs
di: Gu, Xiangming, et al.
Pubblicazione: (2026)
di: Gu, Xiangming, et al.
Pubblicazione: (2026)
The broader spectrum of in-context learning
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2024)
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2024)
Linear representations in language models can change dramatically over a conversation
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2026)
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2026)
The in-context inductive biases of vision-language models differ across modalities
di: Allen, Kelsey, et al.
Pubblicazione: (2025)
di: Allen, Kelsey, et al.
Pubblicazione: (2025)
Transformers need glasses! Information over-squashing in language tasks
di: Barbero, Federico, et al.
Pubblicazione: (2024)
di: Barbero, Federico, et al.
Pubblicazione: (2024)
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
di: Schmied, Thomas, et al.
Pubblicazione: (2025)
di: Schmied, Thomas, et al.
Pubblicazione: (2025)
Round and Round We Go! What makes Rotary Positional Encodings useful?
di: Barbero, Federico, et al.
Pubblicazione: (2024)
di: Barbero, Federico, et al.
Pubblicazione: (2024)
Perplexity Cannot Always Tell Right from Wrong
di: Veličković, Petar, et al.
Pubblicazione: (2026)
di: Veličković, Petar, et al.
Pubblicazione: (2026)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
di: Orvieto, Antonio, et al.
Pubblicazione: (2023)
di: Orvieto, Antonio, et al.
Pubblicazione: (2023)
Latent learning: episodic memory complements parametric learning by enabling flexible reuse of experiences
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2025)
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2025)
Learned feature representations are biased by complexity, learning order, position, and more
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2024)
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2024)
Transformers meet Neural Algorithmic Reasoners
di: Bounsi, Wilfried, et al.
Pubblicazione: (2024)
di: Bounsi, Wilfried, et al.
Pubblicazione: (2024)
Just-in-time and distributed task representations in language models
di: Li, Yuxuan, et al.
Pubblicazione: (2025)
di: Li, Yuxuan, et al.
Pubblicazione: (2025)
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
di: Gu, Xiangming, et al.
Pubblicazione: (2026)
di: Gu, Xiangming, et al.
Pubblicazione: (2026)
Fine-Tuned In-Context Learners for Efficient Adaptation
di: Bornschein, Jorg, et al.
Pubblicazione: (2025)
di: Bornschein, Jorg, et al.
Pubblicazione: (2025)
Language models show human-like content effects on reasoning tasks
di: Dasgupta, Ishita, et al.
Pubblicazione: (2022)
di: Dasgupta, Ishita, et al.
Pubblicazione: (2022)
Filter Equivariant Functions: A symmetric account of length-general extrapolation on lists
di: Lewis, Owen, et al.
Pubblicazione: (2025)
di: Lewis, Owen, et al.
Pubblicazione: (2025)
Interpretability Illusions in the Generalization of Simplified Models
di: Friedman, Dan, et al.
Pubblicazione: (2023)
di: Friedman, Dan, et al.
Pubblicazione: (2023)
Kalman Filter for Online Classification of Non-Stationary Data
di: Titsias, Michalis K., et al.
Pubblicazione: (2023)
di: Titsias, Michalis K., et al.
Pubblicazione: (2023)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
di: De, Soham, et al.
Pubblicazione: (2024)
di: De, Soham, et al.
Pubblicazione: (2024)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
di: Lee, Jin Hwa, et al.
Pubblicazione: (2025)
di: Lee, Jin Hwa, et al.
Pubblicazione: (2025)
Survey on reinforcement learning for language processing
di: Uc-Cetina, Victor, et al.
Pubblicazione: (2021)
di: Uc-Cetina, Victor, et al.
Pubblicazione: (2021)
Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency
di: Goel, Aman, et al.
Pubblicazione: (2025)
di: Goel, Aman, et al.
Pubblicazione: (2025)
The representation landscape of few-shot learning and fine-tuning in large language models
di: Doimo, Diego, et al.
Pubblicazione: (2024)
di: Doimo, Diego, et al.
Pubblicazione: (2024)
Meta-learning how to Share Credit among Macro-Actions
di: Hosu, Ionel-Alexandru, et al.
Pubblicazione: (2025)
di: Hosu, Ionel-Alexandru, et al.
Pubblicazione: (2025)
An evolutionary perspective on modes of learning in Transformers
di: Ku, Alexander Y., et al.
Pubblicazione: (2025)
di: Ku, Alexander Y., et al.
Pubblicazione: (2025)
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
di: Raposo, David, et al.
Pubblicazione: (2024)
di: Raposo, David, et al.
Pubblicazione: (2024)
Enhancing ASD detection accuracy: a combined approach of machine learning and deep learning models with natural language processing
di: Rubio-Martín, Sergio, et al.
Pubblicazione: (2024)
di: Rubio-Martín, Sergio, et al.
Pubblicazione: (2024)
Reducing hallucination in structured outputs via Retrieval-Augmented Generation
di: Béchard, Patrice, et al.
Pubblicazione: (2024)
di: Béchard, Patrice, et al.
Pubblicazione: (2024)
Large language models reorganize representational geometry during in-context learning
di: Xiong, Hua-Dong, et al.
Pubblicazione: (2026)
di: Xiong, Hua-Dong, et al.
Pubblicazione: (2026)
A meta-analysis on the performance of machine-learning based language models for sentiment analysis
di: Rohde, Elena, et al.
Pubblicazione: (2025)
di: Rohde, Elena, et al.
Pubblicazione: (2025)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
di: Zucchet, Nicolas, et al.
Pubblicazione: (2024)
di: Zucchet, Nicolas, et al.
Pubblicazione: (2024)
Evaluating Representations with Readout Model Switching
di: Li, Yazhe, et al.
Pubblicazione: (2023)
di: Li, Yazhe, et al.
Pubblicazione: (2023)
Denoising Autoregressive Representation Learning
di: Li, Yazhe, et al.
Pubblicazione: (2024)
di: Li, Yazhe, et al.
Pubblicazione: (2024)
DevBench: A multimodal developmental benchmark for language learning
di: Tan, Alvin Wei Ming, et al.
Pubblicazione: (2024)
di: Tan, Alvin Wei Ming, et al.
Pubblicazione: (2024)
Perturbation: A simple and efficient adversarial tracer for representation learning in language models
di: Rozner, Joshua, et al.
Pubblicazione: (2026)
di: Rozner, Joshua, et al.
Pubblicazione: (2026)
Aligning language models with human preferences
di: Korbak, Tomasz
Pubblicazione: (2024)
di: Korbak, Tomasz
Pubblicazione: (2024)
Documenti analoghi
-
On the generalization of language models from in-context learning and finetuning: a controlled study
di: Lampinen, Andrew K., et al.
Pubblicazione: (2025) -
Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models
di: Rannen-Triki, Amal, et al.
Pubblicazione: (2024) -
The emergence of sparse attention: impact of data distribution and benefits of repetition
di: Zucchet, Nicolas, et al.
Pubblicazione: (2025) -
The Illusion of Stochasticity in LLMs
di: Gu, Xiangming, et al.
Pubblicazione: (2026) -
The broader spectrum of in-context learning
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2024)