Decomposing Prediction Mechanisms for In-Context Recall
Fuente:
arXiv
Salvato in:
| Autori principali: | Daniels, Sultan, Davis, Dylan, Gautam, Dhruv, Liao, Wentinn, Ranade, Gireeja, Sahai, Anant |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Thermodynamic Costs of Simple Linear Regression
di: D'Ambrosia, Samuel H., et al.
Pubblicazione: (2026)
di: D'Ambrosia, Samuel H., et al.
Pubblicazione: (2026)
Polynomial Regression as a Task for Understanding In-context Learning Through Finetuning and Alignment
di: Wilcoxson, Max, et al.
Pubblicazione: (2024)
di: Wilcoxson, Max, et al.
Pubblicazione: (2024)
Provable Weak-to-Strong Generalization via Benign Overfitting
di: Wu, David X., et al.
Pubblicazione: (2024)
di: Wu, David X., et al.
Pubblicazione: (2024)
Precise Asymptotic Generalization for Multiclass Classification with Overparameterized Linear Models
di: Wu, David X., et al.
Pubblicazione: (2023)
di: Wu, David X., et al.
Pubblicazione: (2023)
LeanTutor: Towards a Verified AI Mathematical Proof Tutor
di: Patel, Manooshree, et al.
Pubblicazione: (2026)
di: Patel, Manooshree, et al.
Pubblicazione: (2026)
Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
di: Wu, David X., et al.
Pubblicazione: (2025)
di: Wu, David X., et al.
Pubblicazione: (2025)
Can Custom Models Learn In-Context? An Exploration of Hybrid Architecture Performance on In-Context Learning Tasks
di: Campbell, Ryan, et al.
Pubblicazione: (2024)
di: Campbell, Ryan, et al.
Pubblicazione: (2024)
LLM In-Context Recall is Prompt Dependent
di: Machlab, Daniel, et al.
Pubblicazione: (2024)
di: Machlab, Daniel, et al.
Pubblicazione: (2024)
Phase Transitions of Diversity in Stochastic Block Model Dynamics
di: Brânzei, Simina, et al.
Pubblicazione: (2023)
di: Brânzei, Simina, et al.
Pubblicazione: (2023)
Fine-Tuning Dynamics of In-Context Factual Recall in Transformers
di: Huang, Ruomin, et al.
Pubblicazione: (2026)
di: Huang, Ruomin, et al.
Pubblicazione: (2026)
DoMINO: A Decomposable Multi-scale Iterative Neural Operator for Modeling Large Scale Engineering Simulations
di: Ranade, Rishikesh, et al.
Pubblicazione: (2025)
di: Ranade, Rishikesh, et al.
Pubblicazione: (2025)
Necessary and Sufficient Oracles: Toward a Computational Taxonomy For Reinforcement Learning
di: Rohatgi, Dhruv, et al.
Pubblicazione: (2025)
di: Rohatgi, Dhruv, et al.
Pubblicazione: (2025)
Constrained Contextual Bandits with Adversarial Contexts
di: Sarkar, Dhruv, et al.
Pubblicazione: (2026)
di: Sarkar, Dhruv, et al.
Pubblicazione: (2026)
Decomposing multimodal embedding spaces with group-sparse autoencoders
di: Kaushik, Chiraag, et al.
Pubblicazione: (2026)
di: Kaushik, Chiraag, et al.
Pubblicazione: (2026)
Stabilizing a linear system using phone calls when time is information
di: Khojasteh, Mohammad Javad, et al.
Pubblicazione: (2018)
di: Khojasteh, Mohammad Javad, et al.
Pubblicazione: (2018)
A Multi-Component AI Framework for Computational Psychology: From Robust Predictive Modeling to Deployed Generative Dialogue
di: Pareek, Anant
Pubblicazione: (2025)
di: Pareek, Anant
Pubblicazione: (2025)
A Simple Reduction Scheme for Constrained Contextual Bandits with Adversarial Contexts via Regression
di: Sarkar, Dhruv, et al.
Pubblicazione: (2026)
di: Sarkar, Dhruv, et al.
Pubblicazione: (2026)
Decomposing Attention To Find Context-Sensitive Neurons
di: Gibson, Alex
Pubblicazione: (2025)
di: Gibson, Alex
Pubblicazione: (2025)
How Transformers Learn In-Context Recall Tasks? Optimality, Training Dynamics and Generalization
di: Nguyen, Quan, et al.
Pubblicazione: (2025)
di: Nguyen, Quan, et al.
Pubblicazione: (2025)
Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification
di: Rohatgi, Dhruv, et al.
Pubblicazione: (2025)
di: Rohatgi, Dhruv, et al.
Pubblicazione: (2025)
The Entropy of Floating-Point Numbers
di: Daniels, Sultan, et al.
Pubblicazione: (2026)
di: Daniels, Sultan, et al.
Pubblicazione: (2026)
Decomposing and Editing Predictions by Modeling Model Computation
di: Shah, Harshay, et al.
Pubblicazione: (2024)
di: Shah, Harshay, et al.
Pubblicazione: (2024)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
di: Chughtai, Bilal, et al.
Pubblicazione: (2024)
di: Chughtai, Bilal, et al.
Pubblicazione: (2024)
ThinkSwitch: Context Distillation with LoRA and Weight Interpolation for Specific-Purpose Reasoning Tasks
di: Saini, Dhruv, et al.
Pubblicazione: (2026)
di: Saini, Dhruv, et al.
Pubblicazione: (2026)
Secure Generalization through Stochastic Bidirectional Parameter Updates Using Dual-Gradient Mechanism
di: Goel, Shourya, et al.
Pubblicazione: (2025)
di: Goel, Shourya, et al.
Pubblicazione: (2025)
Probably Approximately Precision and Recall Learning
di: Cohen, Lee, et al.
Pubblicazione: (2024)
di: Cohen, Lee, et al.
Pubblicazione: (2024)
Precision and Recall Reject Curves for Classification
di: Fischer, Lydia, et al.
Pubblicazione: (2023)
di: Fischer, Lydia, et al.
Pubblicazione: (2023)
Self-Consolidating Language Models: Continual Knowledge Incorporation from Context
di: Wang, Zekun, et al.
Pubblicazione: (2026)
di: Wang, Zekun, et al.
Pubblicazione: (2026)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2026)
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2026)
Exploring Context Window of Large Language Models via Decomposed Positional Vectors
di: Dong, Zican, et al.
Pubblicazione: (2024)
di: Dong, Zican, et al.
Pubblicazione: (2024)
Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Reasoning
di: Štorek, Adam, et al.
Pubblicazione: (2025)
di: Štorek, Adam, et al.
Pubblicazione: (2025)
Heterogeneous Federated Learning Systems for Time-Series Power Consumption Prediction with Multi-Head Embedding Mechanism
di: Syu, Jia-Hao, et al.
Pubblicazione: (2025)
di: Syu, Jia-Hao, et al.
Pubblicazione: (2025)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
di: Lv, Ang, et al.
Pubblicazione: (2024)
di: Lv, Ang, et al.
Pubblicazione: (2024)
GeoTransolver: Learning Physics on Irregular Domains Using Multi-scale Geometry Aware Physics Attention Transformer
di: Adams, Corey, et al.
Pubblicazione: (2025)
di: Adams, Corey, et al.
Pubblicazione: (2025)
Learning to Recall with Transformers Beyond Orthogonal Embeddings
di: Vural, Nuri Mert, et al.
Pubblicazione: (2026)
di: Vural, Nuri Mert, et al.
Pubblicazione: (2026)
FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning
di: Han, Xing, et al.
Pubblicazione: (2026)
di: Han, Xing, et al.
Pubblicazione: (2026)
FullRecall: A Semantic Search-Based Ranking Approach for Maximizing Recall in Patent Retrieval
di: Ali, Amna, et al.
Pubblicazione: (2025)
di: Ali, Amna, et al.
Pubblicazione: (2025)
Decomposable Neuro Symbolic Regression
di: Morales, Giorgio, et al.
Pubblicazione: (2025)
di: Morales, Giorgio, et al.
Pubblicazione: (2025)
Decomposable Transformer Point Processes
di: Panos, Aristeidis
Pubblicazione: (2024)
di: Panos, Aristeidis
Pubblicazione: (2024)
Self-Improvement in Language Models: The Sharpening Mechanism
di: Huang, Audrey, et al.
Pubblicazione: (2024)
di: Huang, Audrey, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Thermodynamic Costs of Simple Linear Regression
di: D'Ambrosia, Samuel H., et al.
Pubblicazione: (2026) -
Polynomial Regression as a Task for Understanding In-context Learning Through Finetuning and Alignment
di: Wilcoxson, Max, et al.
Pubblicazione: (2024) -
Provable Weak-to-Strong Generalization via Benign Overfitting
di: Wu, David X., et al.
Pubblicazione: (2024) -
Precise Asymptotic Generalization for Multiclass Classification with Overparameterized Linear Models
di: Wu, David X., et al.
Pubblicazione: (2023) -
LeanTutor: Towards a Verified AI Mathematical Proof Tutor
di: Patel, Manooshree, et al.
Pubblicazione: (2026)