Filtering Beats Fine Tuning: A Bayesian Kalman View of In Context Learning in LLMs
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Kiruluta, Andrew |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Gradients to Riccati Geometry: Kalman World Models for Single-Pass Learning
von: Kiruluta, Andrew
Veröffentlicht: (2026)
von: Kiruluta, Andrew
Veröffentlicht: (2026)
Data-Driven Variational Basis Learning Beyond Neural Networks: A Non-Neural Framework for Adaptive Basis Discovery
von: Kiruluta, Andrew
Veröffentlicht: (2026)
von: Kiruluta, Andrew
Veröffentlicht: (2026)
Spectral Generative Flow Models: A Physics-Inspired Replacement for Vectorized Large Language Models
von: Kiruluta, Andrew
Veröffentlicht: (2026)
von: Kiruluta, Andrew
Veröffentlicht: (2026)
Adaptive Two Sided Laplace Transforms: A Learnable, Interpretable, and Scalable Replacement for Self-Attention
von: Kiruluta, Andrew
Veröffentlicht: (2025)
von: Kiruluta, Andrew
Veröffentlicht: (2025)
Entropic-Time Inference: Self-Organizing Large Language Model Decoding Beyond Attention
von: Kiruluta, Andrew
Veröffentlicht: (2026)
von: Kiruluta, Andrew
Veröffentlicht: (2026)
Kalman Filter Aided Federated Koopman Learning
von: Chen, Yutao, et al.
Veröffentlicht: (2025)
von: Chen, Yutao, et al.
Veröffentlicht: (2025)
Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
Effective Context in Transformers: An Analysis of Fragmentation and Tokenization
von: Fesharaki, Amirmehdi Jafari, et al.
Veröffentlicht: (2026)
von: Fesharaki, Amirmehdi Jafari, et al.
Veröffentlicht: (2026)
Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning
von: Mitra, Purbesh, et al.
Veröffentlicht: (2025)
von: Mitra, Purbesh, et al.
Veröffentlicht: (2025)
What Makes the Preferred Thinking Direction for LLMs in Multiple-choice Questions?
von: Zhang, Yizhe, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhe, et al.
Veröffentlicht: (2025)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
von: Yang, Tong, et al.
Veröffentlicht: (2024)
von: Yang, Tong, et al.
Veröffentlicht: (2024)
State Fourier Diffusion Language Model (SFDLM): A Scalable, Novel Iterative Approach to Language Modeling
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression
von: Elias, Noel, et al.
Veröffentlicht: (2024)
von: Elias, Noel, et al.
Veröffentlicht: (2024)
MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs
von: Mitra, Purbesh, et al.
Veröffentlicht: (2025)
von: Mitra, Purbesh, et al.
Veröffentlicht: (2025)
A Mathematical Theory for Learning Semantic Languages by Abstract Learners
von: Liao, Kuo-Yu, et al.
Veröffentlicht: (2024)
von: Liao, Kuo-Yu, et al.
Veröffentlicht: (2024)
Efficient Learned Data Compression via Dual-Stream Feature Decoupling
von: Ma, Huidong, et al.
Veröffentlicht: (2026)
von: Ma, Huidong, et al.
Veröffentlicht: (2026)
An Information-theoretic Multi-task Representation Learning Framework for Natural Language Understanding
von: Hu, Dou, et al.
Veröffentlicht: (2025)
von: Hu, Dou, et al.
Veröffentlicht: (2025)
SPEX: Scaling Feature Interaction Explanations for LLMs
von: Kang, Justin Singh, et al.
Veröffentlicht: (2025)
von: Kang, Justin Singh, et al.
Veröffentlicht: (2025)
Learnable Multi-Scale Wavelet Transformer: A Novel Alternative to Self-Attention
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
Beyond Self Attention: A Subquadratic Fourier Wavelet Transformer with Multi Modal Fusion
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2021)
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2021)
FourierNAT: A Fourier-Mixing-Based Non-Autoregressive Transformer for Parallel Sequence Generation
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
von: Català, Mar Gonzàlez I, et al.
Veröffentlicht: (2026)
von: Català, Mar Gonzàlez I, et al.
Veröffentlicht: (2026)
From Eigenmodes to Proofs: Integrating Graph Spectral Operators with Symbolic Interpretable Reasoning
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
von: Kiruluta, Andrew, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Generation as Noisy In-Context Learning: A Unified Theory and Risk Bounds
von: Guo, Yang, et al.
Veröffentlicht: (2025)
von: Guo, Yang, et al.
Veröffentlicht: (2025)
A Rate-Distortion Framework for Summarization
von: Arda, Enes, et al.
Veröffentlicht: (2025)
von: Arda, Enes, et al.
Veröffentlicht: (2025)
GraphRAFT: Retrieval Augmented Fine-Tuning for Knowledge Graphs on Graph Databases
von: Clemedtson, Alfred, et al.
Veröffentlicht: (2025)
von: Clemedtson, Alfred, et al.
Veröffentlicht: (2025)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
In-Context Learning and Fine-Tuning GPT for Argument Mining
von: Cabessa, Jérémie, et al.
Veröffentlicht: (2024)
von: Cabessa, Jérémie, et al.
Veröffentlicht: (2024)
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
von: Nagle, Alliot, et al.
Veröffentlicht: (2024)
von: Nagle, Alliot, et al.
Veröffentlicht: (2024)
Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations
von: Sitdhipol, Supawich, et al.
Veröffentlicht: (2025)
von: Sitdhipol, Supawich, et al.
Veröffentlicht: (2025)
Proposal and study of statistical features for string similarity computation and classification
von: Rodrigues, E. O., et al.
Veröffentlicht: (2026)
von: Rodrigues, E. O., et al.
Veröffentlicht: (2026)
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory
von: Zuo, Fei, et al.
Veröffentlicht: (2026)
von: Zuo, Fei, et al.
Veröffentlicht: (2026)
Speculative Decoding Scaling Laws (SDSL): Throughput Optimization Made Simple
von: Bozorgkhoo, Amirhossein, et al.
Veröffentlicht: (2026)
von: Bozorgkhoo, Amirhossein, et al.
Veröffentlicht: (2026)
Theoretical guarantees on the best-of-n alignment policy
von: Beirami, Ahmad, et al.
Veröffentlicht: (2024)
von: Beirami, Ahmad, et al.
Veröffentlicht: (2024)
Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
Understanding Factual Recall in Transformers via Associative Memories
von: Nichani, Eshaan, et al.
Veröffentlicht: (2024)
von: Nichani, Eshaan, et al.
Veröffentlicht: (2024)
InfAlign: Inference-aware language model alignment
von: Balashankar, Ananth, et al.
Veröffentlicht: (2024)
von: Balashankar, Ananth, et al.
Veröffentlicht: (2024)
Transformers on Markov Data: Constant Depth Suffices
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
An Enhanced Text Compression Approach Using Transformer-based Language Models
von: Rahman, Chowdhury Mofizur, et al.
Veröffentlicht: (2024)
von: Rahman, Chowdhury Mofizur, et al.
Veröffentlicht: (2024)
Cost-aware LLM-based Online Dataset Annotation
von: Elumar, Eray Can, et al.
Veröffentlicht: (2025)
von: Elumar, Eray Can, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Gradients to Riccati Geometry: Kalman World Models for Single-Pass Learning
von: Kiruluta, Andrew
Veröffentlicht: (2026) -
Data-Driven Variational Basis Learning Beyond Neural Networks: A Non-Neural Framework for Adaptive Basis Discovery
von: Kiruluta, Andrew
Veröffentlicht: (2026) -
Spectral Generative Flow Models: A Physics-Inspired Replacement for Vectorized Large Language Models
von: Kiruluta, Andrew
Veröffentlicht: (2026) -
Adaptive Two Sided Laplace Transforms: A Learnable, Interpretable, and Scalable Replacement for Self-Attention
von: Kiruluta, Andrew
Veröffentlicht: (2025) -
Entropic-Time Inference: Self-Organizing Large Language Model Decoding Beyond Attention
von: Kiruluta, Andrew
Veröffentlicht: (2026)