Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yingcong, Rawat, Ankit Singh, Oymak, Samet |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
Mechanics of Next Token Prediction with Self-Attention
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
Transformers as Support Vector Machines
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2023)
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2023)
When and How Unlabeled Data Provably Improve In-Context Learning
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
von: Li, Yingcong, et al.
Veröffentlicht: (2025)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
von: Gautam, Tanmay, et al.
Veröffentlicht: (2024)
von: Gautam, Tanmay, et al.
Veröffentlicht: (2024)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
von: Pan, Rui, et al.
Veröffentlicht: (2024)
von: Pan, Rui, et al.
Veröffentlicht: (2024)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
Causal LLM Routing: End-to-End Regret Minimization from Observational Data
von: Tsiourvas, Asterios, et al.
Veröffentlicht: (2025)
von: Tsiourvas, Asterios, et al.
Veröffentlicht: (2025)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
von: Huang, Xinmeng, et al.
Veröffentlicht: (2024)
von: Huang, Xinmeng, et al.
Veröffentlicht: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
von: Zhang, Zhixin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhixin, et al.
Veröffentlicht: (2025)
Solving General Natural-Language-Description Optimization Problems with Large Language Models
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
DiaBlo: Diagonal Blocks Are Sufficient For Finetuning
von: Gurses, Selcuk, et al.
Veröffentlicht: (2025)
von: Gurses, Selcuk, et al.
Veröffentlicht: (2025)
Provable Benefits of Task-Specific Prompts for In-context Learning
von: Chang, Xiangyu, et al.
Veröffentlicht: (2025)
von: Chang, Xiangyu, et al.
Veröffentlicht: (2025)
Reinforcement Learning from Human Feedback with Active Queries
von: Ji, Kaixuan, et al.
Veröffentlicht: (2024)
von: Ji, Kaixuan, et al.
Veröffentlicht: (2024)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
Variational Learning is Effective for Large Deep Networks
von: Shen, Yuesong, et al.
Veröffentlicht: (2024)
von: Shen, Yuesong, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models for Solving Rare MIP Challenges
von: Wang, Teng, et al.
Veröffentlicht: (2024)
von: Wang, Teng, et al.
Veröffentlicht: (2024)
Reward Collapse in Aligning Large Language Models
von: Song, Ziang, et al.
Veröffentlicht: (2023)
von: Song, Ziang, et al.
Veröffentlicht: (2023)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
A Statistical Framework for Data-dependent Retrieval-Augmented Models
von: Basu, Soumya, et al.
Veröffentlicht: (2024)
von: Basu, Soumya, et al.
Veröffentlicht: (2024)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
von: Kharrat, Salma, et al.
Veröffentlicht: (2024)
von: Kharrat, Salma, et al.
Veröffentlicht: (2024)
ControlAgent: Automating Control System Design via Novel Integration of LLM Agents and Domain Expertise
von: Guo, Xingang, et al.
Veröffentlicht: (2024)
von: Guo, Xingang, et al.
Veröffentlicht: (2024)
New Hybrid Fine-Tuning Paradigm for LLMs: Algorithm Design and Convergence Analysis Framework
von: Ma, Shaocong, et al.
Veröffentlicht: (2026)
von: Ma, Shaocong, et al.
Veröffentlicht: (2026)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
How Well Can Transformers Emulate In-context Newton's Method?
von: Giannou, Angeliki, et al.
Veröffentlicht: (2024)
von: Giannou, Angeliki, et al.
Veröffentlicht: (2024)
Boosting Jailbreak Attack with Momentum
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
Exploring the Robustness of In-Context Learning with Noisy Labels
von: Cheng, Chen, et al.
Veröffentlicht: (2024)
von: Cheng, Chen, et al.
Veröffentlicht: (2024)
Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
Absorber LLM: Harnessing Causal Synchronization for Test-Time Training
von: Zhang, Zhixin, et al.
Veröffentlicht: (2026)
von: Zhang, Zhixin, et al.
Veröffentlicht: (2026)
RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning
von: Wei, Zeming, et al.
Veröffentlicht: (2026)
von: Wei, Zeming, et al.
Veröffentlicht: (2026)
PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
von: Lu, Yao, et al.
Veröffentlicht: (2026)
von: Lu, Yao, et al.
Veröffentlicht: (2026)
Can Transformers Learn Optimal Filtering for Unknown Systems?
von: Balim, Haldun, et al.
Veröffentlicht: (2023)
von: Balim, Haldun, et al.
Veröffentlicht: (2023)
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
von: Kan, Kelvin, et al.
Veröffentlicht: (2025)
von: Kan, Kelvin, et al.
Veröffentlicht: (2025)
How Multimodal Integration Boost the Performance of LLM for Optimization: Case Study on Capacitated Vehicle Routing Problems
von: Huang, Yuxiao, et al.
Veröffentlicht: (2024)
von: Huang, Yuxiao, et al.
Veröffentlicht: (2024)
Data-Driven Exploration for a Class of Continuous-Time Indefinite Linear--Quadratic Reinforcement Learning Problems
von: Huang, Yilie, et al.
Veröffentlicht: (2025)
von: Huang, Yilie, et al.
Veröffentlicht: (2025)
Federated Majorize-Minimization: Beyond Parameter Aggregation
von: Dieuleveut, Aymeric, et al.
Veröffentlicht: (2025)
von: Dieuleveut, Aymeric, et al.
Veröffentlicht: (2025)
Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control
von: Skifstad, Julian, et al.
Veröffentlicht: (2026)
von: Skifstad, Julian, et al.
Veröffentlicht: (2026)
Stronger Approximation Guarantees for Non-Monotone γ-Weakly DR-Submodular Maximization
von: Jadav, Hareshkumar, et al.
Veröffentlicht: (2026)
von: Jadav, Hareshkumar, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
von: Li, Yingcong, et al.
Veröffentlicht: (2025) -
Mechanics of Next Token Prediction with Self-Attention
von: Li, Yingcong, et al.
Veröffentlicht: (2024) -
Transformers as Support Vector Machines
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2023) -
When and How Unlabeled Data Provably Improve In-Context Learning
von: Li, Yingcong, et al.
Veröffentlicht: (2025) -
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)