Transformers as Support Vector Machines
Fuente:
arXiv
Saved in:
| Main Authors: | Tarzanagh, Davoud Ataee, Li, Yingcong, Thrampoulidis, Christos, Oymak, Samet |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
Mechanics of Next Token Prediction with Self-Attention
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
When and How Unlabeled Data Provably Improve In-Context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
Identification and Adaptive Control of Markov Jump Systems: Sample Complexity and Regret Bounds
by: Sattar, Yahya, et al.
Published: (2021)
by: Sattar, Yahya, et al.
Published: (2021)
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
A Penalty-Based Method for Communication-Efficient Decentralized Bilevel Programming
by: Nazari, Parvin, et al.
Published: (2022)
by: Nazari, Parvin, et al.
Published: (2022)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
by: Ildiz, M. Emrullah, et al.
Published: (2024)
by: Ildiz, M. Emrullah, et al.
Published: (2024)
Online Bilevel Optimization: Regret Analysis of Online Alternating Gradient Methods
by: Tarzanagh, Davoud Ataee, et al.
Published: (2022)
by: Tarzanagh, Davoud Ataee, et al.
Published: (2022)
Stochastic Regret Guarantees for Online Zeroth- and First-Order Bilevel Optimization
by: Nazari, Parvin, et al.
Published: (2025)
by: Nazari, Parvin, et al.
Published: (2025)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
A Safe Screening Rule with Bi-level Optimization of $ν$ Support Vector Machine
by: Yang, Zhiji, et al.
Published: (2024)
by: Yang, Zhiji, et al.
Published: (2024)
Kernel-Free Universum Quadratic Surface Twin Support Vector Machines for Imbalanced Data
by: Moosaei, Hossein, et al.
Published: (2024)
by: Moosaei, Hossein, et al.
Published: (2024)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
by: Huang, Xinmeng, et al.
Published: (2024)
by: Huang, Xinmeng, et al.
Published: (2024)
Solving General Natural-Language-Description Optimization Problems with Large Language Models
by: Zhang, Jihai, et al.
Published: (2024)
by: Zhang, Jihai, et al.
Published: (2024)
DiaBlo: Diagonal Blocks Are Sufficient For Finetuning
by: Gurses, Selcuk, et al.
Published: (2025)
by: Gurses, Selcuk, et al.
Published: (2025)
Reward Collapse in Aligning Large Language Models
by: Song, Ziang, et al.
Published: (2023)
by: Song, Ziang, et al.
Published: (2023)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
by: Fernando, Heshan, et al.
Published: (2024)
by: Fernando, Heshan, et al.
Published: (2024)
Causal LLM Routing: End-to-End Regret Minimization from Observational Data
by: Tsiourvas, Asterios, et al.
Published: (2025)
by: Tsiourvas, Asterios, et al.
Published: (2025)
Reinforcement Learning from Human Feedback with Active Queries
by: Ji, Kaixuan, et al.
Published: (2024)
by: Ji, Kaixuan, et al.
Published: (2024)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
by: Gautam, Tanmay, et al.
Published: (2024)
by: Gautam, Tanmay, et al.
Published: (2024)
Variational Learning is Effective for Large Deep Networks
by: Shen, Yuesong, et al.
Published: (2024)
by: Shen, Yuesong, et al.
Published: (2024)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
by: Pan, Rui, et al.
Published: (2024)
by: Pan, Rui, et al.
Published: (2024)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
by: Tao, Hongyi, et al.
Published: (2026)
by: Tao, Hongyi, et al.
Published: (2026)
Leveraging Large Language Models for Solving Rare MIP Challenges
by: Wang, Teng, et al.
Published: (2024)
by: Wang, Teng, et al.
Published: (2024)
Can Transformers Learn Optimal Filtering for Unknown Systems?
by: Balim, Haldun, et al.
Published: (2023)
by: Balim, Haldun, et al.
Published: (2023)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
by: Deora, Puneesh, et al.
Published: (2025)
by: Deora, Puneesh, et al.
Published: (2025)
Truncated Matrix Completion - An Empirical Study
by: Naik, Rishhabh, et al.
Published: (2025)
by: Naik, Rishhabh, et al.
Published: (2025)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
by: Kharrat, Salma, et al.
Published: (2024)
by: Kharrat, Salma, et al.
Published: (2024)
ControlAgent: Automating Control System Design via Novel Integration of LLM Agents and Domain Expertise
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
by: Garrod, Connall, et al.
Published: (2025)
by: Garrod, Connall, et al.
Published: (2025)
On the Optimization and Generalization of Multi-head Attention
by: Deora, Puneesh, et al.
Published: (2023)
by: Deora, Puneesh, et al.
Published: (2023)
Absorber LLM: Harnessing Causal Synchronization for Test-Time Training
by: Zhang, Zhixin, et al.
Published: (2026)
by: Zhang, Zhixin, et al.
Published: (2026)
Secure LLM Fine-Tuning via Safety-Aware Probing
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
Boosting Jailbreak Attack with Momentum
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
by: Zhang, Zhixin, et al.
Published: (2025)
by: Zhang, Zhixin, et al.
Published: (2025)
RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
Exploring the Robustness of In-Context Learning with Noisy Labels
by: Cheng, Chen, et al.
Published: (2024)
by: Cheng, Chen, et al.
Published: (2024)
Similar Items
-
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
by: Li, Yingcong, et al.
Published: (2025) -
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
by: Li, Yingcong, et al.
Published: (2024) -
Mechanics of Next Token Prediction with Self-Attention
by: Li, Yingcong, et al.
Published: (2024) -
When and How Unlabeled Data Provably Improve In-Context Learning
by: Li, Yingcong, et al.
Published: (2025) -
Identification and Adaptive Control of Markov Jump Systems: Sample Complexity and Regret Bounds
by: Sattar, Yahya, et al.
Published: (2021)