Lower bounds on transformers with infinite precision
Fuente:
arXiv
Salvato in:
| Autore principale: | Kozachinskiy, Alexander |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A completely uniform transformer for parity
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2025)
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2025)
Optimal bounds for dissatisfaction in perpetual voting
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2024)
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2024)
Simple online learning with consistent oracle
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2023)
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2023)
Ehrenfeucht-Haussler Rank and Chain of Thought
di: Barceló, Pablo, et al.
Pubblicazione: (2025)
di: Barceló, Pablo, et al.
Pubblicazione: (2025)
Parity, Sensitivity, and Transformers
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2026)
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2026)
Explaining k-Nearest Neighbors: Abductive and Counterfactual Explanations
di: Barceló, Pablo, et al.
Pubblicazione: (2025)
di: Barceló, Pablo, et al.
Pubblicazione: (2025)
Message Passing on the Edge: Towards Scalable and Expressive GNNs
di: Barceló, Pablo, et al.
Pubblicazione: (2025)
di: Barceló, Pablo, et al.
Pubblicazione: (2025)
Continuity and Isolation Lead to Doubts or Dilemmas in Large Language Models
di: Pasten, Hector, et al.
Pubblicazione: (2025)
di: Pasten, Hector, et al.
Pubblicazione: (2025)
Language Generation: Complexity Barriers and Implications for Learning
di: Arenas, Marcelo, et al.
Pubblicazione: (2025)
di: Arenas, Marcelo, et al.
Pubblicazione: (2025)
Strassen Attention, Split VC Dimension and Compositionality in Transformers
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2025)
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2025)
Estimation of Energy-dissipation Lower-bounds for Neuromorphic Learning-in-memory
di: Chen, Zihao, et al.
Pubblicazione: (2024)
di: Chen, Zihao, et al.
Pubblicazione: (2024)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
di: Xu, Yinggan, et al.
Pubblicazione: (2026)
di: Xu, Yinggan, et al.
Pubblicazione: (2026)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
di: Goodall, Alexander W., et al.
Pubblicazione: (2025)
di: Goodall, Alexander W., et al.
Pubblicazione: (2025)
On the inductive bias of infinite-depth ResNets and the bottleneck rank
di: Boix-Adsera, Enric
Pubblicazione: (2025)
di: Boix-Adsera, Enric
Pubblicazione: (2025)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
di: Shah, Jay, et al.
Pubblicazione: (2024)
di: Shah, Jay, et al.
Pubblicazione: (2024)
LoKA: Low-precision Kernel Applications for Recommendation Models At Scale
di: Luo, Liang, et al.
Pubblicazione: (2026)
di: Luo, Liang, et al.
Pubblicazione: (2026)
Upper Entropy for 2-Monotone Lower Probabilities
di: Vu, Tuan-Anh, et al.
Pubblicazione: (2026)
di: Vu, Tuan-Anh, et al.
Pubblicazione: (2026)
Carrying over algorithm in transformers
di: Kruthoff, Jorrit
Pubblicazione: (2024)
di: Kruthoff, Jorrit
Pubblicazione: (2024)
Tight Lower Bounds and Improved Convergence in Performative Prediction
di: Khorsandi, Pedram, et al.
Pubblicazione: (2024)
di: Khorsandi, Pedram, et al.
Pubblicazione: (2024)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
di: He, Jiafan, et al.
Pubblicazione: (2025)
di: He, Jiafan, et al.
Pubblicazione: (2025)
Small transformer architectures for task switching
di: Gros, Claudius
Pubblicazione: (2025)
di: Gros, Claudius
Pubblicazione: (2025)
AgGym: An agricultural biotic stress simulation environment for ultra-precision management planning
di: Khosravi, Mahsa, et al.
Pubblicazione: (2024)
di: Khosravi, Mahsa, et al.
Pubblicazione: (2024)
MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
di: Duanmu, Haojie, et al.
Pubblicazione: (2025)
di: Duanmu, Haojie, et al.
Pubblicazione: (2025)
Step-resolved data attribution for looped transformers
di: Kaissis, Georgios, et al.
Pubblicazione: (2026)
di: Kaissis, Georgios, et al.
Pubblicazione: (2026)
Weight-sparse transformers have interpretable circuits
di: Gao, Leo, et al.
Pubblicazione: (2025)
di: Gao, Leo, et al.
Pubblicazione: (2025)
Upper and Lower Bounds for Distributionally Robust Off-Dynamics Reinforcement Learning
di: Liu, Zhishuai, et al.
Pubblicazione: (2024)
di: Liu, Zhishuai, et al.
Pubblicazione: (2024)
Model-based Offline Reinforcement Learning with Lower Expectile Q-Learning
di: Park, Kwanyoung, et al.
Pubblicazione: (2024)
di: Park, Kwanyoung, et al.
Pubblicazione: (2024)
Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers
di: Tang, Mohan, et al.
Pubblicazione: (2026)
di: Tang, Mohan, et al.
Pubblicazione: (2026)
Learning the greatest common divisor: explaining transformer predictions
di: Charton, François
Pubblicazione: (2023)
di: Charton, François
Pubblicazione: (2023)
Incorporating Domain Differential Equations into Graph Convolutional Networks to Lower Generalization Discrepancy
di: Sun, Yue, et al.
Pubblicazione: (2024)
di: Sun, Yue, et al.
Pubblicazione: (2024)
Bitformer: An efficient Transformer with bitwise operation-based attention for Big Data Analytics at low-cost low-precision devices
di: Duan, Gaoxiang, et al.
Pubblicazione: (2023)
di: Duan, Gaoxiang, et al.
Pubblicazione: (2023)
Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic Bandits
di: Xue, Bo, et al.
Pubblicazione: (2025)
di: Xue, Bo, et al.
Pubblicazione: (2025)
Arnold: a generalist muscle transformer policy
di: Chiappa, Alberto Silvio, et al.
Pubblicazione: (2025)
di: Chiappa, Alberto Silvio, et al.
Pubblicazione: (2025)
Static and multivariate-temporal attentive fusion transformer for readmission risk prediction
di: Sun, Zhe, et al.
Pubblicazione: (2024)
di: Sun, Zhe, et al.
Pubblicazione: (2024)
Don't be lazy: CompleteP enables compute-efficient deep transformers
di: Dey, Nolan, et al.
Pubblicazione: (2025)
di: Dey, Nolan, et al.
Pubblicazione: (2025)
Financial time series augmentation using transformer based GAN architecture
di: Podobiński, Andrzej, et al.
Pubblicazione: (2026)
di: Podobiński, Andrzej, et al.
Pubblicazione: (2026)
A Lower Bound for the Number of Linear Regions of Ternary ReLU Regression Neural Networks
di: Nakahara, Yuta, et al.
Pubblicazione: (2025)
di: Nakahara, Yuta, et al.
Pubblicazione: (2025)
Structural Positional Encoding for knowledge integration in transformer-based medical process monitoring
di: Irwin, Christopher, et al.
Pubblicazione: (2024)
di: Irwin, Christopher, et al.
Pubblicazione: (2024)
$σ$-PCA: a building block for neural learning of identifiable linear transformations
di: Kanavati, Fahdi, et al.
Pubblicazione: (2023)
di: Kanavati, Fahdi, et al.
Pubblicazione: (2023)
Decomposition-based multi-scale transformer framework for time series anomaly detection
di: Zhang, Wenxin, et al.
Pubblicazione: (2025)
di: Zhang, Wenxin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A completely uniform transformer for parity
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2025) -
Optimal bounds for dissatisfaction in perpetual voting
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2024) -
Simple online learning with consistent oracle
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2023) -
Ehrenfeucht-Haussler Rank and Chain of Thought
di: Barceló, Pablo, et al.
Pubblicazione: (2025) -
Parity, Sensitivity, and Transformers
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2026)