Unlocking the Theory Behind Scaling 1-Bit Neural Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Daliri, Majid, Song, Zhao, Yang, Chiwun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
von: Yang, Chiwun
Veröffentlicht: (2025)
von: Yang, Chiwun
Veröffentlicht: (2025)
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
von: Huang, Zekai, et al.
Veröffentlicht: (2025)
von: Huang, Zekai, et al.
Veröffentlicht: (2025)
RoPE Attention Can Be Trained in Almost Linear Time
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
On Fine-Grained I/O Complexity of Attention Backward Passes
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
Circuit Complexity Bounds for Visual Autoregressive Model
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
Towards Infinite-Long Prefix in Transformer
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective
von: Li, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2025)
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
von: Zhang, Yuyang, et al.
Veröffentlicht: (2026)
von: Zhang, Yuyang, et al.
Veröffentlicht: (2026)
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
von: Fan, Lizhou, et al.
Veröffentlicht: (2023)
von: Fan, Lizhou, et al.
Veröffentlicht: (2023)
Demystifying the unreasonable effectiveness of online alignment methods
von: Kang, Enoch Hyunwook
Veröffentlicht: (2026)
von: Kang, Enoch Hyunwook
Veröffentlicht: (2026)
Position: Scaling LLM Agents Requires Asymptotic Analysis with LLM Primitives
von: Meyerson, Elliot, et al.
Veröffentlicht: (2025)
von: Meyerson, Elliot, et al.
Veröffentlicht: (2025)
Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent
von: Chen, Bo, et al.
Veröffentlicht: (2025)
von: Chen, Bo, et al.
Veröffentlicht: (2025)
When Can We Solve the Weighted Low Rank Approximation Problem in Truly Subquadratic Time?
von: Li, Chenyang, et al.
Veröffentlicht: (2025)
von: Li, Chenyang, et al.
Veröffentlicht: (2025)
A Theory of Learning with Autoregressive Chain of Thought
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
Looped ReLU MLPs May Be All You Need as Practical Programmable Computers
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
Transformers Can Represent $n$-gram Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
von: Svete, Anej, et al.
Veröffentlicht: (2026)
von: Svete, Anej, et al.
Veröffentlicht: (2026)
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
von: Liu, Dongrui, et al.
Veröffentlicht: (2026)
von: Liu, Dongrui, et al.
Veröffentlicht: (2026)
Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation
von: Cao, Yang, et al.
Veröffentlicht: (2025)
von: Cao, Yang, et al.
Veröffentlicht: (2025)
Infinite Time Turing Machines and their Applications
von: Weerawarana, Rukmal, et al.
Veröffentlicht: (2025)
von: Weerawarana, Rukmal, et al.
Veröffentlicht: (2025)
Fundamental Limits of Crystalline Equivariant Graph Neural Networks: A Circuit Complexity Perspective
von: Cao, Yang, et al.
Veröffentlicht: (2025)
von: Cao, Yang, et al.
Veröffentlicht: (2025)
CMAT: A Multi-Agent Collaboration Tuning Framework for Enhancing Small Language Models
von: Liang, Xuechen, et al.
Veröffentlicht: (2024)
von: Liang, Xuechen, et al.
Veröffentlicht: (2024)
Journalists, Emotions, and the Introduction of Generative AI Chatbots: A Large-Scale Analysis of Tweets Before and After the Launch of ChatGPT
von: Lewis, Seth C., et al.
Veröffentlicht: (2024)
von: Lewis, Seth C., et al.
Veröffentlicht: (2024)
Inference Scaling vs Reasoning: An Empirical Analysis of Compute-Optimal LLM Problem-Solving
von: AbdElhameed, Marwan, et al.
Veröffentlicht: (2024)
von: AbdElhameed, Marwan, et al.
Veröffentlicht: (2024)
Lossless Model Compression via Joint Low-Rank Factorization Optimization
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
Mathematical Formalism for Memory Compression in Selective State Space Models
von: Bhat, Siddhanth
Veröffentlicht: (2024)
von: Bhat, Siddhanth
Veröffentlicht: (2024)
Mathematical Algorithm Design for Deep Learning under Societal and Judicial Constraints: The Algorithmic Transparency Requirement
von: Boche, Holger, et al.
Veröffentlicht: (2024)
von: Boche, Holger, et al.
Veröffentlicht: (2024)
Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
von: Hazra, Rishi, et al.
Veröffentlicht: (2025)
von: Hazra, Rishi, et al.
Veröffentlicht: (2025)
Learning to Think from Multiple Thinkers
von: Joshi, Nirmit, et al.
Veröffentlicht: (2026)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2026)
Provably Overwhelming Transformer Models with Designed Inputs
von: Stambler, Lev, et al.
Veröffentlicht: (2025)
von: Stambler, Lev, et al.
Veröffentlicht: (2025)
A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
von: Ye, Xiaowei, et al.
Veröffentlicht: (2026)
von: Ye, Xiaowei, et al.
Veröffentlicht: (2026)
On the Expressive Power and Limitations of Multi-Layer SSMs
von: Zubić, Nikola, et al.
Veröffentlicht: (2026)
von: Zubić, Nikola, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024) -
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
von: Yang, Chiwun
Veröffentlicht: (2025) -
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
von: Cao, Yang, et al.
Veröffentlicht: (2024) -
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
von: Huang, Zekai, et al.
Veröffentlicht: (2025) -
RoPE Attention Can Be Trained in Almost Linear Time
von: Cao, Yang, et al.
Veröffentlicht: (2024)