On the Expressive Power of Floating-Point Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Park, Sejun, Park, Yeachan, Hwang, Geonho |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Expressive Power of ReLU and Step Networks under Floating-Point Operations
por: Park, Yeachan, et al.
Publicado: (2024)
por: Park, Yeachan, et al.
Publicado: (2024)
Expressive Power of Floating-Point Neural Networks with Arbitrary Reduction Orders and Inexact Activation Implementations
por: Park, Yeachan, et al.
Publicado: (2026)
por: Park, Yeachan, et al.
Publicado: (2026)
On Expressive Power of Quantized Neural Networks under Fixed-Point Arithmetic
por: Park, Yeachan, et al.
Publicado: (2024)
por: Park, Yeachan, et al.
Publicado: (2024)
Floating-Point Networks with Automatic Differentiation Can Represent Almost All Floating-Point Functions and Their Gradients
por: Park, Sejun, et al.
Publicado: (2026)
por: Park, Sejun, et al.
Publicado: (2026)
Floating-Point Neural Networks Are Provably Robust Universal Approximators
por: Hwang, Geonho, et al.
Publicado: (2025)
por: Hwang, Geonho, et al.
Publicado: (2025)
Minimum width for universal approximation using squashable activation functions
por: Shin, Jonghyun, et al.
Publicado: (2025)
por: Shin, Jonghyun, et al.
Publicado: (2025)
Intrinsic Task Symmetry Drives Generalization in Algorithmic Tasks
por: Hwang, Hyeonbin, et al.
Publicado: (2026)
por: Hwang, Hyeonbin, et al.
Publicado: (2026)
Absence of Closed-Form Descriptions for Gradient Flow in Two-Layer Narrow Networks
por: Park, Yeachan
Publicado: (2024)
por: Park, Yeachan
Publicado: (2024)
A Kernel Perspective on Distillation-based Collaborative Learning
por: Park, Sejun, et al.
Publicado: (2024)
por: Park, Sejun, et al.
Publicado: (2024)
Deep Latent Variable Model based Vertical Federated Learning with Flexible Alignment and Labeling Scenarios
por: Hong, Kihun, et al.
Publicado: (2025)
por: Hong, Kihun, et al.
Publicado: (2025)
IMPaCT GNN: Imposing invariance with Message Passing in Chronological split Temporal Graphs
por: Park, Sejun, et al.
Publicado: (2024)
por: Park, Sejun, et al.
Publicado: (2024)
Minimum width for universal approximation using ReLU networks on compact domain
por: Kim, Namjun, et al.
Publicado: (2023)
por: Kim, Namjun, et al.
Publicado: (2023)
Acceleration of Grokking in Learning Arithmetic Operations via Kolmogorov-Arnold Representation
por: Park, Yeachan, et al.
Publicado: (2024)
por: Park, Yeachan, et al.
Publicado: (2024)
Bridging the Gap Between Molecule and Textual Descriptions via Substructure-aware Alignment
por: Park, Hyuntae, et al.
Publicado: (2025)
por: Park, Hyuntae, et al.
Publicado: (2025)
On the Expressive Power of Contextual Relations in Transformers
por: Fraiman, Demián
Publicado: (2026)
por: Fraiman, Demián
Publicado: (2026)
Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling
por: Wang, Mingze, et al.
Publicado: (2024)
por: Wang, Mingze, et al.
Publicado: (2024)
FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
por: Choi, Kanghyun, et al.
Publicado: (2025)
por: Choi, Kanghyun, et al.
Publicado: (2025)
Transformers for Green Semantic Communication: Less Energy, More Semantics
por: Mukherjee, Shubhabrata, et al.
Publicado: (2023)
por: Mukherjee, Shubhabrata, et al.
Publicado: (2023)
Transformers are Expressive, But Are They Expressive Enough for Regression?
por: Nath, Swaroop, et al.
Publicado: (2024)
por: Nath, Swaroop, et al.
Publicado: (2024)
MetaGreen: Meta-Learning Inspired Transformer Selection for Green Semantic Communication
por: Mukherjee, Shubhabrata, et al.
Publicado: (2024)
por: Mukherjee, Shubhabrata, et al.
Publicado: (2024)
Expressivity of deterministic quantum computation with one qubit
por: Kim, Yujin, et al.
Publicado: (2024)
por: Kim, Yujin, et al.
Publicado: (2024)
MolTRES: Improving Chemical Language Representation Learning for Molecular Property Prediction
por: Park, Jun-Hyung, et al.
Publicado: (2024)
por: Park, Jun-Hyung, et al.
Publicado: (2024)
The Hidden Power of Pure 16-bit Floating-Point Neural Networks
por: Yun, Juyoung, et al.
Publicado: (2023)
por: Yun, Juyoung, et al.
Publicado: (2023)
Exact Expressive Power of Transformers with Padding
por: Merrill, William, et al.
Publicado: (2025)
por: Merrill, William, et al.
Publicado: (2025)
The Expressive Power of Transformers with Chain of Thought
por: Merrill, William, et al.
Publicado: (2023)
por: Merrill, William, et al.
Publicado: (2023)
CAdam: Context-Adaptive Moment Estimation for 3D Gaussian Densification in Generative Distillation
por: Chung, SeungJeh, et al.
Publicado: (2026)
por: Chung, SeungJeh, et al.
Publicado: (2026)
On Expressive Power of Looped Transformers: Theoretical Analysis and Enhancement via Timestep Encoding
por: Xu, Kevin, et al.
Publicado: (2024)
por: Xu, Kevin, et al.
Publicado: (2024)
On The Expressive Power of GNN Derivatives
por: Eitan, Yam, et al.
Publicado: (2025)
por: Eitan, Yam, et al.
Publicado: (2025)
On the Expressive Power of Transformers for Maxout Networks and Continuous Piecewise Linear Functions
por: Gu, Linyan, et al.
Publicado: (2026)
por: Gu, Linyan, et al.
Publicado: (2026)
Expressive Power of Temporal Message Passing
por: Wałęga, Przemysław Andrzej, et al.
Publicado: (2024)
por: Wałęga, Przemysław Andrzej, et al.
Publicado: (2024)
Neural Shortest Path for Surface Reconstruction from Point Clouds
por: Park, Yesom, et al.
Publicado: (2025)
por: Park, Yesom, et al.
Publicado: (2025)
MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization
por: Kim, Daeun, et al.
Publicado: (2025)
por: Kim, Daeun, et al.
Publicado: (2025)
C2A: Client-Customized Adaptation for Parameter-Efficient Federated Learning
por: Kim, Yeachan, et al.
Publicado: (2024)
por: Kim, Yeachan, et al.
Publicado: (2024)
On the Expressive Power of GNNs to Solve Linear SDPs
por: Qian, Chendi, et al.
Publicado: (2026)
por: Qian, Chendi, et al.
Publicado: (2026)
Retrieval-Augmented Generation with Estimation of Source Reliability
por: Hwang, Jeongyeon, et al.
Publicado: (2024)
por: Hwang, Jeongyeon, et al.
Publicado: (2024)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
por: Brösamle, Moritz, et al.
Publicado: (2026)
por: Brösamle, Moritz, et al.
Publicado: (2026)
Search Your Block Floating Point Scales!
por: Gupta, Tanmaey, et al.
Publicado: (2026)
por: Gupta, Tanmaey, et al.
Publicado: (2026)
The AetherFloat Family: Block-Scale-Free Quad-Radix Floating-Point Architectures for AI Accelerators
por: Morisaki, Keita
Publicado: (2026)
por: Morisaki, Keita
Publicado: (2026)
On the Theoretical Expressive Power and the Design Space of Higher-Order Graph Transformers
por: Zhou, Cai, et al.
Publicado: (2024)
por: Zhou, Cai, et al.
Publicado: (2024)
k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS
por: De Schouwer, Jonas, et al.
Publicado: (2026)
por: De Schouwer, Jonas, et al.
Publicado: (2026)
Ejemplares similares
-
Expressive Power of ReLU and Step Networks under Floating-Point Operations
por: Park, Yeachan, et al.
Publicado: (2024) -
Expressive Power of Floating-Point Neural Networks with Arbitrary Reduction Orders and Inexact Activation Implementations
por: Park, Yeachan, et al.
Publicado: (2026) -
On Expressive Power of Quantized Neural Networks under Fixed-Point Arithmetic
por: Park, Yeachan, et al.
Publicado: (2024) -
Floating-Point Networks with Automatic Differentiation Can Represent Almost All Floating-Point Functions and Their Gradients
por: Park, Sejun, et al.
Publicado: (2026) -
Floating-Point Neural Networks Are Provably Robust Universal Approximators
por: Hwang, Geonho, et al.
Publicado: (2025)