Training Transformers in Cosine Coefficient Space
Fuente:
arXiv
Saved in:
| Main Author: | Bergach, Mohamed Amine |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Shortest-Path FFT: Optimal SIMD Instruction Scheduling via Graph Search
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Dual-Select FMA Butterfly for FFT: Eliminating Twiddle Factor Singularities with Bounded Precomputed Ratios
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
From 8 Seconds to 370ms: Kernel-Fused SAR Imaging on Apple Silicon via Single-Dispatch FFT Pipelines
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
EcoSpa: Efficient Transformer Training with Coupled Sparsity
by: Xiao, Jinqi, et al.
Published: (2025)
by: Xiao, Jinqi, et al.
Published: (2025)
Stochastic Q-learning for Large Discrete Action Spaces
by: Fourati, Fares, et al.
Published: (2024)
by: Fourati, Fares, et al.
Published: (2024)
Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
by: Liu, Yi
Published: (2026)
by: Liu, Yi
Published: (2026)
LLM-Driven Design Space Exploration of FPGA-based Accelerators
by: Sharma, Vinamra, et al.
Published: (2026)
by: Sharma, Vinamra, et al.
Published: (2026)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
by: Yi, Qingao, et al.
Published: (2025)
by: Yi, Qingao, et al.
Published: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
by: Liu, Guangda, et al.
Published: (2024)
by: Liu, Guangda, et al.
Published: (2024)
The Power of Training: How Different Neural Network Setups Influence the Energy Demand
by: Geißler, Daniel, et al.
Published: (2024)
by: Geißler, Daniel, et al.
Published: (2024)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
by: Kermani, Arshia, et al.
Published: (2025)
by: Kermani, Arshia, et al.
Published: (2025)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
by: Qiao, Liang, et al.
Published: (2025)
by: Qiao, Liang, et al.
Published: (2025)
KWT-Tiny: RISC-V Accelerated, Embedded Keyword Spotting Transformer
by: Al-Qawlaq, Aness, et al.
Published: (2024)
by: Al-Qawlaq, Aness, et al.
Published: (2024)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2025)
by: Chiang, Hung-Yueh, et al.
Published: (2025)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
by: Yin, Yishu, et al.
Published: (2025)
by: Yin, Yishu, et al.
Published: (2025)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
by: Wu, Wenhao, et al.
Published: (2026)
by: Wu, Wenhao, et al.
Published: (2026)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
by: Zhao, Yushang, et al.
Published: (2025)
by: Zhao, Yushang, et al.
Published: (2025)
Reliability by design: quantifying and eliminating fabrication risk in LLMs. From generative to consultative AI: a comparative analysis in the legal domain and lessons for high-stakes knowledge bases
by: Dantart, Alex
Published: (2026)
by: Dantart, Alex
Published: (2026)
SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference
by: Cavagna, Hiari Pizzini, et al.
Published: (2026)
by: Cavagna, Hiari Pizzini, et al.
Published: (2026)
Learning, Potential, and Retention: An Approach for Evaluating Adaptive AI-Enabled Medical Devices
by: Burgon, Alexis, et al.
Published: (2026)
by: Burgon, Alexis, et al.
Published: (2026)
Improving LLM Performance Through Black-Box Online Tuning: A Case for Adding System Specs to Factsheets for Trusted AI
by: Atinafu, Yonas, et al.
Published: (2026)
by: Atinafu, Yonas, et al.
Published: (2026)
Personalized Model-Based Design of Human Centric AI enabled CPS for Long term usage
by: Ngabonziza, Bernard, et al.
Published: (2026)
by: Ngabonziza, Bernard, et al.
Published: (2026)
Impact of Data-Oriented and Object-Oriented Design on Performance and Cache Utilization with Artificial Intelligence Algorithms in Multi-Threaded CPUs
by: Arantes, Gabriel M., et al.
Published: (2025)
by: Arantes, Gabriel M., et al.
Published: (2025)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
by: Ivanov, Andrei, et al.
Published: (2025)
by: Ivanov, Andrei, et al.
Published: (2025)
Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
by: Prieto, Pablo, et al.
Published: (2025)
by: Prieto, Pablo, et al.
Published: (2025)
PixLift: Accelerating Web Browsing via AI Upscaling
by: Atinafu, Yonas, et al.
Published: (2025)
by: Atinafu, Yonas, et al.
Published: (2025)
Tiny-QMoE
by: Cashman, Jack, et al.
Published: (2025)
by: Cashman, Jack, et al.
Published: (2025)
Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
by: Pfister, Rolf, et al.
Published: (2025)
by: Pfister, Rolf, et al.
Published: (2025)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
by: Chai, Yuji, et al.
Published: (2025)
by: Chai, Yuji, et al.
Published: (2025)
Faster LLM Inference using DBMS-Inspired Preemption and Cache Replacement Policies
by: Kim, Kyoungmin, et al.
Published: (2024)
by: Kim, Kyoungmin, et al.
Published: (2024)
XTC, A Research Platform for Optimizing AI Workload Operators
by: Hugo, Pompougnac, et al.
Published: (2025)
by: Hugo, Pompougnac, et al.
Published: (2025)
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
by: Cai, Yanan, et al.
Published: (2025)
by: Cai, Yanan, et al.
Published: (2025)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
by: Liu, Xiaoxuan, et al.
Published: (2024)
by: Liu, Xiaoxuan, et al.
Published: (2024)
WANDER: An Explainable Decision-Support Framework for HPC
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
by: Zhao, Youpeng, et al.
Published: (2025)
by: Zhao, Youpeng, et al.
Published: (2025)
ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
DeepContext: A Context-aware, Cross-platform, and Cross-framework Tool for Performance Profiling and Analysis of Deep Learning Workloads
by: Zhao, Qidong, et al.
Published: (2024)
by: Zhao, Qidong, et al.
Published: (2024)
Similar Items
-
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026) -
Shortest-Path FFT: Optimal SIMD Instruction Scheduling via Graph Search
by: Bergach, Mohamed Amine
Published: (2026) -
Dual-Select FMA Butterfly for FFT: Eliminating Twiddle Factor Singularities with Bounded Precomputed Ratios
by: Bergach, Mohamed Amine
Published: (2026) -
From 8 Seconds to 370ms: Kernel-Fused SAR Imaging on Apple Silicon via Single-Dispatch FFT Pipelines
by: Bergach, Mohamed Amine
Published: (2026) -
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)