Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Jinming, Tian, Jiayi, Zhao, Yequan, Li, Hai, Zhang, Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
von: Zhang, Jinsong, et al.
Veröffentlicht: (2025)
von: Zhang, Jinsong, et al.
Veröffentlicht: (2025)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
Experimental Demonstration of an Optical Neural PDE Solver via On-Chip PINN Training
von: Zhao, Yequan, et al.
Veröffentlicht: (2025)
von: Zhao, Yequan, et al.
Veröffentlicht: (2025)
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
von: Lu, Jinming, et al.
Veröffentlicht: (2025)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
von: Huang, Wei, et al.
Veröffentlicht: (2023)
von: Huang, Wei, et al.
Veröffentlicht: (2023)
Improving Quantization with Post-Training Model Expansion
von: Franco, Giuseppe, et al.
Veröffentlicht: (2025)
von: Franco, Giuseppe, et al.
Veröffentlicht: (2025)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
von: Qiao, Ye, et al.
Veröffentlicht: (2026)
von: Qiao, Ye, et al.
Veröffentlicht: (2026)
Neural Network Quantization for Microcontrollers: A Comprehensive Survey of Methods, Platforms, and Applications
von: Abushahla, Hamza A., et al.
Veröffentlicht: (2025)
von: Abushahla, Hamza A., et al.
Veröffentlicht: (2025)
RESQ: A Unified Framework for REliability- and Security Enhancement of Quantized Deep Neural Networks
von: Mohammadi, Ali Soltan, et al.
Veröffentlicht: (2026)
von: Mohammadi, Ali Soltan, et al.
Veröffentlicht: (2026)
LEGO: Spatial Accelerator Generation and Optimization for Tensor Applications
von: Lin, Yujun, et al.
Veröffentlicht: (2025)
von: Lin, Yujun, et al.
Veröffentlicht: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
NeuralMatrix: Compute the Entire Neural Networks with Linear Matrix Operations for Efficient Inference
von: Sun, Ruiqi, et al.
Veröffentlicht: (2023)
von: Sun, Ruiqi, et al.
Veröffentlicht: (2023)
SmartQuant: CXL-based AI Model Store in Support of Runtime Configurable Weight Quantization
von: Xie, Rui, et al.
Veröffentlicht: (2024)
von: Xie, Rui, et al.
Veröffentlicht: (2024)
Hybrid JIT-CUDA Graph Optimization for Low-Latency Large Language Model Inference
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
von: Matsushima, Kosuke, et al.
Veröffentlicht: (2026)
von: Matsushima, Kosuke, et al.
Veröffentlicht: (2026)
MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing Architectures
von: Kang, Do Yeong, et al.
Veröffentlicht: (2025)
von: Kang, Do Yeong, et al.
Veröffentlicht: (2025)
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
von: Ma, Shaobo, et al.
Veröffentlicht: (2025)
von: Ma, Shaobo, et al.
Veröffentlicht: (2025)
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
von: Ma, Shaobo, et al.
Veröffentlicht: (2024)
von: Ma, Shaobo, et al.
Veröffentlicht: (2024)
LayerPipe2: Multistage Pipelining and Weight Recompute via Improved Exponential Moving Average for Training Neural Networks
von: Unnikrishnan, Nanda K., et al.
Veröffentlicht: (2025)
von: Unnikrishnan, Nanda K., et al.
Veröffentlicht: (2025)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
von: Chhugani, Jatin, et al.
Veröffentlicht: (2026)
von: Chhugani, Jatin, et al.
Veröffentlicht: (2026)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
von: Kim, Jiyoon, et al.
Veröffentlicht: (2025)
von: Kim, Jiyoon, et al.
Veröffentlicht: (2025)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
von: Gimenes, Pedro, et al.
Veröffentlicht: (2025)
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
Late Breaking Results: Quamba-SE: Soft-edge Quantizer for Activations in State Space Models
von: Chen, Yizhi, et al.
Veröffentlicht: (2026)
von: Chen, Yizhi, et al.
Veröffentlicht: (2026)
MonoSparse-CAM: Efficient Tree Model Processing via Monotonicity and Sparsity in CAMs
von: Molom-Ochir, Tergel, et al.
Veröffentlicht: (2024)
von: Molom-Ochir, Tergel, et al.
Veröffentlicht: (2024)
vTrain: A Simulation Framework for Evaluating Cost-effective and Compute-optimal Large Language Model Training
von: Bang, Jehyeon, et al.
Veröffentlicht: (2023)
von: Bang, Jehyeon, et al.
Veröffentlicht: (2023)
Heterogeneous Acceleration Pipeline for Recommendation System Training
von: Adnan, Muhammad, et al.
Veröffentlicht: (2022)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2022)
MCU-MixQ: A HW/SW Co-optimized Mixed-precision Neural Network Design Framework for MCUs
von: Gong, Junfeng, et al.
Veröffentlicht: (2024)
von: Gong, Junfeng, et al.
Veröffentlicht: (2024)
TransPlace: Transferable Circuit Global Placement via Graph Neural Network
von: Hou, Yunbo, et al.
Veröffentlicht: (2025)
von: Hou, Yunbo, et al.
Veröffentlicht: (2025)
Mirage: An RNS-Based Photonic Accelerator for DNN Training
von: Demirkiran, Cansu, et al.
Veröffentlicht: (2023)
von: Demirkiran, Cansu, et al.
Veröffentlicht: (2023)
Extending Straight-Through Estimation for Robust Neural Networks on Analog CIM Hardware
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing Units
von: Das, Arghadip, et al.
Veröffentlicht: (2025)
von: Das, Arghadip, et al.
Veröffentlicht: (2025)
LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
TRAM: Training Approximate Multiplier Structures for Low-Power AI Accelerators
von: Meng, Chang, et al.
Veröffentlicht: (2026)
von: Meng, Chang, et al.
Veröffentlicht: (2026)
PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
On-Sensor Convolutional Neural Networks with Early-Exits
von: Shalby, Hazem Hesham Yousef, et al.
Veröffentlicht: (2025)
von: Shalby, Hazem Hesham Yousef, et al.
Veröffentlicht: (2025)
HiVeGen -- Hierarchical LLM-based Verilog Generation for Scalable Chip Design
von: Tang, Jinwei, et al.
Veröffentlicht: (2024)
von: Tang, Jinwei, et al.
Veröffentlicht: (2024)
NeFT: Negative Feedback Training to Improve Robustness of Compute-In-Memory DNN Accelerators
von: Qin, Yifan, et al.
Veröffentlicht: (2023)
von: Qin, Yifan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
von: Zhang, Jinsong, et al.
Veröffentlicht: (2025) -
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
von: Tian, Jiayi, et al.
Veröffentlicht: (2025) -
Experimental Demonstration of an Optical Neural PDE Solver via On-Chip PINN Training
von: Zhao, Yequan, et al.
Veröffentlicht: (2025) -
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
von: Lu, Jinming, et al.
Veröffentlicht: (2025) -
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
von: Huang, Wei, et al.
Veröffentlicht: (2023)