Saved in:
| Main Authors: | Sunaga, Kazuki, Sugiura, Keisuke, Matsutani, Hiroki |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2312.15138 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InstantFT: An FPGA-Based Runtime Subsecond Fine-tuning of CNN Models
by: Sugiura, Keisuke, et al.
Published: (2025)
by: Sugiura, Keisuke, et al.
Published: (2025)
A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODE
by: Okubo, Ikumi, et al.
Published: (2024)
by: Okubo, Ikumi, et al.
Published: (2024)
ElasticZO: A Memory-Efficient On-Device Learning with Combined Zeroth- and First-Order Optimization
by: Sugiura, Keisuke, et al.
Published: (2025)
by: Sugiura, Keisuke, et al.
Published: (2025)
FPGA-Accelerated Correspondence-free Point Cloud Registration with PointNet Features
by: Sugiura, Keisuke, et al.
Published: (2024)
by: Sugiura, Keisuke, et al.
Published: (2024)
Skip2-LoRA: A Lightweight On-device DNN Fine-tuning Method for Low-cost Edge Devices
by: Matsutani, Hiroki, et al.
Published: (2024)
by: Matsutani, Hiroki, et al.
Published: (2024)
PointODE: Lightweight Point Cloud Learning with Neural Ordinary Differential Equations on Edge
by: Sugiura, Keisuke, et al.
Published: (2025)
by: Sugiura, Keisuke, et al.
Published: (2025)
Accelerating Local LLMs on Resource-Constrained Edge Devices via Distributed Prompt Caching
by: Matsutani, Hiroki, et al.
Published: (2026)
by: Matsutani, Hiroki, et al.
Published: (2026)
A Tiny Supervised ODL Core with Auto Data Pruning for Human Activity Recognition
by: Matsutani, Hiroki, et al.
Published: (2024)
by: Matsutani, Hiroki, et al.
Published: (2024)
Sinkhorn Algorithm for Sequentially Composed Optimal Transports
by: Watanabe, Kazuki, et al.
Published: (2024)
by: Watanabe, Kazuki, et al.
Published: (2024)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
by: Matsutani, Kohsei, et al.
Published: (2026)
by: Matsutani, Kohsei, et al.
Published: (2026)
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
by: Shao, Haikuo, et al.
Published: (2024)
by: Shao, Haikuo, et al.
Published: (2024)
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
by: Kawata, Ryotaro, et al.
Published: (2025)
by: Kawata, Ryotaro, et al.
Published: (2025)
An FPGA-Based Accelerator Enabling Efficient Support for CNNs with Arbitrary Kernel Sizes
by: Wang, Miaoxin, et al.
Published: (2024)
by: Wang, Miaoxin, et al.
Published: (2024)
Accelerating Storage-Based Training for Graph Neural Networks
by: Jang, Myung-Hwan, et al.
Published: (2026)
by: Jang, Myung-Hwan, et al.
Published: (2026)
Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
by: Hajimolahoseini, Habib, et al.
Published: (2023)
by: Hajimolahoseini, Habib, et al.
Published: (2023)
Accelerating Large Language Model Training with 4D Parallelism and Memory Consumption Estimator
by: Fujii, Kazuki, et al.
Published: (2024)
by: Fujii, Kazuki, et al.
Published: (2024)
Exploring Parallelism in FPGA-Based Accelerators for Machine Learning Applications
by: Centeno, Sed, et al.
Published: (2025)
by: Centeno, Sed, et al.
Published: (2025)
Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer
by: Li, Richie, et al.
Published: (2025)
by: Li, Richie, et al.
Published: (2025)
Accelerating Recommender Model Training by Dynamically Skipping Stale Embeddings
by: Maboud, Yassaman Ebrahimzadeh, et al.
Published: (2024)
by: Maboud, Yassaman Ebrahimzadeh, et al.
Published: (2024)
BitLogic: Training Framework for Gradient-Based FPGA-Native Neural Networks
by: Bührer, Simon, et al.
Published: (2026)
by: Bührer, Simon, et al.
Published: (2026)
Real-Time Cell Sorting with Scalable In Situ FPGA-Accelerated Deep Learning
by: Islam, Khayrul, et al.
Published: (2025)
by: Islam, Khayrul, et al.
Published: (2025)
Metalearning Continual Learning Algorithms
by: Irie, Kazuki, et al.
Published: (2023)
by: Irie, Kazuki, et al.
Published: (2023)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
by: Chen, Hongzheng, et al.
Published: (2023)
by: Chen, Hongzheng, et al.
Published: (2023)
Sequential-Parallel Duality in Prefix Scannable Models
by: Yau, Morris, et al.
Published: (2025)
by: Yau, Morris, et al.
Published: (2025)
Hardware-Accelerated Event-Graph Neural Networks for Low-Latency Time-Series Classification on SoC FPGA
by: Nakano, Hiroshi, et al.
Published: (2025)
by: Nakano, Hiroshi, et al.
Published: (2025)
ProTEA: Programmable Transformer Encoder Acceleration on FPGA
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
Accelerating Transposed Convolutions on FPGA-based Edge Devices
by: Haris, Jude, et al.
Published: (2025)
by: Haris, Jude, et al.
Published: (2025)
Practicable Black-box Evasion Attacks on Link Prediction in Dynamic Graphs -- A Graph Sequential Embedding Method
by: Li, Jiate, et al.
Published: (2024)
by: Li, Jiate, et al.
Published: (2024)
EmbedPart: Embedding-Driven Graph Partitioning for Scalable Graph Neural Network Training
by: Merkel, Nikolai, et al.
Published: (2026)
by: Merkel, Nikolai, et al.
Published: (2026)
Sequential Regression Learning with Randomized Algorithms
by: Leão, Dorival, et al.
Published: (2025)
by: Leão, Dorival, et al.
Published: (2025)
StarMAP: Global Neighbor Embedding for Faithful Data Visualization
by: Watanabe, Koshi, et al.
Published: (2025)
by: Watanabe, Koshi, et al.
Published: (2025)
FPGA-based Acceleration for Convolutional Neural Networks: A Comprehensive Review
by: Jiang, Junye, et al.
Published: (2025)
by: Jiang, Junye, et al.
Published: (2025)
Integer-only Quantized Transformers for Embedded FPGA-based Time-series Forecasting in AIoT
by: Ling, Tianheng, et al.
Published: (2024)
by: Ling, Tianheng, et al.
Published: (2024)
FPGA-Based Neural Thrust Controller for UAVs
by: Azem, Sharif, et al.
Published: (2024)
by: Azem, Sharif, et al.
Published: (2024)
A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
by: Gupta, Neelesh, et al.
Published: (2026)
by: Gupta, Neelesh, et al.
Published: (2026)
GCoD: Graph Convolutional Network Acceleration via Dedicated Algorithm and Accelerator Co-Design
by: You, Haoran, et al.
Published: (2021)
by: You, Haoran, et al.
Published: (2021)
PEFSL: A deployment Pipeline for Embedded Few-Shot Learning on a FPGA SoC
by: Ribeiro, Lucas Grativol, et al.
Published: (2024)
by: Ribeiro, Lucas Grativol, et al.
Published: (2024)
Machine Learning-Based Graph Simplification for Symbolic Accelerators
by: Yu, Tiffany, et al.
Published: (2026)
by: Yu, Tiffany, et al.
Published: (2026)
Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
by: Fujii, Kazuki, et al.
Published: (2024)
by: Fujii, Kazuki, et al.
Published: (2024)
Accelerated Sequential Flow Matching: A Bayesian Filtering Perspective
by: Huang, Yinan, et al.
Published: (2026)
by: Huang, Yinan, et al.
Published: (2026)
Similar Items
-
InstantFT: An FPGA-Based Runtime Subsecond Fine-tuning of CNN Models
by: Sugiura, Keisuke, et al.
Published: (2025) -
A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODE
by: Okubo, Ikumi, et al.
Published: (2024) -
ElasticZO: A Memory-Efficient On-Device Learning with Combined Zeroth- and First-Order Optimization
by: Sugiura, Keisuke, et al.
Published: (2025) -
FPGA-Accelerated Correspondence-free Point Cloud Registration with PointNet Features
by: Sugiura, Keisuke, et al.
Published: (2024) -
Skip2-LoRA: A Lightweight On-device DNN Fine-tuning Method for Low-cost Edge Devices
by: Matsutani, Hiroki, et al.
Published: (2024)