Design of A Low-Latency and Parallelizable SVD Dataflow Architecture on FPGA
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Fangqiang, Chong, Sixuan, Huang, Zixuan, Qin, Rui, Mi, Fengnan, Hu, Caibao, Chen, Jiangang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimal Expert Selection for Distributed Mixture-of-Experts at the Wireless Edge
by: Qin, Shengling, et al.
Published: (2025)
by: Qin, Shengling, et al.
Published: (2025)
Mapping Gemma3 onto an Edge Dataflow Architecture
by: Du, Shouyu, et al.
Published: (2026)
by: Du, Shouyu, et al.
Published: (2026)
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
by: Jia, Wenqi, et al.
Published: (2025)
by: Jia, Wenqi, et al.
Published: (2025)
Mangrove: Fast and Parallelizable State Replication for Blockchains
by: Paramonov, Anton, et al.
Published: (2025)
by: Paramonov, Anton, et al.
Published: (2025)
Reconfigurable Intelligent Computational Surfaces for MEC-Assisted Autonomous Driving Networks: Design Optimization and Analysis
by: Zhang, Xueyao, et al.
Published: (2024)
by: Zhang, Xueyao, et al.
Published: (2024)
Efficiently Parallelizable Strassen-Based Multiplication of a Matrix by its Transpose
by: Arrigoni, Viviana, et al.
Published: (2021)
by: Arrigoni, Viviana, et al.
Published: (2021)
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
by: Zheng, Ce, et al.
Published: (2026)
by: Zheng, Ce, et al.
Published: (2026)
Reconfigurable Intelligent Computational Surfaces for MEC-Assisted Autonomous Driving Networks
by: Yang, Bo, et al.
Published: (2024)
by: Yang, Bo, et al.
Published: (2024)
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
by: Berg, Benjamin, et al.
Published: (2024)
by: Berg, Benjamin, et al.
Published: (2024)
Exploration of Energy and Throughput Tradeoffs for Dataflow Networks
by: Karim, Abrarul, et al.
Published: (2026)
by: Karim, Abrarul, et al.
Published: (2026)
LAAFD: LLM-based Agents for Accelerated FPGA Design
by: Moraru, Maxim, et al.
Published: (2026)
by: Moraru, Maxim, et al.
Published: (2026)
DGNNFlow: A Streaming Dataflow Architecture for Real-Time Edge-based Dynamic GNN Inference in HL-LHC Trigger Systems
by: Maharaj, Davendra, et al.
Published: (2026)
by: Maharaj, Davendra, et al.
Published: (2026)
New Improvements in Solving Large LABS Instances Using Massively Parallelizable Memetic Tabu Search
by: Zhang, Zhiwei, et al.
Published: (2025)
by: Zhang, Zhiwei, et al.
Published: (2025)
APWA: A Distributed Architecture for Parallelizable Agentic Workflows
by: Rose, Evan, et al.
Published: (2026)
by: Rose, Evan, et al.
Published: (2026)
IRS Aided Federated Learning: Multiple Access and Fundamental Tradeoff
by: Chen, Guangji, et al.
Published: (2024)
by: Chen, Guangji, et al.
Published: (2024)
An Experimental Exploration of In-Memory Computing for Multi-Layer Perceptrons
by: Carrinho, Pedro, et al.
Published: (2025)
by: Carrinho, Pedro, et al.
Published: (2025)
RASC: Region-Aware Self-Calibration for Dense 2D Sensor Arrays
by: Ma, Yinglei, et al.
Published: (2026)
by: Ma, Yinglei, et al.
Published: (2026)
Computation Offloading Strategies in Integrated Terrestrial and Non-Terrestrial Networks
by: Mohsin, Muhammad Ahmed, et al.
Published: (2025)
by: Mohsin, Muhammad Ahmed, et al.
Published: (2025)
PowerTrain: Fast, Generalizable Time and Power Prediction Models to Optimize DNN Training on Accelerated Edges
by: K., Prashanthi S., et al.
Published: (2024)
by: K., Prashanthi S., et al.
Published: (2024)
Clustering-Based User Selection in Federated Learning: Metadata Exploitation for 3GPP Networks
by: Zheng, Ce, et al.
Published: (2026)
by: Zheng, Ce, et al.
Published: (2026)
Communication-Efficient Distributed On-Device LLM Inference Over Wireless Networks
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Parallel-in-Time Kalman Smoothing Using Orthogonal Transformations
by: Gargir, Shahaf, et al.
Published: (2025)
by: Gargir, Shahaf, et al.
Published: (2025)
Over-the-Air Federated Learning with Phase Noise: Analysis and Countermeasures
by: Dahl, Martin, et al.
Published: (2024)
by: Dahl, Martin, et al.
Published: (2024)
Revisiting Speculative Leaderless Protocols for Low-Latency BFT Replication
by: Qian, Daniel, et al.
Published: (2026)
by: Qian, Daniel, et al.
Published: (2026)
Low Latency, High Bandwidth Streaming of Experimental Data with EJFAT
by: Baldin, Ilya, et al.
Published: (2025)
by: Baldin, Ilya, et al.
Published: (2025)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
by: Wu, Siyu, et al.
Published: (2025)
by: Wu, Siyu, et al.
Published: (2025)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
by: Zhu, Yu, et al.
Published: (2025)
by: Zhu, Yu, et al.
Published: (2025)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
by: Li, Zixuan, et al.
Published: (2026)
by: Li, Zixuan, et al.
Published: (2026)
The Renoir Dataflow Platform: Efficient Data Processing without Complexity
by: De Martini, Luca, et al.
Published: (2023)
by: De Martini, Luca, et al.
Published: (2023)
Satellite Federated Edge Learning: Architecture Design and Convergence Analysis
by: Shi, Yuanming, et al.
Published: (2024)
by: Shi, Yuanming, et al.
Published: (2024)
Dataflow-Oriented Classification and Performance Analysis of GPU-Accelerated Homomorphic Encryption
by: Nozaki, Ai, et al.
Published: (2026)
by: Nozaki, Ai, et al.
Published: (2026)
Real-Time Diagnostic Integrity Meets Efficiency: A Novel Platform-Agnostic Architecture for Physiological Signal Compression
by: Vora, Neel R, et al.
Published: (2023)
by: Vora, Neel R, et al.
Published: (2023)
BBCA-CHAIN: Low Latency, High Throughput BFT Consensus on a DAG
by: Malkhi, Dahlia, et al.
Published: (2023)
by: Malkhi, Dahlia, et al.
Published: (2023)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
by: Hidayetoglu, Mert, et al.
Published: (2025)
by: Hidayetoglu, Mert, et al.
Published: (2025)
Low-Latency Layer-Aware Proactive and Passive Container Migration in Meta Computing
by: Liu, Mengjie, et al.
Published: (2024)
by: Liu, Mengjie, et al.
Published: (2024)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
by: Yu, Minchen, et al.
Published: (2023)
by: Yu, Minchen, et al.
Published: (2023)
Chasing the Speed of Light: Low-Latency Planetary-Scale Adaptive Byzantine Consensus
by: Berger, Christian, et al.
Published: (2023)
by: Berger, Christian, et al.
Published: (2023)
DEEP: Edge-based Dataflow Processing with Hybrid Docker Hub and Regional Registries
by: Mehran, Narges, et al.
Published: (2025)
by: Mehran, Narges, et al.
Published: (2025)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
by: Wang, Wenfeng, et al.
Published: (2026)
by: Wang, Wenfeng, et al.
Published: (2026)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
by: Yuan, Yitao, et al.
Published: (2025)
by: Yuan, Yitao, et al.
Published: (2025)
Similar Items
-
Optimal Expert Selection for Distributed Mixture-of-Experts at the Wireless Edge
by: Qin, Shengling, et al.
Published: (2025) -
Mapping Gemma3 onto an Edge Dataflow Architecture
by: Du, Shouyu, et al.
Published: (2026) -
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
by: Jia, Wenqi, et al.
Published: (2025) -
Mangrove: Fast and Parallelizable State Replication for Blockchains
by: Paramonov, Anton, et al.
Published: (2025) -
Reconfigurable Intelligent Computational Surfaces for MEC-Assisted Autonomous Driving Networks: Design Optimization and Analysis
by: Zhang, Xueyao, et al.
Published: (2024)