TileLoom: Automatic Dataflow Planning for Tile-Based Languages on Spatial Dataflow Accelerators
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Wei, Bai, Zhenyu, Wang, Heru, Dangi, Pranav, Zhang, Zhiqiang, Tan, Cheng, Lan, Huiying, Wong, Weng-Fai, Mitra, Tulika |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
por: Bai, Zhenyu, et al.
Publicado: (2025)
por: Bai, Zhenyu, et al.
Publicado: (2025)
Mapping Gemma3 onto an Edge Dataflow Architecture
por: Du, Shouyu, et al.
Publicado: (2026)
por: Du, Shouyu, et al.
Publicado: (2026)
Dataflow-Oriented Classification and Performance Analysis of GPU-Accelerated Homomorphic Encryption
por: Nozaki, Ai, et al.
Publicado: (2026)
por: Nozaki, Ai, et al.
Publicado: (2026)
Accelerating Sparse DNNs Based on Tiled GEMM
por: Guo, Cong, et al.
Publicado: (2024)
por: Guo, Cong, et al.
Publicado: (2024)
Exploration of Energy and Throughput Tradeoffs for Dataflow Networks
por: Karim, Abrarul, et al.
Publicado: (2026)
por: Karim, Abrarul, et al.
Publicado: (2026)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
por: Shen, Aofeng, et al.
Publicado: (2025)
por: Shen, Aofeng, et al.
Publicado: (2025)
The Renoir Dataflow Platform: Efficient Data Processing without Complexity
por: De Martini, Luca, et al.
Publicado: (2023)
por: De Martini, Luca, et al.
Publicado: (2023)
Styx: Transactional Stateful Functions on Streaming Dataflows
por: Psarakis, Kyriakos, et al.
Publicado: (2023)
por: Psarakis, Kyriakos, et al.
Publicado: (2023)
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
por: Zheng, Size, et al.
Publicado: (2025)
por: Zheng, Size, et al.
Publicado: (2025)
Kitsune: Enabling Dataflow Execution on GPUs
por: Davies, Michael, et al.
Publicado: (2025)
por: Davies, Michael, et al.
Publicado: (2025)
Suki: Choreographed Distributed Dataflow in Rust
por: Laddad, Shadaj, et al.
Publicado: (2024)
por: Laddad, Shadaj, et al.
Publicado: (2024)
CheckMate: Evaluating Checkpointing Protocols for Streaming Dataflows
por: Siachamis, George, et al.
Publicado: (2024)
por: Siachamis, George, et al.
Publicado: (2024)
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
por: Fang, Jiahao, et al.
Publicado: (2024)
por: Fang, Jiahao, et al.
Publicado: (2024)
DEEP: Edge-based Dataflow Processing with Hybrid Docker Hub and Regional Registries
por: Mehran, Narges, et al.
Publicado: (2025)
por: Mehran, Narges, et al.
Publicado: (2025)
Stateful Entities: Object-oriented Cloud Applications as Distributed Dataflows
por: Psarakis, Kyriakos, et al.
Publicado: (2021)
por: Psarakis, Kyriakos, et al.
Publicado: (2021)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
por: Shi, Man, et al.
Publicado: (2024)
por: Shi, Man, et al.
Publicado: (2024)
Xorbits: Automating Operator Tiling for Distributed Data Science
por: Lu, Weizheng, et al.
Publicado: (2023)
por: Lu, Weizheng, et al.
Publicado: (2023)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
por: Li, Jonathan, et al.
Publicado: (2025)
por: Li, Jonathan, et al.
Publicado: (2025)
Democratizing Scalable Cloud Applications: Transactional Stateful Functions on Streaming Dataflows
por: Psarakis, Kyriakos
Publicado: (2025)
por: Psarakis, Kyriakos
Publicado: (2025)
Failure Transparency in Stateful Dataflow Systems (Technical Report)
por: Veresov, Aleksey, et al.
Publicado: (2024)
por: Veresov, Aleksey, et al.
Publicado: (2024)
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
por: Jia, Wenqi, et al.
Publicado: (2025)
por: Jia, Wenqi, et al.
Publicado: (2025)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
por: Zhu, Yu, et al.
Publicado: (2025)
por: Zhu, Yu, et al.
Publicado: (2025)
FlowUnits: Extending Dataflow for the Edge-to-Cloud Computing Continuum
por: Chini, Fabio, et al.
Publicado: (2025)
por: Chini, Fabio, et al.
Publicado: (2025)
Beyond Exascale: Dataflow Domain Translation on a Cerebras Cluster
por: Oppelstrup, Tomas, et al.
Publicado: (2025)
por: Oppelstrup, Tomas, et al.
Publicado: (2025)
An Efficient and Adaptive Watermark Detection System with Tile-based Error Correction
por: Zhong, Xinrui, et al.
Publicado: (2025)
por: Zhong, Xinrui, et al.
Publicado: (2025)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
por: Negi, Shubham, et al.
Publicado: (2025)
por: Negi, Shubham, et al.
Publicado: (2025)
DOPPLER: Dual-Policy Learning for Device Assignment in Asynchronous Dataflow Graphs
por: Yao, Xinyu, et al.
Publicado: (2025)
por: Yao, Xinyu, et al.
Publicado: (2025)
Design of A Low-Latency and Parallelizable SVD Dataflow Architecture on FPGA
por: Du, Fangqiang, et al.
Publicado: (2025)
por: Du, Fangqiang, et al.
Publicado: (2025)
DGNNFlow: A Streaming Dataflow Architecture for Real-Time Edge-based Dynamic GNN Inference in HL-LHC Trigger Systems
por: Maharaj, Davendra, et al.
Publicado: (2026)
por: Maharaj, Davendra, et al.
Publicado: (2026)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
por: Chen, Qian, et al.
Publicado: (2024)
por: Chen, Qian, et al.
Publicado: (2024)
IsoSched: Preemptive Tile Cascaded Scheduling of Multi-DNN via Subgraph Isomorphism
por: Zhao, Boran, et al.
Publicado: (2025)
por: Zhao, Boran, et al.
Publicado: (2025)
Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach
por: Zhang, Qiao, et al.
Publicado: (2025)
por: Zhang, Qiao, et al.
Publicado: (2025)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
por: Scheinert, Dominik, et al.
Publicado: (2023)
por: Scheinert, Dominik, et al.
Publicado: (2023)
Learned Cost Model for Placement on Reconfigurable Dataflow Hardware
por: Guha, Etash, et al.
Publicado: (2025)
por: Guha, Etash, et al.
Publicado: (2025)
Can Asymmetric Tile Buffering Be Beneficial?
por: Wang, Chengyue, et al.
Publicado: (2025)
por: Wang, Chengyue, et al.
Publicado: (2025)
PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel
por: Yi, Jinjun, et al.
Publicado: (2025)
por: Yi, Jinjun, et al.
Publicado: (2025)
Incremental GNN Embedding Computation on Streaming Graphs
por: Wang, Qiange, et al.
Publicado: (2026)
por: Wang, Qiange, et al.
Publicado: (2026)
DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference
por: Zhang, Yujie, et al.
Publicado: (2024)
por: Zhang, Yujie, et al.
Publicado: (2024)
DABench-LLM: Standardized and In-Depth Benchmarking of Post-Moore Dataflow AI Accelerators for LLMs
por: Hu, Ziyu, et al.
Publicado: (2025)
por: Hu, Ziyu, et al.
Publicado: (2025)
Union: An Automatic Workload Manager for Accelerating Network Simulation
por: Wang, Xin, et al.
Publicado: (2024)
por: Wang, Xin, et al.
Publicado: (2024)
Ejemplares similares
-
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
por: Bai, Zhenyu, et al.
Publicado: (2025) -
Mapping Gemma3 onto an Edge Dataflow Architecture
por: Du, Shouyu, et al.
Publicado: (2026) -
Dataflow-Oriented Classification and Performance Analysis of GPU-Accelerated Homomorphic Encryption
por: Nozaki, Ai, et al.
Publicado: (2026) -
Accelerating Sparse DNNs Based on Tiled GEMM
por: Guo, Cong, et al.
Publicado: (2024) -
Exploration of Energy and Throughput Tradeoffs for Dataflow Networks
por: Karim, Abrarul, et al.
Publicado: (2026)