TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
Fuente:
arXiv
Guardado en:
| Autores principales: | Kwak, Hyunseok, Lee, Kyeongwon, Min, Kyeongpil, Jung, Chaebin, Lee, Woojoo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
por: Wijeratne, Sasindu, et al.
Publicado: (2024)
por: Wijeratne, Sasindu, et al.
Publicado: (2024)
LoRA-Edge: Tensor-Train-Assisted LoRA for Practical CNN Fine-Tuning on Edge Devices
por: Kwak, Hyunseok, et al.
Publicado: (2025)
por: Kwak, Hyunseok, et al.
Publicado: (2025)
Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -- A Critical Review
por: Majeed, Ashiyana Abdul, et al.
Publicado: (2025)
por: Majeed, Ashiyana Abdul, et al.
Publicado: (2025)
CMAX-CAMEL: A Coarse-to-Fine Adaptive, Memory-Efficient, and Low-Power Edge Processor for Contrast Maximization
por: Min, Kyeongpil, et al.
Publicado: (2026)
por: Min, Kyeongpil, et al.
Publicado: (2026)
A Reliable, Time-Predictable Heterogeneous SoC for AI-Enhanced Mixed-Criticality Edge Applications
por: Garofalo, Angelo, et al.
Publicado: (2025)
por: Garofalo, Angelo, et al.
Publicado: (2025)
NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing
por: Zou, Cheng, et al.
Publicado: (2026)
por: Zou, Cheng, et al.
Publicado: (2026)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
por: Adnan, Muhammad, et al.
Publicado: (2024)
por: Adnan, Muhammad, et al.
Publicado: (2024)
Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
por: Makin, Yashasvi, et al.
Publicado: (2025)
por: Makin, Yashasvi, et al.
Publicado: (2025)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
por: Ma, Ke, et al.
Publicado: (2025)
por: Ma, Ke, et al.
Publicado: (2025)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
por: Qu, Huanyu, et al.
Publicado: (2025)
por: Qu, Huanyu, et al.
Publicado: (2025)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
por: Zhou, Zhuoshan, et al.
Publicado: (2026)
por: Zhou, Zhuoshan, et al.
Publicado: (2026)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
por: Garg, Raveesh, et al.
Publicado: (2023)
por: Garg, Raveesh, et al.
Publicado: (2023)
Next-generation Probabilistic Computing Hardware with 3D MOSAICs, Illusion Scale-up, and Co-design
por: Srimani, Tathagata, et al.
Publicado: (2024)
por: Srimani, Tathagata, et al.
Publicado: (2024)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
por: Xu, Weihong, et al.
Publicado: (2025)
por: Xu, Weihong, et al.
Publicado: (2025)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
por: Kubwimana, Benjamin, et al.
Publicado: (2025)
por: Kubwimana, Benjamin, et al.
Publicado: (2025)
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
por: Yeo, Gwangoo, et al.
Publicado: (2024)
por: Yeo, Gwangoo, et al.
Publicado: (2024)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
por: Renney, Harri, et al.
Publicado: (2026)
por: Renney, Harri, et al.
Publicado: (2026)
Tascade: Hardware Support for Atomic-free, Asynchronous and Efficient Reduction Trees
por: Orenes-Vera, Marcelo, et al.
Publicado: (2023)
por: Orenes-Vera, Marcelo, et al.
Publicado: (2023)
Switchboard: An Open-Source Framework for Modular Simulation of Large Hardware Systems
por: Herbst, Steven, et al.
Publicado: (2024)
por: Herbst, Steven, et al.
Publicado: (2024)
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
por: Peccia, Federico Nicolas, et al.
Publicado: (2024)
por: Peccia, Federico Nicolas, et al.
Publicado: (2024)
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
por: Panigrahy, Deepak, et al.
Publicado: (2026)
por: Panigrahy, Deepak, et al.
Publicado: (2026)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
por: Qiu, Tong Dong, et al.
Publicado: (2023)
por: Qiu, Tong Dong, et al.
Publicado: (2023)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
por: Feng, Weigang, et al.
Publicado: (2025)
por: Feng, Weigang, et al.
Publicado: (2025)
Dynamic Simultaneous Multithreaded Architecture
por: Ortiz-Arroyo, Daniel, et al.
Publicado: (2024)
por: Ortiz-Arroyo, Daniel, et al.
Publicado: (2024)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
por: Zhang, Chen, et al.
Publicado: (2026)
por: Zhang, Chen, et al.
Publicado: (2026)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
por: Mo, Zhiwen, et al.
Publicado: (2026)
por: Mo, Zhiwen, et al.
Publicado: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
por: Meng, William, et al.
Publicado: (2025)
por: Meng, William, et al.
Publicado: (2025)
Accelerating Transposed Convolutions on FPGA-based Edge Devices
por: Haris, Jude, et al.
Publicado: (2025)
por: Haris, Jude, et al.
Publicado: (2025)
MANOJAVAM: A Scalable, Unified FPGA Accelerator for Matrix Multiplication and Singular Value Decomposition in Principal Component Analysis
por: Ramasubramanian, Srivaths, et al.
Publicado: (2026)
por: Ramasubramanian, Srivaths, et al.
Publicado: (2026)
CRAFT: Latency and Cost-Aware Genetic-Based Framework for Node Placement in Edge-Fog Environments
por: Mahdizadeh, Soheil, et al.
Publicado: (2025)
por: Mahdizadeh, Soheil, et al.
Publicado: (2025)
Debunking the CUDA Myth Towards GPU-based AI Systems
por: Lee, Yunjae, et al.
Publicado: (2024)
por: Lee, Yunjae, et al.
Publicado: (2024)
DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages
por: Kanani, Alish, et al.
Publicado: (2026)
por: Kanani, Alish, et al.
Publicado: (2026)
Enabling Mixed criticality applications for the Versal AI-Engines
por: Sprave, Vincent, et al.
Publicado: (2026)
por: Sprave, Vincent, et al.
Publicado: (2026)
Achieving Dependability of AI Execution with Radiation Hardened Processors
por: Taquichiri, Carlos Rafael Tordoya, et al.
Publicado: (2025)
por: Taquichiri, Carlos Rafael Tordoya, et al.
Publicado: (2025)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
por: Kurzynski, Marco, et al.
Publicado: (2025)
por: Kurzynski, Marco, et al.
Publicado: (2025)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
por: McDaniel, Adam, et al.
Publicado: (2026)
por: McDaniel, Adam, et al.
Publicado: (2026)
DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2024)
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2024)
TeraPool-SDR: An 1.89TOPS 1024 RV-Cores 4MiB Shared-L1 Cluster for Next-Generation Open-Source Software-Defined Radios
por: Zhang, Yichao, et al.
Publicado: (2024)
por: Zhang, Yichao, et al.
Publicado: (2024)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
por: Li, Jiamin, et al.
Publicado: (2025)
por: Li, Jiamin, et al.
Publicado: (2025)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
por: Chu, Xiaoyu, et al.
Publicado: (2024)
por: Chu, Xiaoyu, et al.
Publicado: (2024)
Ejemplares similares
-
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
por: Wijeratne, Sasindu, et al.
Publicado: (2024) -
LoRA-Edge: Tensor-Train-Assisted LoRA for Practical CNN Fine-Tuning on Edge Devices
por: Kwak, Hyunseok, et al.
Publicado: (2025) -
Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -- A Critical Review
por: Majeed, Ashiyana Abdul, et al.
Publicado: (2025) -
CMAX-CAMEL: A Coarse-to-Fine Adaptive, Memory-Efficient, and Low-Power Edge Processor for Contrast Maximization
por: Min, Kyeongpil, et al.
Publicado: (2026) -
A Reliable, Time-Predictable Heterogeneous SoC for AI-Enhanced Mixed-Criticality Edge Applications
por: Garofalo, Angelo, et al.
Publicado: (2025)