Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Ke, Xie, Junfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Decentralized Network Topology Design for Task Offloading in Mobile Edge Computing
von: Ma, Ke, et al.
Veröffentlicht: (2024)
von: Ma, Ke, et al.
Veröffentlicht: (2024)
MOFCO: Mobility- and Migration-Aware Task Offloading in Three-Layer Fog Computing Environments
von: Mahdizadeh, Soheil, et al.
Veröffentlicht: (2025)
von: Mahdizadeh, Soheil, et al.
Veröffentlicht: (2025)
Optimizing Offload Performance in Heterogeneous MPSoCs
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
von: Meng, William, et al.
Veröffentlicht: (2025)
von: Meng, William, et al.
Veröffentlicht: (2025)
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
von: Pati, Suchita, et al.
Veröffentlicht: (2025)
von: Pati, Suchita, et al.
Veröffentlicht: (2025)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
von: Sirjani, Mohammad Sadegh, et al.
Veröffentlicht: (2025)
von: Sirjani, Mohammad Sadegh, et al.
Veröffentlicht: (2025)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
A Multi-Layered Distributed Computing Framework for Enhanced Edge Computing
von: Ma, Ke, et al.
Veröffentlicht: (2024)
von: Ma, Ke, et al.
Veröffentlicht: (2024)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
von: Lin, Bin, et al.
Veröffentlicht: (2024)
von: Lin, Bin, et al.
Veröffentlicht: (2024)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
Memory-Centric Computing: Solving Computing's Memory Problem
von: Mutlu, Onur, et al.
Veröffentlicht: (2025)
von: Mutlu, Onur, et al.
Veröffentlicht: (2025)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
von: Feng, Weigang, et al.
Veröffentlicht: (2025)
von: Feng, Weigang, et al.
Veröffentlicht: (2025)
Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -- A Critical Review
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
Part-time Power Measurements: nvidia-smi's Lack of Attention
von: Yang, Zeyu, et al.
Veröffentlicht: (2023)
von: Yang, Zeyu, et al.
Veröffentlicht: (2023)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
von: Wang, Zeke, et al.
Veröffentlicht: (2025)
von: Wang, Zeke, et al.
Veröffentlicht: (2025)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
A Reliable, Time-Predictable Heterogeneous SoC for AI-Enhanced Mixed-Criticality Edge Applications
von: Garofalo, Angelo, et al.
Veröffentlicht: (2025)
von: Garofalo, Angelo, et al.
Veröffentlicht: (2025)
Revisiting Computational Storage for Data Integrity and Security
von: Shi, Chao, et al.
Veröffentlicht: (2025)
von: Shi, Chao, et al.
Veröffentlicht: (2025)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
von: Mutlu, Onur, et al.
Veröffentlicht: (2024)
von: Mutlu, Onur, et al.
Veröffentlicht: (2024)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
von: Font, Martí Llopart, et al.
Veröffentlicht: (2026)
von: Font, Martí Llopart, et al.
Veröffentlicht: (2026)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
von: Punniyamurthy, Kishore, et al.
Veröffentlicht: (2023)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
von: Agrawal, Anirudha, et al.
Veröffentlicht: (2024)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
UniFormer: Unified and Efficient Transformer for Reasoning Across General and Custom Computing
von: Ran, Zhuoheng, et al.
Veröffentlicht: (2025)
von: Ran, Zhuoheng, et al.
Veröffentlicht: (2025)
BlockAMC: Scalable In-Memory Analog Matrix Computing for Solving Linear Systems
von: Pan, Lunshuai, et al.
Veröffentlicht: (2024)
von: Pan, Lunshuai, et al.
Veröffentlicht: (2024)
RapidOMS: FPGA-based Open Modification Spectral Library Searching with HD Computing
von: Pinge, Sumukh, et al.
Veröffentlicht: (2024)
von: Pinge, Sumukh, et al.
Veröffentlicht: (2024)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2023)
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2023)
Next-generation Probabilistic Computing Hardware with 3D MOSAICs, Illusion Scale-up, and Co-design
von: Srimani, Tathagata, et al.
Veröffentlicht: (2024)
von: Srimani, Tathagata, et al.
Veröffentlicht: (2024)
Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State Drives
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
Pooling Engram Conditional Memory in Large Language Models using CXL
von: Ma, Ruiyang, et al.
Veröffentlicht: (2026)
von: Ma, Ruiyang, et al.
Veröffentlicht: (2026)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
von: Qu, Huanyu, et al.
Veröffentlicht: (2025)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
von: Zou, An, et al.
Veröffentlicht: (2025)
von: Zou, An, et al.
Veröffentlicht: (2025)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
von: Zheng, Xianzhe, et al.
Veröffentlicht: (2026)
von: Zheng, Xianzhe, et al.
Veröffentlicht: (2026)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Decentralized Network Topology Design for Task Offloading in Mobile Edge Computing
von: Ma, Ke, et al.
Veröffentlicht: (2024) -
MOFCO: Mobility- and Migration-Aware Task Offloading in Three-Layer Fog Computing Environments
von: Mahdizadeh, Soheil, et al.
Veröffentlicht: (2025) -
Optimizing Offload Performance in Heterogeneous MPSoCs
von: Colagrande, Luca, et al.
Veröffentlicht: (2024) -
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
von: Meng, William, et al.
Veröffentlicht: (2025) -
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
von: Pati, Suchita, et al.
Veröffentlicht: (2025)