t-READi: Transformer-Powered Robust and Efficient Multimodal Inference for Autonomous Driving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Pengfei, Qian, Yuhang, Zheng, Tianyue, Li, Ang, Chen, Zhe, Gao, Yue, Cheng, Xiuzhen, Luo, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
PIR-DSN: A Decentralized Storage Network Supporting Private Information Retrieval
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
von: Yang, Zheming, et al.
Veröffentlicht: (2026)
von: Yang, Zheming, et al.
Veröffentlicht: (2026)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
von: Ma, Bin, et al.
Veröffentlicht: (2026)
von: Ma, Bin, et al.
Veröffentlicht: (2026)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
OrbitBFT: Enabling Scalable and Robust BFT Consensus in LEO Constellations
von: Sun, Tianyi, et al.
Veröffentlicht: (2026)
von: Sun, Tianyi, et al.
Veröffentlicht: (2026)
Power Aware Dynamic Reallocation For Inference
von: Jiang, Yiwei, et al.
Veröffentlicht: (2026)
von: Jiang, Yiwei, et al.
Veröffentlicht: (2026)
SkyMemory: A LEO Edge Cache for Transformer Inference Optimization and Scale Out
von: Sandholm, Thomas, et al.
Veröffentlicht: (2025)
von: Sandholm, Thomas, et al.
Veröffentlicht: (2025)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
von: Yang, Zheming, et al.
Veröffentlicht: (2025)
von: Yang, Zheming, et al.
Veröffentlicht: (2025)
Efficient Distributed MLLM Training with Cornstarch
von: Jang, Insu, et al.
Veröffentlicht: (2025)
von: Jang, Insu, et al.
Veröffentlicht: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
Communication Efficient and Provable Federated Unlearning
von: Tao, Youming, et al.
Veröffentlicht: (2024)
von: Tao, Youming, et al.
Veröffentlicht: (2024)
Distributed Bilevel Optimization with Dual Pruning for Resource-limited Clients
von: Li, Mingyi, et al.
Veröffentlicht: (2025)
von: Li, Mingyi, et al.
Veröffentlicht: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
FourierCompress: Layer-Aware Spectral Activation Compression for Efficient and Accurate Collaborative LLM Inference
von: Ma, Jian, et al.
Veröffentlicht: (2025)
von: Ma, Jian, et al.
Veröffentlicht: (2025)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
von: Zhang, Runhua, et al.
Veröffentlicht: (2025)
von: Zhang, Runhua, et al.
Veröffentlicht: (2025)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
von: Shyam, Vasu, et al.
Veröffentlicht: (2026)
von: Shyam, Vasu, et al.
Veröffentlicht: (2026)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
Metronome: Efficient Scheduling for Periodic Traffic Jobs with Network and Priority Awareness
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
von: Wang, Yingping, et al.
Veröffentlicht: (2026)
RServe: Overlapping Encoding and Prefill for Efficient LMM Inference
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2025)
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2025)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
RIPPLE++: An Incremental Framework for Efficient GNN Inference on Evolving Graphs
von: Naman, Pranjal, et al.
Veröffentlicht: (2026)
von: Naman, Pranjal, et al.
Veröffentlicht: (2026)
Parallax: Efficient LLM Inference Service over Decentralized Environment
von: Tong, Chris, et al.
Veröffentlicht: (2025)
von: Tong, Chris, et al.
Veröffentlicht: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
von: Zhang, Songge, et al.
Veröffentlicht: (2026)
von: Zhang, Songge, et al.
Veröffentlicht: (2026)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
von: Wu, Z., et al.
Veröffentlicht: (2025)
von: Wu, Z., et al.
Veröffentlicht: (2025)
BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
von: Sun, Ao, et al.
Veröffentlicht: (2025)
von: Sun, Ao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026) -
PIR-DSN: A Decentralized Storage Network Supporting Private Information Retrieval
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025) -
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
von: Yang, Zheming, et al.
Veröffentlicht: (2026) -
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
von: Ma, Bin, et al.
Veröffentlicht: (2026) -
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)