RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Lianming, Wu, Shangyu, Cui, Yufei, Xiong, Ying, Hu, Haibo, Liu, Xue, Kuo, Tei-Wei, Guan, Nan, Xue, Chun Jason |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeeAD: Dynamic Early Exit of Vision-Language Action for Efficient Autonomous Driving
by: HU, Haibo, et al.
Published: (2025)
by: HU, Haibo, et al.
Published: (2025)
Retrieval-Augmented Generation for Natural Language Processing: A Survey
by: Wu, Shangyu, et al.
Published: (2024)
by: Wu, Shangyu, et al.
Published: (2024)
Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving
by: Hu, Haibo, et al.
Published: (2025)
by: Hu, Haibo, et al.
Published: (2025)
AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
by: Huang, Lianming, et al.
Published: (2025)
by: Huang, Lianming, et al.
Published: (2025)
ReFusion: Improving Natural Language Understanding with Computation-Efficient Retrieval Representation Fusion
by: Wu, Shangyu, et al.
Published: (2024)
by: Wu, Shangyu, et al.
Published: (2024)
EvoP: Robust LLM Inference via Evolutionary Pruning
by: Wu, Shangyu, et al.
Published: (2025)
by: Wu, Shangyu, et al.
Published: (2025)
ReFilter: Improving Robustness of Retrieval-Augmented Generation via Gated Filter
by: Chen, Yixin, et al.
Published: (2026)
by: Chen, Yixin, et al.
Published: (2026)
Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever
by: Chen, Yixin, et al.
Published: (2025)
by: Chen, Yixin, et al.
Published: (2025)
On-Demand Multi-Task Sparsity for Efficient Large-Model Deployment on Edge Devices
by: Huang, Lianming, et al.
Published: (2025)
by: Huang, Lianming, et al.
Published: (2025)
GM-Skip: Metric-Guided Transformer Block Skipping for Efficient Vision-Language Models
by: Huang, Lianming, et al.
Published: (2025)
by: Huang, Lianming, et al.
Published: (2025)
FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference
by: Du, Hongchao, et al.
Published: (2025)
by: Du, Hongchao, et al.
Published: (2025)
GeneQuery: A General QA-based Framework for Spatial Gene Expression Predictions from Histology Images
by: Xiong, Ying, et al.
Published: (2024)
by: Xiong, Ying, et al.
Published: (2024)
RALAD: Bridging the Real-to-Sim Domain Gap in Autonomous Driving with Retrieval-Augmented Learning
by: Zuo, Jiacheng, et al.
Published: (2025)
by: Zuo, Jiacheng, et al.
Published: (2025)
Easz: An Agile Transformer-based Image Compression Framework for Resource-constrained IoTs
by: Mao, Yu, et al.
Published: (2025)
by: Mao, Yu, et al.
Published: (2025)
BAHOP: Similarity-based Basin Hopping for A fast hyper-parameter search in WSI classification
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
IHC Matters: Incorporating IHC analysis to H&E Whole Slide Image Analysis for Improved Cancer Grading via Two-stage Multimodal Bilinear Pooling Fusion
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge Devices
by: Peng, Xian, et al.
Published: (2025)
by: Peng, Xian, et al.
Published: (2025)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
by: He, Junhui, et al.
Published: (2024)
by: He, Junhui, et al.
Published: (2024)
VLM-C4L: Continual Core Dataset Learning with Corner Case Optimization via Vision-Language Models for Autonomous Driving
by: Hu, Haibo, et al.
Published: (2025)
by: Hu, Haibo, et al.
Published: (2025)
WISE: A Framework for Gigapixel Whole-Slide-Image Lossless Compression
by: Mao, Yu, et al.
Published: (2025)
by: Mao, Yu, et al.
Published: (2025)
HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
by: Kumar, Avinash, et al.
Published: (2025)
by: Kumar, Avinash, et al.
Published: (2025)
Dynamic Rebatching for Efficient Early-Exit Inference with DREX
by: Liu, Xuting, et al.
Published: (2025)
by: Liu, Xuting, et al.
Published: (2025)
Early-Exit with Class Exclusion for Efficient Inference of Neural Networks
by: Wang, Jingcun, et al.
Published: (2023)
by: Wang, Jingcun, et al.
Published: (2023)
Efficient Distributed Retrieval-Augmented Generation for Enhancing Language Model Performance
by: Liu, Shangyu, et al.
Published: (2025)
by: Liu, Shangyu, et al.
Published: (2025)
A$^2$ATS: Retrieval-Based KV Cache Reduction via Windowed Rotary Position Embedding and Query-Aware Vector Quantization
by: He, Junhui, et al.
Published: (2025)
by: He, Junhui, et al.
Published: (2025)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
by: Zou, Jing, et al.
Published: (2026)
by: Zou, Jing, et al.
Published: (2026)
Advances in Multiple Instance Learning for Whole Slide Image Analysis: Techniques, Challenges, and Future Directions
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
SHAP-CAT: A interpretable multi-modal framework enhancing WSI classification via virtual staining and shapley-value-based multimodal fusion
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Grade Like a Human: Rethinking Automated Assessment with Large Language Models
by: Xie, Wenjing, et al.
Published: (2024)
by: Xie, Wenjing, et al.
Published: (2024)
PTEENet: Post-Trained Early-Exit Neural Networks Augmentation for Inference Cost Optimization
by: Lahiany, Assaf, et al.
Published: (2025)
by: Lahiany, Assaf, et al.
Published: (2025)
When Compression Meets Model Compression: Memory-Efficient Double Compression for Large Language Models
by: Wang, Weilan, et al.
Published: (2025)
by: Wang, Weilan, et al.
Published: (2025)
FlashThink: An Early Exit Method For Efficient Reasoning
by: Jiang, Guochao, et al.
Published: (2025)
by: Jiang, Guochao, et al.
Published: (2025)
CoT-VLM4Tar: Chain-of-Thought Guided Vision-Language Models for Traffic Anomaly Resolution
by: Ren, Tianchi, et al.
Published: (2025)
by: Ren, Tianchi, et al.
Published: (2025)
Mode-as-Sequence: Translating Multimodal Motion Prediction into Unified Sequential Mode Modeling
by: Zhou, Zikang, et al.
Published: (2026)
by: Zhou, Zikang, et al.
Published: (2026)
CEEBERT: Cross-Domain Inference in Early Exit BERT
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
Timely Fusion of Surround Radar/Lidar for Object Detection in Autonomous Driving Systems
by: Xie, Wenjing, et al.
Published: (2023)
by: Xie, Wenjing, et al.
Published: (2023)
On the Compressibility of Quantized Large Language Models
by: Mao, Yu, et al.
Published: (2024)
by: Mao, Yu, et al.
Published: (2024)
BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch Prediction
by: Zhou, Zikang, et al.
Published: (2024)
by: Zhou, Zikang, et al.
Published: (2024)
BEExformer: A Fast Inferencing Binarized Transformer with Early Exits
by: Ansar, Wazib, et al.
Published: (2024)
by: Ansar, Wazib, et al.
Published: (2024)
Early-Exit meets Model-Distributed Inference at Edge Networks
by: Colocrese, Marco, et al.
Published: (2024)
by: Colocrese, Marco, et al.
Published: (2024)
Similar Items
-
DeeAD: Dynamic Early Exit of Vision-Language Action for Efficient Autonomous Driving
by: HU, Haibo, et al.
Published: (2025) -
Retrieval-Augmented Generation for Natural Language Processing: A Survey
by: Wu, Shangyu, et al.
Published: (2024) -
Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving
by: Hu, Haibo, et al.
Published: (2025) -
AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
by: Huang, Lianming, et al.
Published: (2025) -
ReFusion: Improving Natural Language Understanding with Computation-Efficient Retrieval Representation Fusion
by: Wu, Shangyu, et al.
Published: (2024)