Value-Driven Mixed-Precision Quantization for Patch-Based Inference on Microcontrollers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tao, Wei, He, Shenglin, Lu, Kai, Qu, Xiaoyang, Li, Guokuan, Wan, Jiguang, Wang, Jianzong, Xiao, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
von: Tao, Wei, et al.
Veröffentlicht: (2026)
von: Tao, Wei, et al.
Veröffentlicht: (2026)
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
von: Tao, Wei, et al.
Veröffentlicht: (2025)
von: Tao, Wei, et al.
Veröffentlicht: (2025)
PRENet: A Plane-Fit Redundancy Encoding Point Cloud Sequence Network for Real-Time 3D Action Recognition
von: He, Shenglin, et al.
Veröffentlicht: (2024)
von: He, Shenglin, et al.
Veröffentlicht: (2024)
BAGNet: A Boundary-Aware Graph Attention Network for 3D Point Cloud Semantic Segmentation
von: Tao, Wei, et al.
Veröffentlicht: (2025)
von: Tao, Wei, et al.
Veröffentlicht: (2025)
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
RUNA: Object-level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
von: Lu, Haocheng, et al.
Veröffentlicht: (2026)
von: Lu, Haocheng, et al.
Veröffentlicht: (2026)
RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
von: Wang, Anmin, et al.
Veröffentlicht: (2026)
von: Wang, Anmin, et al.
Veröffentlicht: (2026)
Cocktail: Chunk-Adaptive Mixed-Precision Quantization for Long-Context LLM Inference
von: Tao, Wei, et al.
Veröffentlicht: (2025)
von: Tao, Wei, et al.
Veröffentlicht: (2025)
MADLLM: Multivariate Anomaly Detection via Pre-trained LLMs
von: Tao, Wei, et al.
Veröffentlicht: (2025)
von: Tao, Wei, et al.
Veröffentlicht: (2025)
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
von: Liu, Chuhang, et al.
Veröffentlicht: (2026)
von: Liu, Chuhang, et al.
Veröffentlicht: (2026)
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
Federated Domain Generalization with Domain-specific Soft Prompts Generation
von: Wu, Jianhan, et al.
Veröffentlicht: (2025)
von: Wu, Jianhan, et al.
Veröffentlicht: (2025)
Enhancing Multi-Agent Systems via Reinforcement Learning with LLM-based Planner and Graph-based Policy
von: Jia, Ziqi, et al.
Veröffentlicht: (2025)
von: Jia, Ziqi, et al.
Veröffentlicht: (2025)
MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
von: Lu, Renjie, et al.
Veröffentlicht: (2026)
Mix-QViT: Mixed-Precision Vision Transformer Quantization Driven by Layer Importance and Quantization Sensitivity
von: Ranjan, Navin, et al.
Veröffentlicht: (2025)
von: Ranjan, Navin, et al.
Veröffentlicht: (2025)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
von: Shi, Jiaqi, et al.
Veröffentlicht: (2026)
ESARM: 3D Emotional Speech-to-Animation via Reward Model from Automatically-Ranked Demonstrations
von: Zhang, Xulong, et al.
Veröffentlicht: (2024)
von: Zhang, Xulong, et al.
Veröffentlicht: (2024)
Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual Learning
von: Jia, Ziqi, et al.
Veröffentlicht: (2025)
von: Jia, Ziqi, et al.
Veröffentlicht: (2025)
6Bit-Diffusion: Inference-Time Mixed-Precision Quantization for Video Diffusion Models
von: Su, Rundong, et al.
Veröffentlicht: (2026)
von: Su, Rundong, et al.
Veröffentlicht: (2026)
Mix-QSAM: Mixed-Precision Quantization of the Segment Anything Model
von: Ranjan, Navin, et al.
Veröffentlicht: (2025)
von: Ranjan, Navin, et al.
Veröffentlicht: (2025)
MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models
von: Feng, Weilun, et al.
Veröffentlicht: (2024)
von: Feng, Weilun, et al.
Veröffentlicht: (2024)
Efficient and Effective Methods for Mixed Precision Neural Network Quantization for Faster, Energy-efficient Inference
von: Bablani, Deepika, et al.
Veröffentlicht: (2023)
von: Bablani, Deepika, et al.
Veröffentlicht: (2023)
Adaptive Distribution-aware Quantization for Mixed-Precision Neural Networks
von: Jia, Shaohang, et al.
Veröffentlicht: (2025)
von: Jia, Shaohang, et al.
Veröffentlicht: (2025)
Learning from Loss Landscape: Generalizable Mixed-Precision Quantization via Adaptive Sharpness-Aware Gradient Aligning
von: Ma, Lianbo, et al.
Veröffentlicht: (2025)
von: Ma, Lianbo, et al.
Veröffentlicht: (2025)
MPQ-Diff: Mixed Precision Quantization for Diffusion Models
von: Maruzzelli, Rocco Manz, et al.
Veröffentlicht: (2024)
von: Maruzzelli, Rocco Manz, et al.
Veröffentlicht: (2024)
MPTQ-ViT: Mixed-Precision Post-Training Quantization for Vision Transformer
von: Tai, Yu-Shan, et al.
Veröffentlicht: (2024)
von: Tai, Yu-Shan, et al.
Veröffentlicht: (2024)
MetaMix: Meta-state Precision Searcher for Mixed-precision Activation Quantization
von: Kim, Han-Byul, et al.
Veröffentlicht: (2023)
von: Kim, Han-Byul, et al.
Veröffentlicht: (2023)
GAIA: Delving into Gradient-based Attribution Abnormality for Out-of-distribution Detection
von: Chen, Jinggang, et al.
Veröffentlicht: (2023)
von: Chen, Jinggang, et al.
Veröffentlicht: (2023)
LoRA Patching: Exposing the Fragility of Proactive Defenses against Deepfakes
von: Qu, Zuomin, et al.
Veröffentlicht: (2025)
von: Qu, Zuomin, et al.
Veröffentlicht: (2025)
TreeQ: Pushing the Quantization Boundary of Diffusion Transformer via Tree-Structured Mixed-Precision Search
von: Yang, Kaicheng, et al.
Veröffentlicht: (2025)
von: Yang, Kaicheng, et al.
Veröffentlicht: (2025)
On-Device Vision Training, Deployment, and Inference on a Thumb-Sized Microcontroller
von: Ellis, Jeremy
Veröffentlicht: (2026)
von: Ellis, Jeremy
Veröffentlicht: (2026)
Animal Re-Identification on Microcontrollers
von: Chen, Yubo, et al.
Veröffentlicht: (2025)
von: Chen, Yubo, et al.
Veröffentlicht: (2025)
MixA-Q: Revisiting Activation Sparsity for Vision Transformers from a Mixed-Precision Quantization Perspective
von: Wang, Weitian, et al.
Veröffentlicht: (2025)
von: Wang, Weitian, et al.
Veröffentlicht: (2025)
Patch-aware Vector Quantized Codebook Learning for Unsupervised Visual Defect Detection
von: Cheng, Qisen, et al.
Veröffentlicht: (2025)
von: Cheng, Qisen, et al.
Veröffentlicht: (2025)
LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision Transformers
von: Kim, Minjun, et al.
Veröffentlicht: (2025)
von: Kim, Minjun, et al.
Veröffentlicht: (2025)
Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language
von: Chen, Yicheng, et al.
Veröffentlicht: (2024)
von: Chen, Yicheng, et al.
Veröffentlicht: (2024)
Precision Neural Network Quantization via Learnable Adaptive Modules
von: Zhou, Wenqiang, et al.
Veröffentlicht: (2025)
von: Zhou, Wenqiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
von: Tao, Wei, et al.
Veröffentlicht: (2026) -
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
von: Tao, Wei, et al.
Veröffentlicht: (2025) -
PRENet: A Plane-Fit Redundancy Encoding Point Cloud Sequence Network for Real-Time 3D Action Recognition
von: He, Shenglin, et al.
Veröffentlicht: (2024) -
BAGNet: A Boundary-Aware Graph Attention Network for 3D Point Cloud Semantic Segmentation
von: Tao, Wei, et al.
Veröffentlicht: (2025) -
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
von: Zhang, Bin, et al.
Veröffentlicht: (2025)