Saved in:
| Main Authors: | Xiao, He, Yang, Qingyao, Xie, Dirui, Xu, Wendong, Su, Zunhai, yang, Runming, Zhou, Wenyong, Liu, Haobo, Liu, Zhengwu, Wong, Ngai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.03332 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
by: Xiao, He, et al.
Published: (2025)
by: Xiao, He, et al.
Published: (2025)
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
by: Zhou, Wenyong, et al.
Published: (2025)
by: Zhou, Wenyong, et al.
Published: (2025)
Beyond Outliers: A Data-Free Layer-wise Mixed-Precision Quantization Approach Driven by Numerical and Structural Dual-Sensitivity
by: Zhang, Hengyuan, et al.
Published: (2026)
by: Zhang, Hengyuan, et al.
Published: (2026)
Distribution-Aware Hadamard Quantization for Hardware-Efficient Implicit Neural Representations
by: Zhou, Wenyong, et al.
Published: (2025)
by: Zhou, Wenyong, et al.
Published: (2025)
Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality
by: Wu, Taiqiang, et al.
Published: (2026)
by: Wu, Taiqiang, et al.
Published: (2026)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
Enhancing Robustness of Implicit Neural Representations Against Weight Perturbations
by: Zhou, Wenyong, et al.
Published: (2025)
by: Zhou, Wenyong, et al.
Published: (2025)
MINR: Efficient Implicit Neural Representations for Multi-Image Encoding
by: Zhou, Wenyong, et al.
Published: (2025)
by: Zhou, Wenyong, et al.
Published: (2025)
ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems
by: Zhou, Wenyong, et al.
Published: (2026)
by: Zhou, Wenyong, et al.
Published: (2026)
Timber: Training-free Instruct Model Refining with Base via Effective Rank
by: Wu, Taiqiang, et al.
Published: (2025)
by: Wu, Taiqiang, et al.
Published: (2025)
Decomposing Densification in Gaussian Splatting for Faster 3D Scene Reconstruction
by: Huang, Binxiao, et al.
Published: (2025)
by: Huang, Binxiao, et al.
Published: (2025)
HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CIM Hardware
by: Feng, Yuannuo, et al.
Published: (2025)
by: Feng, Yuannuo, et al.
Published: (2025)
HaLoRA: Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture
by: Wu, Taiqiang, et al.
Published: (2025)
by: Wu, Taiqiang, et al.
Published: (2025)
Quantization Meets Reasoning: Exploring and Mitigating Degradation of Low-Bit LLMs in Mathematical Reasoning
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
Extending Straight-Through Estimation for Robust Neural Networks on Analog CIM Hardware
by: Feng, Yuannuo, et al.
Published: (2025)
by: Feng, Yuannuo, et al.
Published: (2025)
Re-Activating Frozen Primitives for 3D Gaussian Splatting
by: Cheng, Yuxin, et al.
Published: (2025)
by: Cheng, Yuxin, et al.
Published: (2025)
QuadINR: Hardware-Efficient Implicit Neural Representations Through Quadratic Activation
by: Zhou, Wenyong, et al.
Published: (2025)
by: Zhou, Wenyong, et al.
Published: (2025)
CktFormalizer: Autoformalization of Natural Language into Circuit Representations
by: Xiong, Jing, et al.
Published: (2026)
by: Xiong, Jing, et al.
Published: (2026)
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
by: Su, Zunhai, et al.
Published: (2025)
by: Su, Zunhai, et al.
Published: (2025)
Perspective-aware 3D Gaussian Inpainting with Multi-view Consistency
by: Cheng, Yuxin, et al.
Published: (2025)
by: Cheng, Yuxin, et al.
Published: (2025)
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
by: Guan, Ziyi, et al.
Published: (2024)
by: Guan, Ziyi, et al.
Published: (2024)
Revisiting Model Interpolation for Efficient Reasoning
by: Wu, Taiqiang, et al.
Published: (2025)
by: Wu, Taiqiang, et al.
Published: (2025)
DoPE: Denoising Rotary Position Embedding
by: Xiong, Jing, et al.
Published: (2025)
by: Xiong, Jing, et al.
Published: (2025)
Comparing point‐wise and pair‐wise relevance judgment with brain signals
by: Shuqi Zhu, et al.
Published: (2024)
by: Shuqi Zhu, et al.
Published: (2024)
P$^2$-ViT: Power-of-Two Post-Training Quantization and Acceleration for Fully Quantized Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
by: Xu, Wendong, et al.
Published: (2025)
by: Xu, Wendong, et al.
Published: (2025)
Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
by: Wu, Taiqiang, et al.
Published: (2025)
by: Wu, Taiqiang, et al.
Published: (2025)
Nonparametric Teaching for Graph Property Learners
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
Unveiling Super Experts in Mixture-of-Experts Large Language Models
by: Su, Zunhai, et al.
Published: (2025)
by: Su, Zunhai, et al.
Published: (2025)
A Time- and Energy-Efficient CNN with Dense Connections on Memristor-Based Chips
by: Zhou, Wenyong, et al.
Published: (2025)
by: Zhou, Wenyong, et al.
Published: (2025)
Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models
by: Wu, Taiqiang, et al.
Published: (2024)
by: Wu, Taiqiang, et al.
Published: (2024)
Weight Group-wise Post-Training Quantization for Medical Foundation Model
by: Chen, Yineng, et al.
Published: (2026)
by: Chen, Yineng, et al.
Published: (2026)
Layer-wise Quantization for Quantized Optimistic Dual Averaging
by: Nguyen, Anh Duc, et al.
Published: (2025)
by: Nguyen, Anh Duc, et al.
Published: (2025)
Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization
by: Arai, Yamato, et al.
Published: (2025)
by: Arai, Yamato, et al.
Published: (2025)
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
by: Lee, Jung Hyun, et al.
Published: (2023)
by: Lee, Jung Hyun, et al.
Published: (2023)
EPTQ: Enhanced Post-Training Quantization via Hessian-guided Network-wise Optimization
by: Gordon, Ofir, et al.
Published: (2023)
by: Gordon, Ofir, et al.
Published: (2023)
Similar Items
-
PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
by: Xiao, He, et al.
Published: (2025) -
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
by: Zhou, Wenyong, et al.
Published: (2025) -
Beyond Outliers: A Data-Free Layer-wise Mixed-Precision Quantization Approach Driven by Numerical and Structural Dual-Sensitivity
by: Zhang, Hengyuan, et al.
Published: (2026) -
Distribution-Aware Hadamard Quantization for Hardware-Efficient Implicit Neural Representations
by: Zhou, Wenyong, et al.
Published: (2025) -
Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality
by: Wu, Taiqiang, et al.
Published: (2026)