Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tan, Qitao, Song, Xiaoying, Cheng, Ningxi, Liu, Ninghao, Zhai, Xiaoming, Hong, Lingzi, Wang, Yanzhi, Xiang, Zhen, Yuan, Geng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs
von: Tan, Qitao, et al.
Veröffentlicht: (2026)
von: Tan, Qitao, et al.
Veröffentlicht: (2026)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
A Framework for Human-AI Q-Matrix Refinement: A NeuralCDM Evaluation
von: Zhang, Ying, et al.
Veröffentlicht: (2026)
von: Zhang, Ying, et al.
Veröffentlicht: (2026)
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
von: Lee, Deokjae, et al.
Veröffentlicht: (2025)
von: Lee, Deokjae, et al.
Veröffentlicht: (2025)
Outcome-Constrained Large Language Models for Countering Hate Speech
von: Hong, Lingzi, et al.
Veröffentlicht: (2024)
von: Hong, Lingzi, et al.
Veröffentlicht: (2024)
Assessing the Human Likeness of AI-Generated Counterspeech
von: Song, Xiaoying, et al.
Veröffentlicht: (2024)
von: Song, Xiaoying, et al.
Veröffentlicht: (2024)
Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuning
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
von: Wang, Yun, et al.
Veröffentlicht: (2026)
von: Wang, Yun, et al.
Veröffentlicht: (2026)
From Bits to Chips: An LLM-based Hardware-Aware Quantization Agent for Streamlined Deployment of LLMs
von: Deng, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Deng, Kaiyuan, et al.
Veröffentlicht: (2026)
Piggyback Camera: Easy-to-Deploy Visual Surveillance by Mobile Sensing on Commercial Robot Vacuums
von: Yonetani, Ryo
Veröffentlicht: (2025)
von: Yonetani, Ryo
Veröffentlicht: (2025)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025)
Joint Design of Piggyback and Conjugate Transformation Functions for Repair Bandwidth Reduction in Piggybacking Codes
von: Shi, Hao, et al.
Veröffentlicht: (2026)
von: Shi, Hao, et al.
Veröffentlicht: (2026)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
Echoes of Discord: Forecasting Hater Reactions to Counterspeech
von: Song, Xiaoying, et al.
Veröffentlicht: (2025)
von: Song, Xiaoying, et al.
Veröffentlicht: (2025)
Towards Fast LLM Fine-tuning through Zeroth-Order Optimization with Projected Gradient-Aligned Perturbations
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
A Hybrid Framework for Subject Analysis: Integrating Embedding-Based Regression Models with Large Language Models
von: Liu, Jinyu, et al.
Veröffentlicht: (2025)
von: Liu, Jinyu, et al.
Veröffentlicht: (2025)
A Dynamic Fusion Model for Consistent Crisis Response
von: Song, Xiaoying, et al.
Veröffentlicht: (2025)
von: Song, Xiaoying, et al.
Veröffentlicht: (2025)
A Hybrid Framework for Subject Analysis: Integrating Embedding‐Based Regression Models with Large Language Models
von: Jinyu Liu, et al.
Veröffentlicht: (2025)
von: Jinyu Liu, et al.
Veröffentlicht: (2025)
One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments
von: Yi, Ke, et al.
Veröffentlicht: (2024)
von: Yi, Ke, et al.
Veröffentlicht: (2024)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
Multi-Agent Retrieval-Augmented Framework for Evidence-Based Counterspeech Against Health Misinformation
von: Anik, Anirban Saha, et al.
Veröffentlicht: (2025)
von: Anik, Anirban Saha, et al.
Veröffentlicht: (2025)
Speaking at the Right Level: Literacy-Controlled Counterspeech Generation with RAG-RL
von: Song, Xiaoying, et al.
Veröffentlicht: (2025)
von: Song, Xiaoying, et al.
Veröffentlicht: (2025)
Separability criteria based on realignment
von: Lu, Yu, et al.
Veröffentlicht: (2024)
von: Lu, Yu, et al.
Veröffentlicht: (2024)
Generic realignments in Maxillariinae (Orchidaceae)
von: Mario A Blanco
Veröffentlicht: (2007)
von: Mario A Blanco
Veröffentlicht: (2007)
Rethinking the Potential of Layer Freezing for Efficient DNN Training
von: Yang, Chence, et al.
Veröffentlicht: (2025)
von: Yang, Chence, et al.
Veröffentlicht: (2025)
GLOVE: Global Verifier for LLM Memory-Environment Realignment
von: Yin, Xingkun, et al.
Veröffentlicht: (2026)
von: Yin, Xingkun, et al.
Veröffentlicht: (2026)
Applying Large Language Models and Chain-of-Thought for Automatic Scoring
von: Lee, Gyeong-Geon, et al.
Veröffentlicht: (2023)
von: Lee, Gyeong-Geon, et al.
Veröffentlicht: (2023)
Chapter 9 “Nonsense Rides Piggyback on Sensible Things”
von: Thorpe, Deborah Ellen
Veröffentlicht: (2021)
von: Thorpe, Deborah Ellen
Veröffentlicht: (2021)
Piggybacking as a Creative Strategy to Teach Technology Skills
von: Russell Carpenter, et al.
Veröffentlicht: (2024)
von: Russell Carpenter, et al.
Veröffentlicht: (2024)
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
von: Wan, Ninghao, et al.
Veröffentlicht: (2026)
von: Wan, Ninghao, et al.
Veröffentlicht: (2026)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
von: Zhang, Tuo, et al.
Veröffentlicht: (2025)
von: Zhang, Tuo, et al.
Veröffentlicht: (2025)
Multipartite entanglement based on realignment moments
von: Zhao, Hui, et al.
Veröffentlicht: (2025)
von: Zhao, Hui, et al.
Veröffentlicht: (2025)
Optimizing Age-of-Information in Piggyback Networks with Recurrent Data Generation
von: Lin, Ching-Chi, et al.
Veröffentlicht: (2025)
von: Lin, Ching-Chi, et al.
Veröffentlicht: (2025)
Piggybacking astronomical hazard investigations on scientific Big Data missions
von: Kleijn, Gijs A. Verdoes, et al.
Veröffentlicht: (2024)
von: Kleijn, Gijs A. Verdoes, et al.
Veröffentlicht: (2024)
AyE-Edge: Automated Deployment Space Search Empowering Accuracy yet Efficient Real-Time Object Detection on the Edge
von: Wu, Chao, et al.
Veröffentlicht: (2024)
von: Wu, Chao, et al.
Veröffentlicht: (2024)
BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
von: Wang, Yun, et al.
Veröffentlicht: (2026)
von: Wang, Yun, et al.
Veröffentlicht: (2026)
AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
Mixed Precision Block-Jacobi Preconditioner: Algorithms, Performance Evaluation and Feature Analysis
von: Tian, Ningxi, et al.
Veröffentlicht: (2024)
von: Tian, Ningxi, et al.
Veröffentlicht: (2024)
Tuning-Free Accountable Intervention for LLM Deployment -- A Metacognitive Approach
von: Tan, Zhen, et al.
Veröffentlicht: (2024)
von: Tan, Zhen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs
von: Tan, Qitao, et al.
Veröffentlicht: (2026) -
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
von: Tan, Qitao, et al.
Veröffentlicht: (2025) -
A Framework for Human-AI Q-Matrix Refinement: A NeuralCDM Evaluation
von: Zhang, Ying, et al.
Veröffentlicht: (2026) -
Self-Regularization with Sparse Autoencoders for Controllable LLM-based Classification
von: Wu, Xuansheng, et al.
Veröffentlicht: (2025) -
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
von: Lee, Deokjae, et al.
Veröffentlicht: (2025)