SQ-format: A Unified Sparse-Quantized Hardware-friendly Data Format for LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Ruixuan, Zeng, Hao, Huang, Hantao, Shi, Jinyuan, Yu, Minghui, Yen, Ian En-Hsu, Wang, Shuai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
di: Guan, Ziyi, et al.
Pubblicazione: (2024)
di: Guan, Ziyi, et al.
Pubblicazione: (2024)
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
di: Sadhukhan, Ranajoy, et al.
Pubblicazione: (2024)
di: Sadhukhan, Ranajoy, et al.
Pubblicazione: (2024)
Endor: Hardware-Friendly Sparse Format for Offloaded LLM Inference
di: Joo, Donghyeon, et al.
Pubblicazione: (2024)
di: Joo, Donghyeon, et al.
Pubblicazione: (2024)
Enhancing Function-Calling Capabilities in LLMs: Strategies for Prompt Formats, Data Integration, and Multilingual Translation
di: Chen, Yi-Chang, et al.
Pubblicazione: (2024)
di: Chen, Yi-Chang, et al.
Pubblicazione: (2024)
Persona-SQ: A Personalized Suggested Question Generation Framework For Real-world Documents
di: Lin, Zihao, et al.
Pubblicazione: (2024)
di: Lin, Zihao, et al.
Pubblicazione: (2024)
DeSQ: Decomposition-based SPARQL Query Generation
di: Diallo, Papa Abdou Karim Karou, et al.
Pubblicazione: (2026)
di: Diallo, Papa Abdou Karim Karou, et al.
Pubblicazione: (2026)
MobileQuant: Mobile-friendly Quantization for On-device Language Models
di: Tan, Fuwen, et al.
Pubblicazione: (2024)
di: Tan, Fuwen, et al.
Pubblicazione: (2024)
Atomic-SNLI: Fine-Grained Natural Language Inference through Atomic Fact Decomposition
di: Huang, Minghui
Pubblicazione: (2026)
di: Huang, Minghui
Pubblicazione: (2026)
DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs
di: Huang, Minghui
Pubblicazione: (2025)
di: Huang, Minghui
Pubblicazione: (2025)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
di: Huang, Ruixuan, et al.
Pubblicazione: (2025)
di: Huang, Ruixuan, et al.
Pubblicazione: (2025)
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs
di: Long, Do Xuan, et al.
Pubblicazione: (2024)
di: Long, Do Xuan, et al.
Pubblicazione: (2024)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
di: Nawrot, Piotr, et al.
Pubblicazione: (2025)
di: Nawrot, Piotr, et al.
Pubblicazione: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
di: Yuan, Jingyang, et al.
Pubblicazione: (2025)
di: Yuan, Jingyang, et al.
Pubblicazione: (2025)
SynCPKL: Harnessing LLMs to Generate Synthetic Data for Commonsense Persona Knowledge Linking
di: Lin, Kuan-Yen
Pubblicazione: (2024)
di: Lin, Kuan-Yen
Pubblicazione: (2024)
For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
Training Data Efficiency in Multimodal Process Reward Models
di: Li, Jinyuan, et al.
Pubblicazione: (2026)
di: Li, Jinyuan, et al.
Pubblicazione: (2026)
SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs
di: Deng, Boyi, et al.
Pubblicazione: (2025)
di: Deng, Boyi, et al.
Pubblicazione: (2025)
Sparse Brains are Also Adaptive Brains: Cognitive-Load-Aware Dynamic Activation for LLMs
di: Yang, Yiheng, et al.
Pubblicazione: (2025)
di: Yang, Yiheng, et al.
Pubblicazione: (2025)
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
di: Long, Lin, et al.
Pubblicazione: (2024)
di: Long, Lin, et al.
Pubblicazione: (2024)
Fundamental Limitations on Subquadratic Alternatives to Transformers
di: Alman, Josh, et al.
Pubblicazione: (2024)
di: Alman, Josh, et al.
Pubblicazione: (2024)
AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs
di: Lin, Wenxiang, et al.
Pubblicazione: (2026)
di: Lin, Wenxiang, et al.
Pubblicazione: (2026)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
Format-Adapter: Improving Reasoning Capability of LLMs by Adapting Suitable Format
di: Wang, Dingzirui, et al.
Pubblicazione: (2025)
di: Wang, Dingzirui, et al.
Pubblicazione: (2025)
From Oracle to Noisy Context: Mitigating Contextual Exposure Bias in Speech-LLMs
di: Guo, Xiaoyong, et al.
Pubblicazione: (2026)
di: Guo, Xiaoyong, et al.
Pubblicazione: (2026)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
SALS: Sparse Attention in Latent Space for KV cache Compression
di: Mu, Junlin, et al.
Pubblicazione: (2025)
di: Mu, Junlin, et al.
Pubblicazione: (2025)
Evaluating LLMs for Hardware Design and Test
di: Blocklove, Jason, et al.
Pubblicazione: (2024)
di: Blocklove, Jason, et al.
Pubblicazione: (2024)
SqueezeLLM: Dense-and-Sparse Quantization
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
Uncovering Safety Risks of Large Language Models through Concept Activation Vector
di: Xu, Zhihao, et al.
Pubblicazione: (2024)
di: Xu, Zhihao, et al.
Pubblicazione: (2024)
Attention Consistency for LLMs Explanation
di: Lan, Tian, et al.
Pubblicazione: (2025)
di: Lan, Tian, et al.
Pubblicazione: (2025)
MathEDU: Feedback Generation on Problem-Solving Processes for Mathematical Learning Support
di: Hsu, Wei-Ling, et al.
Pubblicazione: (2025)
di: Hsu, Wei-Ling, et al.
Pubblicazione: (2025)
Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
di: Song, Yuerong, et al.
Pubblicazione: (2025)
di: Song, Yuerong, et al.
Pubblicazione: (2025)
Tuning LLMs with Contrastive Alignment Instructions for Machine Translation in Unseen, Low-resource Languages
di: Mao, Zhuoyuan, et al.
Pubblicazione: (2024)
di: Mao, Zhuoyuan, et al.
Pubblicazione: (2024)
LightMamba: Efficient Mamba Acceleration on FPGA with Quantization and Hardware Co-design
di: Wei, Renjie, et al.
Pubblicazione: (2025)
di: Wei, Renjie, et al.
Pubblicazione: (2025)
UniSparse: An Intermediate Language for General Sparse Format Customization
di: Liu, Jie, et al.
Pubblicazione: (2024)
di: Liu, Jie, et al.
Pubblicazione: (2024)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
di: Liu, Hongyi, et al.
Pubblicazione: (2025)
di: Liu, Hongyi, et al.
Pubblicazione: (2025)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
di: Sun, Guohao, et al.
Pubblicazione: (2024)
di: Sun, Guohao, et al.
Pubblicazione: (2024)
Foot-In-The-Door: A Multi-turn Jailbreak for LLMs
di: Weng, Zixuan, et al.
Pubblicazione: (2025)
di: Weng, Zixuan, et al.
Pubblicazione: (2025)
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity
di: Zhang, Di, et al.
Pubblicazione: (2026)
di: Zhang, Di, et al.
Pubblicazione: (2026)
Documenti analoghi
-
APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models
di: Guan, Ziyi, et al.
Pubblicazione: (2024) -
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
di: Sadhukhan, Ranajoy, et al.
Pubblicazione: (2024) -
Endor: Hardware-Friendly Sparse Format for Offloaded LLM Inference
di: Joo, Donghyeon, et al.
Pubblicazione: (2024) -
Enhancing Function-Calling Capabilities in LLMs: Strategies for Prompt Formats, Data Integration, and Multilingual Translation
di: Chen, Yi-Chang, et al.
Pubblicazione: (2024) -
Persona-SQ: A Personalized Suggested Question Generation Framework For Real-world Documents
di: Lin, Zihao, et al.
Pubblicazione: (2024)