QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Taesu, Lee, Jongho, Ahn, Daehyun, Kim, Sarang, Choi, Jiwoong, Kim, Minkyu, Kim, Hyungjun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
di: Jung, Yeonjoon, et al.
Pubblicazione: (2025)
di: Jung, Yeonjoon, et al.
Pubblicazione: (2025)
Explicit Feature Interaction-aware Graph Neural Networks
di: Kim, Minkyu, et al.
Pubblicazione: (2022)
di: Kim, Minkyu, et al.
Pubblicazione: (2022)
The Adoption and Efficacy of Large Language Models: Evidence From Consumer Complaints in the Financial Industry
di: Shin, Minkyu, et al.
Pubblicazione: (2023)
di: Shin, Minkyu, et al.
Pubblicazione: (2023)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
di: Kim, Minkyu, et al.
Pubblicazione: (2026)
di: Kim, Minkyu, et al.
Pubblicazione: (2026)
OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
di: Lee, Changhun, et al.
Pubblicazione: (2023)
di: Lee, Changhun, et al.
Pubblicazione: (2023)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
di: Park, Dongmin, et al.
Pubblicazione: (2024)
di: Park, Dongmin, et al.
Pubblicazione: (2024)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
di: Lee, Banseok, et al.
Pubblicazione: (2025)
di: Lee, Banseok, et al.
Pubblicazione: (2025)
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
di: Song, Jiwon, et al.
Pubblicazione: (2024)
di: Song, Jiwon, et al.
Pubblicazione: (2024)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
di: Kim, Dongyoung, et al.
Pubblicazione: (2024)
di: Kim, Dongyoung, et al.
Pubblicazione: (2024)
BoA: Attention-aware Post-training Quantization without Backpropagation
di: Kim, Junhan, et al.
Pubblicazione: (2024)
di: Kim, Junhan, et al.
Pubblicazione: (2024)
HybridRAG: A Practical LLM-based ChatBot Framework based on Pre-Generated Q&A over Raw Unstructured Documents
di: Kim, Sungmoon, et al.
Pubblicazione: (2025)
di: Kim, Sungmoon, et al.
Pubblicazione: (2025)
Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data Augmentation
di: Jin, Kyohoon, et al.
Pubblicazione: (2024)
di: Jin, Kyohoon, et al.
Pubblicazione: (2024)
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring
di: Jung, Hee-Jun, et al.
Pubblicazione: (2022)
di: Jung, Hee-Jun, et al.
Pubblicazione: (2022)
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
di: Kim, Jeonghye, et al.
Pubblicazione: (2025)
di: Kim, Jeonghye, et al.
Pubblicazione: (2025)
Counterfactual-Consistency Prompting for Relative Temporal Understanding in Large Language Models
di: Kim, Jongho, et al.
Pubblicazione: (2025)
di: Kim, Jongho, et al.
Pubblicazione: (2025)
SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling
di: Kim, Dahyun, et al.
Pubblicazione: (2023)
di: Kim, Dahyun, et al.
Pubblicazione: (2023)
MedRep: Medical Concept Representation for General Electronic Health Record Foundation Models
di: Kim, Junmo, et al.
Pubblicazione: (2025)
di: Kim, Junmo, et al.
Pubblicazione: (2025)
Text Change Detection in Multilingual Documents Using Image Comparison
di: Park, Doyoung, et al.
Pubblicazione: (2024)
di: Park, Doyoung, et al.
Pubblicazione: (2024)
Measuring the Depth of LLM Unlearning via Activation Patching
di: Lee, Jaeung, et al.
Pubblicazione: (2026)
di: Lee, Jaeung, et al.
Pubblicazione: (2026)
Becoming Experienced Judges: Selective Test-Time Learning for Evaluators
di: Jwa, Seungyeon, et al.
Pubblicazione: (2025)
di: Jwa, Seungyeon, et al.
Pubblicazione: (2025)
Locality-Aware Redundancy Pruning for LLM Depth Compression
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2026)
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2026)
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
di: Cho, Yoonjun, et al.
Pubblicazione: (2025)
di: Cho, Yoonjun, et al.
Pubblicazione: (2025)
Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
di: Choi, Minsik, et al.
Pubblicazione: (2025)
di: Choi, Minsik, et al.
Pubblicazione: (2025)
Training-free LLM Verification via Recycling Few-shot Examples
di: Lee, Dongseok, et al.
Pubblicazione: (2025)
di: Lee, Dongseok, et al.
Pubblicazione: (2025)
Effective Test-Time Scaling of Discrete Diffusion through Iterative Refinement
di: Lee, Sanghyun, et al.
Pubblicazione: (2025)
di: Lee, Sanghyun, et al.
Pubblicazione: (2025)
Personalizing Large Language Models using Retrieval Augmented Generation and Knowledge Graph
di: Prahlad, Deeksha, et al.
Pubblicazione: (2025)
di: Prahlad, Deeksha, et al.
Pubblicazione: (2025)
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
di: Lim, Junghwan, et al.
Pubblicazione: (2025)
di: Lim, Junghwan, et al.
Pubblicazione: (2025)
AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction
di: Kim, Junsol, et al.
Pubblicazione: (2023)
di: Kim, Junsol, et al.
Pubblicazione: (2023)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
Relevance to Utility: Process-Supervised Rewrite for RAG
di: Kim, Jaeyoung, et al.
Pubblicazione: (2025)
di: Kim, Jaeyoung, et al.
Pubblicazione: (2025)
IPCGRL: Language-Instructed Reinforcement Learning for Procedural Level Generation
di: Baek, In-Chang, et al.
Pubblicazione: (2025)
di: Baek, In-Chang, et al.
Pubblicazione: (2025)
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
di: Kim, Yumin, et al.
Pubblicazione: (2024)
di: Kim, Yumin, et al.
Pubblicazione: (2024)
MELT: Materials-aware Continued Pre-training for Language Model Adaptation to Materials Science
di: Kim, Junho, et al.
Pubblicazione: (2024)
di: Kim, Junho, et al.
Pubblicazione: (2024)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
di: Kim, Minkyu, et al.
Pubblicazione: (2026)
di: Kim, Minkyu, et al.
Pubblicazione: (2026)
Learning to Correct for QA Reasoning with Black-box LLMs
di: Kim, Jaehyung, et al.
Pubblicazione: (2024)
di: Kim, Jaehyung, et al.
Pubblicazione: (2024)
DiTTO-LLM: Framework for Discovering Topic-based Technology Opportunities via Large Language Model
di: Kim, Wonyoung, et al.
Pubblicazione: (2025)
di: Kim, Wonyoung, et al.
Pubblicazione: (2025)
Test-time Alignment of Diffusion Models without Reward Over-optimization
di: Kim, Sunwoo, et al.
Pubblicazione: (2025)
di: Kim, Sunwoo, et al.
Pubblicazione: (2025)
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
di: Koh, Woosung, et al.
Pubblicazione: (2025)
di: Koh, Woosung, et al.
Pubblicazione: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
di: Jung, Yeonjoon, et al.
Pubblicazione: (2025) -
Explicit Feature Interaction-aware Graph Neural Networks
di: Kim, Minkyu, et al.
Pubblicazione: (2022) -
The Adoption and Efficacy of Large Language Models: Evidence From Consumer Complaints in the Financial Industry
di: Shin, Minkyu, et al.
Pubblicazione: (2023) -
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
di: Kim, Minkyu, et al.
Pubblicazione: (2026) -
OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
di: Lee, Changhun, et al.
Pubblicazione: (2023)