Gumbel Distillation for Parallel Text Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Chi, Hu, Xixi, Liu, Bo, Liu, Qiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Distilling Large Language Models for Text-Attributed Graph Learning
di: Pan, Bo, et al.
Pubblicazione: (2024)
di: Pan, Bo, et al.
Pubblicazione: (2024)
Gumbel Counterfactual Generation From Language Models
di: Ravfogel, Shauli, et al.
Pubblicazione: (2024)
di: Ravfogel, Shauli, et al.
Pubblicazione: (2024)
AMO Sampler: Enhancing Text Rendering with Overshooting
di: Hu, Xixi, et al.
Pubblicazione: (2024)
di: Hu, Xixi, et al.
Pubblicazione: (2024)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
di: Dadgarnia, Alireza, et al.
Pubblicazione: (2026)
di: Dadgarnia, Alireza, et al.
Pubblicazione: (2026)
UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations
di: Lu, Qiuyang, et al.
Pubblicazione: (2025)
di: Lu, Qiuyang, et al.
Pubblicazione: (2025)
DiLM: Distilling Dataset into Language Model for Text-level Dataset Distillation
di: Maekawa, Aru, et al.
Pubblicazione: (2024)
di: Maekawa, Aru, et al.
Pubblicazione: (2024)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
di: Pangakis, Nicholas, et al.
Pubblicazione: (2024)
di: Pangakis, Nicholas, et al.
Pubblicazione: (2024)
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
di: Zhou, Yuhang, et al.
Pubblicazione: (2026)
di: Zhou, Yuhang, et al.
Pubblicazione: (2026)
Differentially Private Knowledge Distillation via Synthetic Text Generation
di: Flemings, James, et al.
Pubblicazione: (2024)
di: Flemings, James, et al.
Pubblicazione: (2024)
Distilling Analysis from Generative Models for Investment Decisions
di: Chen, Chung-Chi, et al.
Pubblicazione: (2024)
di: Chen, Chung-Chi, et al.
Pubblicazione: (2024)
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
di: Zhou, Tianyi, et al.
Pubblicazione: (2026)
di: Zhou, Tianyi, et al.
Pubblicazione: (2026)
Bridging Text and Molecule: A Survey on Multimodal Frameworks for Molecule
di: Xiao, Yi, et al.
Pubblicazione: (2024)
di: Xiao, Yi, et al.
Pubblicazione: (2024)
Parallel Scaling Law for Language Models
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
Large Language Models for Conducting Advanced Text Analytics Information Systems Research
di: Ampel, Benjamin M., et al.
Pubblicazione: (2023)
di: Ampel, Benjamin M., et al.
Pubblicazione: (2023)
Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
di: Fang, Luyang, et al.
Pubblicazione: (2025)
di: Fang, Luyang, et al.
Pubblicazione: (2025)
Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
Strong Teacher Not Needed? On Distillation in LLM Pretraining
di: Lu, Taiming, et al.
Pubblicazione: (2026)
di: Lu, Taiming, et al.
Pubblicazione: (2026)
Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean Discrepancy
di: Zhang, Shuhai, et al.
Pubblicazione: (2024)
di: Zhang, Shuhai, et al.
Pubblicazione: (2024)
Multilingual Safety Alignment via Self-Distillation
di: Qin, Ruiyang, et al.
Pubblicazione: (2026)
di: Qin, Ruiyang, et al.
Pubblicazione: (2026)
FlashEvaluator: Expanding Search Space with Parallel Evaluation
di: Feng, Chao, et al.
Pubblicazione: (2026)
di: Feng, Chao, et al.
Pubblicazione: (2026)
You Can Generate It Again: Data-to-Text Generation with Verification and Correction Prompting
di: Ren, Xuan, et al.
Pubblicazione: (2023)
di: Ren, Xuan, et al.
Pubblicazione: (2023)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
di: Zhao, Siyan, et al.
Pubblicazione: (2026)
di: Zhao, Siyan, et al.
Pubblicazione: (2026)
Knowledge Distillation with Training Wheels
di: Liu, Guanlin, et al.
Pubblicazione: (2025)
di: Liu, Guanlin, et al.
Pubblicazione: (2025)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
di: Xiao, Zilin, et al.
Pubblicazione: (2024)
di: Xiao, Zilin, et al.
Pubblicazione: (2024)
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
di: Liu, Chi, et al.
Pubblicazione: (2026)
di: Liu, Chi, et al.
Pubblicazione: (2026)
M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
di: Chen, Jianlv, et al.
Pubblicazione: (2024)
di: Chen, Jianlv, et al.
Pubblicazione: (2024)
Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
di: Li, Chengye, et al.
Pubblicazione: (2025)
di: Li, Chengye, et al.
Pubblicazione: (2025)
DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation
di: Li, Pingzhi, et al.
Pubblicazione: (2025)
di: Li, Pingzhi, et al.
Pubblicazione: (2025)
Text-Guided Molecule Generation with Diffusion Language Model
di: Gong, Haisong, et al.
Pubblicazione: (2024)
di: Gong, Haisong, et al.
Pubblicazione: (2024)
Memory-Efficient LLM Training with Online Subspace Descent
di: Liang, Kaizhao, et al.
Pubblicazione: (2024)
di: Liang, Kaizhao, et al.
Pubblicazione: (2024)
TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning
di: Miok, Kristian, et al.
Pubblicazione: (2025)
di: Miok, Kristian, et al.
Pubblicazione: (2025)
Structured Recurrent Mixers for Massively Parallelized Sequence Generation
di: Badger, Benjamin L.
Pubblicazione: (2026)
di: Badger, Benjamin L.
Pubblicazione: (2026)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
di: Zhang, Songming, et al.
Pubblicazione: (2025)
di: Zhang, Songming, et al.
Pubblicazione: (2025)
Implicit Word Reordering with Knowledge Distillation for Cross-Lingual Dependency Parsing
di: Li, Zhuoran, et al.
Pubblicazione: (2025)
di: Li, Zhuoran, et al.
Pubblicazione: (2025)
MiniDisc: Minimal Distillation Schedule for Language Model Compression
di: Zhang, Chen, et al.
Pubblicazione: (2022)
di: Zhang, Chen, et al.
Pubblicazione: (2022)
A Survey of Large Language Models for Text-Guided Molecular Discovery: from Molecule Generation to Optimization
di: Wang, Ziqing, et al.
Pubblicazione: (2025)
di: Wang, Ziqing, et al.
Pubblicazione: (2025)
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
di: Huang, Mincong, et al.
Pubblicazione: (2024)
di: Huang, Mincong, et al.
Pubblicazione: (2024)
ChatTraffic: Text-to-Traffic Generation via Diffusion Model
di: Zhang, Chengyang, et al.
Pubblicazione: (2023)
di: Zhang, Chengyang, et al.
Pubblicazione: (2023)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
di: Rodionov, Gleb, et al.
Pubblicazione: (2025)
di: Rodionov, Gleb, et al.
Pubblicazione: (2025)
KVSculpt: KV Cache Compression as Distillation
di: Jiang, Bo, et al.
Pubblicazione: (2026)
di: Jiang, Bo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Distilling Large Language Models for Text-Attributed Graph Learning
di: Pan, Bo, et al.
Pubblicazione: (2024) -
Gumbel Counterfactual Generation From Language Models
di: Ravfogel, Shauli, et al.
Pubblicazione: (2024) -
AMO Sampler: Enhancing Text Rendering with Overshooting
di: Hu, Xixi, et al.
Pubblicazione: (2024) -
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
di: Dadgarnia, Alireza, et al.
Pubblicazione: (2026) -
UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations
di: Lu, Qiuyang, et al.
Pubblicazione: (2025)