DeCAL Tokenwise Compression
Fuente:
arXiv
Saved in:
| Main Author: | Panwar, Sameer |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Critical Look At Tokenwise Reward-Guided Text Generation
by: Rashid, Ahmad, et al.
Published: (2024)
by: Rashid, Ahmad, et al.
Published: (2024)
In-Context Learning through the Bayesian Prism
by: Panwar, Madhur, et al.
Published: (2023)
by: Panwar, Madhur, et al.
Published: (2023)
Entropy-Aligned Decoding of LMs for Better Writing and Reasoning
by: Ahmed, Kareem, et al.
Published: (2026)
by: Ahmed, Kareem, et al.
Published: (2026)
TIDE: Textual Identity Detection for Evaluating and Augmenting Classification and Language Models
by: Klu, Emmanuel, et al.
Published: (2023)
by: Klu, Emmanuel, et al.
Published: (2023)
ShifaMind: A Multiplicative Concept Bottleneck for Interpretable ICD-10 Coding
by: Syed, Mohammed Sameer, et al.
Published: (2026)
by: Syed, Mohammed Sameer, et al.
Published: (2026)
InversionView: A General-Purpose Method for Reading Information from Neural Activations
by: Huang, Xinting, et al.
Published: (2024)
by: Huang, Xinting, et al.
Published: (2024)
ENMA: Tokenwise Autoregression for Generative Neural PDE Operators
by: Koupaï, Armand Kassaï, et al.
Published: (2025)
by: Koupaï, Armand Kassaï, et al.
Published: (2025)
Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers
by: Ahuja, Kabir, et al.
Published: (2024)
by: Ahuja, Kabir, et al.
Published: (2024)
Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression
by: Trukhina, Natalia, et al.
Published: (2026)
by: Trukhina, Natalia, et al.
Published: (2026)
Proxy Compression for Language Modeling
by: Zheng, Lin, et al.
Published: (2026)
by: Zheng, Lin, et al.
Published: (2026)
Nudging: Inference-time Alignment of LLMs via Guided Decoding
by: Fei, Yu, et al.
Published: (2024)
by: Fei, Yu, et al.
Published: (2024)
Parallel Token Prediction for Language Models
by: Draxler, Felix, et al.
Published: (2025)
by: Draxler, Felix, et al.
Published: (2025)
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
Tiny Transformers Excel at Sentence Compression
by: Belcak, Peter, et al.
Published: (2024)
by: Belcak, Peter, et al.
Published: (2024)
Multi-word Tokenization for Sequence Compression
by: Gee, Leonidas, et al.
Published: (2024)
by: Gee, Leonidas, et al.
Published: (2024)
Compressed Models are NOT Trust-equivalent to Their Large Counterparts
by: Rai, Rohit Raj, et al.
Published: (2025)
by: Rai, Rohit Raj, et al.
Published: (2025)
Merging Feed-Forward Sublayers for Compressed Transformers
by: Verma, Neha, et al.
Published: (2025)
by: Verma, Neha, et al.
Published: (2025)
Compression Scaling Laws:Unifying Sparsity and Quantization
by: Frantar, Elias, et al.
Published: (2025)
by: Frantar, Elias, et al.
Published: (2025)
Are Compressed Language Models Less Subgroup Robust?
by: Gee, Leonidas, et al.
Published: (2024)
by: Gee, Leonidas, et al.
Published: (2024)
Training LLMs over Neurally Compressed Text
by: Lester, Brian, et al.
Published: (2024)
by: Lester, Brian, et al.
Published: (2024)
LoMA: Lossless Compressed Memory Attention
by: Wang, Yumeng, et al.
Published: (2024)
by: Wang, Yumeng, et al.
Published: (2024)
Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills
by: Nottingham, Kolby, et al.
Published: (2024)
by: Nottingham, Kolby, et al.
Published: (2024)
SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression
by: Wen, Haoming, et al.
Published: (2025)
by: Wen, Haoming, et al.
Published: (2025)
Inference-Time Hyper-Scaling with KV Cache Compression
by: Łańcucki, Adrian, et al.
Published: (2025)
by: Łańcucki, Adrian, et al.
Published: (2025)
Better Prompt Compression Without Multi-Layer Perceptrons
by: Honig, Edouardo, et al.
Published: (2025)
by: Honig, Edouardo, et al.
Published: (2025)
Weight-Inherited Distillation for Task-Agnostic BERT Compression
by: Wu, Taiqiang, et al.
Published: (2023)
by: Wu, Taiqiang, et al.
Published: (2023)
Compressed Context Memory For Online Language Model Interaction
by: Kim, Jang-Hyun, et al.
Published: (2023)
by: Kim, Jang-Hyun, et al.
Published: (2023)
Rethinking LLM Memorization through the Lens of Adversarial Compression
by: Schwarzschild, Avi, et al.
Published: (2024)
by: Schwarzschild, Avi, et al.
Published: (2024)
Characterizing Prompt Compression Methods for Long Context Inference
by: Jha, Siddharth, et al.
Published: (2024)
by: Jha, Siddharth, et al.
Published: (2024)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
by: Jaiswal, Ajay, et al.
Published: (2023)
by: Jaiswal, Ajay, et al.
Published: (2023)
SoftQE: Learned Representations of Queries Expanded by LLMs
by: Pimpalkhute, Varad, et al.
Published: (2024)
by: Pimpalkhute, Varad, et al.
Published: (2024)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
by: Song, Woomin, et al.
Published: (2025)
by: Song, Woomin, et al.
Published: (2025)
Radio: Rate-Distortion Optimization for Large Language Model Compression
by: Young, Sean I.
Published: (2025)
by: Young, Sean I.
Published: (2025)
ProCut: LLM Prompt Compression via Attribution Estimation
by: Xu, Zhentao, et al.
Published: (2025)
by: Xu, Zhentao, et al.
Published: (2025)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
by: Yang, Qingyue, et al.
Published: (2025)
by: Yang, Qingyue, et al.
Published: (2025)
SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression
by: Zhang, Biao, et al.
Published: (2025)
by: Zhang, Biao, et al.
Published: (2025)
ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
by: Grishina, Ekaterina, et al.
Published: (2025)
by: Grishina, Ekaterina, et al.
Published: (2025)
Trellis: Learning to Compress Key-Value Memory in Attention Models
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
CompAct: Compressed Activations for Memory-Efficient LLM Training
by: Shamshoum, Yara, et al.
Published: (2024)
by: Shamshoum, Yara, et al.
Published: (2024)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
by: Xu, Yuzhuang, et al.
Published: (2024)
by: Xu, Yuzhuang, et al.
Published: (2024)
Similar Items
-
A Critical Look At Tokenwise Reward-Guided Text Generation
by: Rashid, Ahmad, et al.
Published: (2024) -
In-Context Learning through the Bayesian Prism
by: Panwar, Madhur, et al.
Published: (2023) -
Entropy-Aligned Decoding of LMs for Better Writing and Reasoning
by: Ahmed, Kareem, et al.
Published: (2026) -
TIDE: Textual Identity Detection for Evaluating and Augmenting Classification and Language Models
by: Klu, Emmanuel, et al.
Published: (2023) -
ShifaMind: A Multiplicative Concept Bottleneck for Interpretable ICD-10 Coding
by: Syed, Mohammed Sameer, et al.
Published: (2026)