Training LLMs over Neurally Compressed Text
Fuente:
arXiv
Saved in:
| Main Authors: | Lester, Brian, Lee, Jaehoon, Alemi, Alex, Pennington, Jeffrey, Roberts, Adam, Sohl-Dickstein, Jascha, Constant, Noah |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The boundary of neural network trainability is fractal
by: Sohl-Dickstein, Jascha
Published: (2024)
by: Sohl-Dickstein, Jascha
Published: (2024)
Scaling Exponents Across Parameterizations and Optimizers
by: Everett, Katie, et al.
Published: (2024)
by: Everett, Katie, et al.
Published: (2024)
General-Purpose In-Context Learning by Meta-Learning Transformers
by: Kirsch, Louis, et al.
Published: (2022)
by: Kirsch, Louis, et al.
Published: (2022)
Training Language Models on the Knowledge Graph: Insights on Hallucinations and Their Detectability
by: Hron, Jiri, et al.
Published: (2024)
by: Hron, Jiri, et al.
Published: (2024)
House of Cards: Massive Weights in LLMs
by: Oh, Jaehoon, et al.
Published: (2024)
by: Oh, Jaehoon, et al.
Published: (2024)
Llamazip: Leveraging LLaMA for Lossless Text Compression and Training Dataset Detection
by: Dréano, Sören, et al.
Published: (2025)
by: Dréano, Sören, et al.
Published: (2025)
Determinants of Training Corpus Size for Clinical Text Classification
by: Chaturvedi, Jaya, et al.
Published: (2026)
by: Chaturvedi, Jaya, et al.
Published: (2026)
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
by: Snell, Charlie, et al.
Published: (2024)
by: Snell, Charlie, et al.
Published: (2024)
How Powerful are Decoder-Only Transformer Neural Models?
by: Roberts, Jesse
Published: (2023)
by: Roberts, Jesse
Published: (2023)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
by: Nakkiran, Preetum, et al.
Published: (2025)
by: Nakkiran, Preetum, et al.
Published: (2025)
Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMC
by: Du, Yilun, et al.
Published: (2023)
by: Du, Yilun, et al.
Published: (2023)
Training Neural Networks as Recognizers of Formal Languages
by: Butoi, Alexandra, et al.
Published: (2024)
by: Butoi, Alexandra, et al.
Published: (2024)
LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring
by: Lee, Unggi, et al.
Published: (2026)
by: Lee, Unggi, et al.
Published: (2026)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
by: Jaiswal, Ajay, et al.
Published: (2023)
by: Jaiswal, Ajay, et al.
Published: (2023)
Residual Matrix Transformers: Scaling the Size of the Residual Stream
by: Mak, Brian, et al.
Published: (2025)
by: Mak, Brian, et al.
Published: (2025)
Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
by: Singh, Avi, et al.
Published: (2023)
by: Singh, Avi, et al.
Published: (2023)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
by: Xu, Yuzhuang, et al.
Published: (2024)
by: Xu, Yuzhuang, et al.
Published: (2024)
ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
by: Grishina, Ekaterina, et al.
Published: (2025)
by: Grishina, Ekaterina, et al.
Published: (2025)
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
by: Wallace, Eric, et al.
Published: (2024)
by: Wallace, Eric, et al.
Published: (2024)
Zero2Text: Zero-Training Cross-Domain Inversion Attacks on Textual Embeddings
by: Kim, Doohyun, et al.
Published: (2026)
by: Kim, Doohyun, et al.
Published: (2026)
CompAct: Compressed Activations for Memory-Efficient LLM Training
by: Shamshoum, Yara, et al.
Published: (2024)
by: Shamshoum, Yara, et al.
Published: (2024)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
by: Bai, Yuyang, et al.
Published: (2026)
by: Bai, Yuyang, et al.
Published: (2026)
MarginSel : Max-Margin Demonstration Selection for LLMs
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
Reliable Decision Support with LLMs: A Framework for Evaluating Consistency in Binary Text Classification Applications
by: Megahed, Fadel M., et al.
Published: (2025)
by: Megahed, Fadel M., et al.
Published: (2025)
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024)
by: Marshall, Thomas, et al.
Published: (2024)
TensorLLM: Tensorising Multi-Head Attention for Enhanced Reasoning and Compression in LLMs
by: Gu, Yuxuan, et al.
Published: (2025)
by: Gu, Yuxuan, et al.
Published: (2025)
Why Are Web AI Agents More Vulnerable Than Standalone LLMs? A Security Analysis
by: Chiang, Jeffrey Yang Fan, et al.
Published: (2025)
by: Chiang, Jeffrey Yang Fan, et al.
Published: (2025)
Accelerated AI Inference via Dynamic Execution Methods
by: Barad, Haim, et al.
Published: (2024)
by: Barad, Haim, et al.
Published: (2024)
How to Train Text Summarization Model with Weak Supervisions
by: Wang, Yanbo, et al.
Published: (2024)
by: Wang, Yanbo, et al.
Published: (2024)
Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts
by: Lee, Sang-Woo, et al.
Published: (2025)
by: Lee, Sang-Woo, et al.
Published: (2025)
Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
by: Yakushev, George, et al.
Published: (2025)
by: Yakushev, George, et al.
Published: (2025)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
by: Mao, Yu, et al.
Published: (2025)
by: Mao, Yu, et al.
Published: (2025)
SemanticZip: A Pilot Framework for Lossy Text Compression with LLMs as Semantic Decompressors
by: Trukhina, Natalia, et al.
Published: (2026)
by: Trukhina, Natalia, et al.
Published: (2026)
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
by: Jiang, Huiqiang, et al.
Published: (2023)
by: Jiang, Huiqiang, et al.
Published: (2023)
Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
by: Datta, Joyeeta, et al.
Published: (2025)
by: Datta, Joyeeta, et al.
Published: (2025)
Concept Algebra for (Score-Based) Text-Controlled Generative Models
by: Wang, Zihao, et al.
Published: (2023)
by: Wang, Zihao, et al.
Published: (2023)
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?
by: Hayase, Jonathan, et al.
Published: (2024)
by: Hayase, Jonathan, et al.
Published: (2024)
A Survey : Neural Networks for AMR-to-Text
by: Hao, Hongyu, et al.
Published: (2022)
by: Hao, Hongyu, et al.
Published: (2022)
TensorBLEU: Vectorized GPU-based BLEU Score Implementation for Per-Sentence In-Training Evaluation
by: Filipek, Adam
Published: (2025)
by: Filipek, Adam
Published: (2025)
Conditioning LLMs with Emotion in Neural Machine Translation
by: Brazier, Charles, et al.
Published: (2024)
by: Brazier, Charles, et al.
Published: (2024)
Similar Items
-
The boundary of neural network trainability is fractal
by: Sohl-Dickstein, Jascha
Published: (2024) -
Scaling Exponents Across Parameterizations and Optimizers
by: Everett, Katie, et al.
Published: (2024) -
General-Purpose In-Context Learning by Meta-Learning Transformers
by: Kirsch, Louis, et al.
Published: (2022) -
Training Language Models on the Knowledge Graph: Insights on Hallucinations and Their Detectability
by: Hron, Jiri, et al.
Published: (2024) -
House of Cards: Massive Weights in LLMs
by: Oh, Jaehoon, et al.
Published: (2024)