From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
Fuente:
arXiv
Salvato in:
| Autori principali: | Shani, Chen, Soffer, Liron, Jurafsky, Dan, LeCun, Yann, Shwartz-Ziv, Ravid |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An Information-Theoretic Perspective on Variance-Invariance-Covariance Regularization
di: Shwartz-Ziv, Ravid, et al.
Pubblicazione: (2023)
di: Shwartz-Ziv, Ravid, et al.
Pubblicazione: (2023)
Learning to Compress: Local Rank and Information Compression in Deep Neural Networks
di: Patel, Niket, et al.
Pubblicazione: (2024)
di: Patel, Niket, et al.
Pubblicazione: (2024)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
di: Skean, Oscar, et al.
Pubblicazione: (2024)
di: Skean, Oscar, et al.
Pubblicazione: (2024)
Video Representation Learning with Joint-Embedding Predictive Architectures
di: Drozdov, Katrina, et al.
Pubblicazione: (2024)
di: Drozdov, Katrina, et al.
Pubblicazione: (2024)
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
di: Goldfeder, Judah, et al.
Pubblicazione: (2026)
di: Goldfeder, Judah, et al.
Pubblicazione: (2026)
The Entropy Enigma: Success and Failure of Entropy Minimization
di: Press, Ori, et al.
Pubblicazione: (2024)
di: Press, Ori, et al.
Pubblicazione: (2024)
Layer by Layer: Uncovering Hidden Representations in Language Models
di: Skean, Oscar, et al.
Pubblicazione: (2025)
di: Skean, Oscar, et al.
Pubblicazione: (2025)
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
di: Arefin, Md Rifat, et al.
Pubblicazione: (2024)
di: Arefin, Md Rifat, et al.
Pubblicazione: (2024)
Variance-Covariance Regularization Improves Representation Learning
di: Zhu, Jiachen, et al.
Pubblicazione: (2023)
di: Zhu, Jiachen, et al.
Pubblicazione: (2023)
On Training in Imagination
di: Timor, Nadav, et al.
Pubblicazione: (2026)
di: Timor, Nadav, et al.
Pubblicazione: (2026)
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
di: Queipo-de-Llano, Enrique, et al.
Pubblicazione: (2025)
di: Queipo-de-Llano, Enrique, et al.
Pubblicazione: (2025)
Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models
di: Zhang, Christine, et al.
Pubblicazione: (2026)
di: Zhang, Christine, et al.
Pubblicazione: (2026)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
di: Ioannides, Georgios, et al.
Pubblicazione: (2025)
di: Ioannides, Georgios, et al.
Pubblicazione: (2025)
Beyond Tokens: Concept-Level Training Objectives for LLMs
di: Iyer, Laya, et al.
Pubblicazione: (2026)
di: Iyer, Laya, et al.
Pubblicazione: (2026)
Just How Flexible are Neural Networks in Practice?
di: Shwartz-Ziv, Ravid, et al.
Pubblicazione: (2024)
di: Shwartz-Ziv, Ravid, et al.
Pubblicazione: (2024)
Rate-In: Information-Driven Adaptive Dropout Rates for Improved Inference-Time Uncertainty Estimation
di: Zeevi, Tal, et al.
Pubblicazione: (2024)
di: Zeevi, Tal, et al.
Pubblicazione: (2024)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
di: Sanyal, Sunny, et al.
Pubblicazione: (2024)
di: Sanyal, Sunny, et al.
Pubblicazione: (2024)
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression
di: Elias, Noel, et al.
Pubblicazione: (2024)
di: Elias, Noel, et al.
Pubblicazione: (2024)
Beyond the Loss Curve: Scaling Laws, Active Learning, and the Limits of Learning from Exact Posteriors
di: Khorasani, Arian, et al.
Pubblicazione: (2026)
di: Khorasani, Arian, et al.
Pubblicazione: (2026)
Improving Pre-trained Self-Supervised Embeddings Through Effective Entropy Maximization
di: Chakraborty, Deep, et al.
Pubblicazione: (2024)
di: Chakraborty, Deep, et al.
Pubblicazione: (2024)
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
di: Huang, Hai, et al.
Pubblicazione: (2025)
di: Huang, Hai, et al.
Pubblicazione: (2025)
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
di: LeVi, Amit, et al.
Pubblicazione: (2025)
di: LeVi, Amit, et al.
Pubblicazione: (2025)
Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models
di: Paech, Samuel, et al.
Pubblicazione: (2025)
di: Paech, Samuel, et al.
Pubblicazione: (2025)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
di: Ioannides, Georgios, et al.
Pubblicazione: (2026)
di: Ioannides, Georgios, et al.
Pubblicazione: (2026)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
di: Janiak, Denis, et al.
Pubblicazione: (2025)
di: Janiak, Denis, et al.
Pubblicazione: (2025)
The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
di: Sun, Shangwen, et al.
Pubblicazione: (2026)
di: Sun, Shangwen, et al.
Pubblicazione: (2026)
How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
di: Bianchi, Federico, et al.
Pubblicazione: (2024)
di: Bianchi, Federico, et al.
Pubblicazione: (2024)
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs
di: Chen, Angelica, et al.
Pubblicazione: (2023)
di: Chen, Angelica, et al.
Pubblicazione: (2023)
Multihead Finite-State Compression
di: Lutz, Neil
Pubblicazione: (2025)
di: Lutz, Neil
Pubblicazione: (2025)
Complexity Agnostic Recursive Decomposition of Thoughts
di: Qasim, Kaleem Ullah, et al.
Pubblicazione: (2025)
di: Qasim, Kaleem Ullah, et al.
Pubblicazione: (2025)
Compression of enumerations and gain
di: Barmpalias, George, et al.
Pubblicazione: (2023)
di: Barmpalias, George, et al.
Pubblicazione: (2023)
Frequency-Ordered Tokenization for Better Text Compression
di: Kalcher, Maximilian
Pubblicazione: (2026)
di: Kalcher, Maximilian
Pubblicazione: (2026)
Measuring Grammatical Diversity from Small Corpora: Derivational Entropy Rates, Mean Length of Utterances, and Annotation Invariance
di: Martin, Fermin Moscoso del Prado
Pubblicazione: (2024)
di: Martin, Fermin Moscoso del Prado
Pubblicazione: (2024)
The Roots of Performance Disparity in Multilingual Language Models: Intrinsic Modeling Difficulty or Design Choices?
di: Shani, Chen, et al.
Pubblicazione: (2026)
di: Shani, Chen, et al.
Pubblicazione: (2026)
Nacrith: Neural Lossless Compression via Ensemble Context Modeling and High-Precision CDF Coding
di: Tacconelli, Roberto
Pubblicazione: (2026)
di: Tacconelli, Roberto
Pubblicazione: (2026)
Effective Context in Transformers: An Analysis of Fragmentation and Tokenization
di: Fesharaki, Amirmehdi Jafari, et al.
Pubblicazione: (2026)
di: Fesharaki, Amirmehdi Jafari, et al.
Pubblicazione: (2026)
Compressing integer lists with Contextual Arithmetic Trits
di: Barsamian, Yann, et al.
Pubblicazione: (2022)
di: Barsamian, Yann, et al.
Pubblicazione: (2022)
Rethinking Word Similarity: Semantic Similarity through Classification Confusion
di: Zhou, Kaitlyn, et al.
Pubblicazione: (2025)
di: Zhou, Kaitlyn, et al.
Pubblicazione: (2025)
Transformers without Normalization
di: Zhu, Jiachen, et al.
Pubblicazione: (2025)
di: Zhu, Jiachen, et al.
Pubblicazione: (2025)
Fast and Exact Enumeration of Deep Networks Partitions Regions
di: Balestriero, Randall, et al.
Pubblicazione: (2024)
di: Balestriero, Randall, et al.
Pubblicazione: (2024)
Documenti analoghi
-
An Information-Theoretic Perspective on Variance-Invariance-Covariance Regularization
di: Shwartz-Ziv, Ravid, et al.
Pubblicazione: (2023) -
Learning to Compress: Local Rank and Information Compression in Deep Neural Networks
di: Patel, Niket, et al.
Pubblicazione: (2024) -
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
di: Skean, Oscar, et al.
Pubblicazione: (2024) -
Video Representation Learning with Joint-Embedding Predictive Architectures
di: Drozdov, Katrina, et al.
Pubblicazione: (2024) -
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
di: Goldfeder, Judah, et al.
Pubblicazione: (2026)