An Enhanced Text Compression Approach Using Transformer-based Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Rahman, Chowdhury Mofizur, Sobhani, Mahbub E, Rodela, Anika Tasnim, Shatabda, Swakkhar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
di: Sobhani, Mahbub E, et al.
Pubblicazione: (2025)
di: Sobhani, Mahbub E, et al.
Pubblicazione: (2025)
Language Modeling Is Compression
di: Delétang, Grégoire, et al.
Pubblicazione: (2023)
di: Delétang, Grégoire, et al.
Pubblicazione: (2023)
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
di: Nagle, Alliot, et al.
Pubblicazione: (2024)
di: Nagle, Alliot, et al.
Pubblicazione: (2024)
Embedding Is (Almost) All You Need: Retrieval-Augmented Inference for Generalizable Genomic Prediction Tasks
di: Datta, Nirjhor, et al.
Pubblicazione: (2025)
di: Datta, Nirjhor, et al.
Pubblicazione: (2025)
Efficient Learned Data Compression via Dual-Stream Feature Decoupling
di: Ma, Huidong, et al.
Pubblicazione: (2026)
di: Ma, Huidong, et al.
Pubblicazione: (2026)
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression
di: Elias, Noel, et al.
Pubblicazione: (2024)
di: Elias, Noel, et al.
Pubblicazione: (2024)
BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce
di: Shamael, Mohammad Nazmush, et al.
Pubblicazione: (2024)
di: Shamael, Mohammad Nazmush, et al.
Pubblicazione: (2024)
Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
di: Badger, Benjamin L., et al.
Pubblicazione: (2025)
di: Badger, Benjamin L., et al.
Pubblicazione: (2025)
Effective Context in Transformers: An Analysis of Fragmentation and Tokenization
di: Fesharaki, Amirmehdi Jafari, et al.
Pubblicazione: (2026)
di: Fesharaki, Amirmehdi Jafari, et al.
Pubblicazione: (2026)
Transformers on Markov Data: Constant Depth Suffices
di: Rajaraman, Nived, et al.
Pubblicazione: (2024)
di: Rajaraman, Nived, et al.
Pubblicazione: (2024)
From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge
di: Chowdhury, Nafis, et al.
Pubblicazione: (2025)
di: Chowdhury, Nafis, et al.
Pubblicazione: (2025)
Understanding Factual Recall in Transformers via Associative Memories
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
di: Nichani, Eshaan, et al.
Pubblicazione: (2024)
Compression Represents Intelligence Linearly
di: Huang, Yuzhen, et al.
Pubblicazione: (2024)
di: Huang, Yuzhen, et al.
Pubblicazione: (2024)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
di: Makkuva, Ashok Vardhan, et al.
Pubblicazione: (2024)
di: Makkuva, Ashok Vardhan, et al.
Pubblicazione: (2024)
Memorization-Compression Cycles Improve Generalization
di: Yu, Fangyuan
Pubblicazione: (2025)
di: Yu, Fangyuan
Pubblicazione: (2025)
AlphaZip: Neural Network-Enhanced Lossless Text Compression
di: Narashiman, Swathi Shree, et al.
Pubblicazione: (2024)
di: Narashiman, Swathi Shree, et al.
Pubblicazione: (2024)
Learning is Forgetting: LLM Training As Lossy Compression
di: Conklin, Henry C., et al.
Pubblicazione: (2026)
di: Conklin, Henry C., et al.
Pubblicazione: (2026)
BN-AuthProf: Benchmarking Machine Learning for Bangla Author Profiling on Social Media Texts
di: Tasnim, Raisa, et al.
Pubblicazione: (2024)
di: Tasnim, Raisa, et al.
Pubblicazione: (2024)
A Mathematical Theory for Learning Semantic Languages by Abstract Learners
di: Liao, Kuo-Yu, et al.
Pubblicazione: (2024)
di: Liao, Kuo-Yu, et al.
Pubblicazione: (2024)
Approaching Rate-Distortion Limits in Neural Compression with Lattice Transform Coding
di: Lei, Eric, et al.
Pubblicazione: (2024)
di: Lei, Eric, et al.
Pubblicazione: (2024)
Cost-aware LLM-based Online Dataset Annotation
di: Elumar, Eray Can, et al.
Pubblicazione: (2025)
di: Elumar, Eray Can, et al.
Pubblicazione: (2025)
BnMMLU: Measuring Massive Multitask Language Understanding in Bengali
di: Joy, Saman Sarker, et al.
Pubblicazione: (2025)
di: Joy, Saman Sarker, et al.
Pubblicazione: (2025)
An Information-theoretic Multi-task Representation Learning Framework for Natural Language Understanding
di: Hu, Dou, et al.
Pubblicazione: (2025)
di: Hu, Dou, et al.
Pubblicazione: (2025)
Quantifying Logical Consistency in Transformers via Query-Key Alignment
di: Tulchinskii, Eduard, et al.
Pubblicazione: (2025)
di: Tulchinskii, Eduard, et al.
Pubblicazione: (2025)
Optimizing Small Transformer-Based Language Models for Multi-Label Sentiment Analysis in Short Texts
di: Neumann, Julius, et al.
Pubblicazione: (2025)
di: Neumann, Julius, et al.
Pubblicazione: (2025)
The Information of Large Language Model Geometry
di: Tan, Zhiquan, et al.
Pubblicazione: (2024)
di: Tan, Zhiquan, et al.
Pubblicazione: (2024)
Theoretical Limits of Language Model Alignment
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2026)
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2026)
A Training-free Method for LLM Text Attribution
di: Radvand, Tara, et al.
Pubblicazione: (2025)
di: Radvand, Tara, et al.
Pubblicazione: (2025)
A Survey on Large Language Models from Concept to Implementation
di: Wang, Chen, et al.
Pubblicazione: (2024)
di: Wang, Chen, et al.
Pubblicazione: (2024)
Integrating Pre-Trained Language Model with Physical Layer Communications
di: Lee, Ju-Hyung, et al.
Pubblicazione: (2024)
di: Lee, Ju-Hyung, et al.
Pubblicazione: (2024)
Geometric Signatures of Compositionality Across a Language Model's Lifetime
di: Lee, Jin Hwa, et al.
Pubblicazione: (2024)
di: Lee, Jin Hwa, et al.
Pubblicazione: (2024)
New Directions in Text Classification Research: Maximizing The Performance of Sentiment Classification from Limited Data
di: Agustian, Surya, et al.
Pubblicazione: (2024)
di: Agustian, Surya, et al.
Pubblicazione: (2024)
Improving Legal Entity Recognition Using a Hybrid Transformer Model and Semantic Filtering Approach
di: Rajamanickam, Duraimurugan
Pubblicazione: (2024)
di: Rajamanickam, Duraimurugan
Pubblicazione: (2024)
Proposal and study of statistical features for string similarity computation and classification
di: Rodrigues, E. O., et al.
Pubblicazione: (2026)
di: Rodrigues, E. O., et al.
Pubblicazione: (2026)
Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
di: Krasnovsky, Anatoly A.
Pubblicazione: (2025)
di: Krasnovsky, Anatoly A.
Pubblicazione: (2025)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
di: Yang, Tong, et al.
Pubblicazione: (2024)
di: Yang, Tong, et al.
Pubblicazione: (2024)
Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models
di: Wei, Lai, et al.
Pubblicazione: (2024)
di: Wei, Lai, et al.
Pubblicazione: (2024)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
di: Wieser, Frederico, et al.
Pubblicazione: (2025)
di: Wieser, Frederico, et al.
Pubblicazione: (2025)
PashtoCorp: A 1.25-Billion-Word Corpus, Evaluation Suite, and Reproducible Pipeline for Low-Resource Language Development
di: Rahman, Hanif
Pubblicazione: (2026)
di: Rahman, Hanif
Pubblicazione: (2026)
Lossless and Near-Lossless Compression for Foundation Models
di: Hershcovitch, Moshik, et al.
Pubblicazione: (2024)
di: Hershcovitch, Moshik, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
di: Sobhani, Mahbub E, et al.
Pubblicazione: (2025) -
Language Modeling Is Compression
di: Delétang, Grégoire, et al.
Pubblicazione: (2023) -
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
di: Nagle, Alliot, et al.
Pubblicazione: (2024) -
Embedding Is (Almost) All You Need: Retrieval-Augmented Inference for Generalizable Genomic Prediction Tasks
di: Datta, Nirjhor, et al.
Pubblicazione: (2025) -
Efficient Learned Data Compression via Dual-Stream Feature Decoupling
di: Ma, Huidong, et al.
Pubblicazione: (2026)