An Enhanced Text Compression Approach Using Transformer-based Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rahman, Chowdhury Mofizur, Sobhani, Mahbub E, Rodela, Anika Tasnim, Shatabda, Swakkhar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
von: Sobhani, Mahbub E, et al.
Veröffentlicht: (2025)
von: Sobhani, Mahbub E, et al.
Veröffentlicht: (2025)
Language Modeling Is Compression
von: Delétang, Grégoire, et al.
Veröffentlicht: (2023)
von: Delétang, Grégoire, et al.
Veröffentlicht: (2023)
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
von: Nagle, Alliot, et al.
Veröffentlicht: (2024)
von: Nagle, Alliot, et al.
Veröffentlicht: (2024)
Embedding Is (Almost) All You Need: Retrieval-Augmented Inference for Generalizable Genomic Prediction Tasks
von: Datta, Nirjhor, et al.
Veröffentlicht: (2025)
von: Datta, Nirjhor, et al.
Veröffentlicht: (2025)
Efficient Learned Data Compression via Dual-Stream Feature Decoupling
von: Ma, Huidong, et al.
Veröffentlicht: (2026)
von: Ma, Huidong, et al.
Veröffentlicht: (2026)
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression
von: Elias, Noel, et al.
Veröffentlicht: (2024)
von: Elias, Noel, et al.
Veröffentlicht: (2024)
BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce
von: Shamael, Mohammad Nazmush, et al.
Veröffentlicht: (2024)
von: Shamael, Mohammad Nazmush, et al.
Veröffentlicht: (2024)
Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
von: Badger, Benjamin L., et al.
Veröffentlicht: (2025)
von: Badger, Benjamin L., et al.
Veröffentlicht: (2025)
Effective Context in Transformers: An Analysis of Fragmentation and Tokenization
von: Fesharaki, Amirmehdi Jafari, et al.
Veröffentlicht: (2026)
von: Fesharaki, Amirmehdi Jafari, et al.
Veröffentlicht: (2026)
Transformers on Markov Data: Constant Depth Suffices
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge
von: Chowdhury, Nafis, et al.
Veröffentlicht: (2025)
von: Chowdhury, Nafis, et al.
Veröffentlicht: (2025)
Understanding Factual Recall in Transformers via Associative Memories
von: Nichani, Eshaan, et al.
Veröffentlicht: (2024)
von: Nichani, Eshaan, et al.
Veröffentlicht: (2024)
Compression Represents Intelligence Linearly
von: Huang, Yuzhen, et al.
Veröffentlicht: (2024)
von: Huang, Yuzhen, et al.
Veröffentlicht: (2024)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
Memorization-Compression Cycles Improve Generalization
von: Yu, Fangyuan
Veröffentlicht: (2025)
von: Yu, Fangyuan
Veröffentlicht: (2025)
AlphaZip: Neural Network-Enhanced Lossless Text Compression
von: Narashiman, Swathi Shree, et al.
Veröffentlicht: (2024)
von: Narashiman, Swathi Shree, et al.
Veröffentlicht: (2024)
Learning is Forgetting: LLM Training As Lossy Compression
von: Conklin, Henry C., et al.
Veröffentlicht: (2026)
von: Conklin, Henry C., et al.
Veröffentlicht: (2026)
BN-AuthProf: Benchmarking Machine Learning for Bangla Author Profiling on Social Media Texts
von: Tasnim, Raisa, et al.
Veröffentlicht: (2024)
von: Tasnim, Raisa, et al.
Veröffentlicht: (2024)
A Mathematical Theory for Learning Semantic Languages by Abstract Learners
von: Liao, Kuo-Yu, et al.
Veröffentlicht: (2024)
von: Liao, Kuo-Yu, et al.
Veröffentlicht: (2024)
Approaching Rate-Distortion Limits in Neural Compression with Lattice Transform Coding
von: Lei, Eric, et al.
Veröffentlicht: (2024)
von: Lei, Eric, et al.
Veröffentlicht: (2024)
Cost-aware LLM-based Online Dataset Annotation
von: Elumar, Eray Can, et al.
Veröffentlicht: (2025)
von: Elumar, Eray Can, et al.
Veröffentlicht: (2025)
BnMMLU: Measuring Massive Multitask Language Understanding in Bengali
von: Joy, Saman Sarker, et al.
Veröffentlicht: (2025)
von: Joy, Saman Sarker, et al.
Veröffentlicht: (2025)
An Information-theoretic Multi-task Representation Learning Framework for Natural Language Understanding
von: Hu, Dou, et al.
Veröffentlicht: (2025)
von: Hu, Dou, et al.
Veröffentlicht: (2025)
Quantifying Logical Consistency in Transformers via Query-Key Alignment
von: Tulchinskii, Eduard, et al.
Veröffentlicht: (2025)
von: Tulchinskii, Eduard, et al.
Veröffentlicht: (2025)
Optimizing Small Transformer-Based Language Models for Multi-Label Sentiment Analysis in Short Texts
von: Neumann, Julius, et al.
Veröffentlicht: (2025)
von: Neumann, Julius, et al.
Veröffentlicht: (2025)
The Information of Large Language Model Geometry
von: Tan, Zhiquan, et al.
Veröffentlicht: (2024)
von: Tan, Zhiquan, et al.
Veröffentlicht: (2024)
Theoretical Limits of Language Model Alignment
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2026)
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2026)
A Training-free Method for LLM Text Attribution
von: Radvand, Tara, et al.
Veröffentlicht: (2025)
von: Radvand, Tara, et al.
Veröffentlicht: (2025)
A Survey on Large Language Models from Concept to Implementation
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
Integrating Pre-Trained Language Model with Physical Layer Communications
von: Lee, Ju-Hyung, et al.
Veröffentlicht: (2024)
von: Lee, Ju-Hyung, et al.
Veröffentlicht: (2024)
Geometric Signatures of Compositionality Across a Language Model's Lifetime
von: Lee, Jin Hwa, et al.
Veröffentlicht: (2024)
von: Lee, Jin Hwa, et al.
Veröffentlicht: (2024)
New Directions in Text Classification Research: Maximizing The Performance of Sentiment Classification from Limited Data
von: Agustian, Surya, et al.
Veröffentlicht: (2024)
von: Agustian, Surya, et al.
Veröffentlicht: (2024)
Improving Legal Entity Recognition Using a Hybrid Transformer Model and Semantic Filtering Approach
von: Rajamanickam, Duraimurugan
Veröffentlicht: (2024)
von: Rajamanickam, Duraimurugan
Veröffentlicht: (2024)
Proposal and study of statistical features for string similarity computation and classification
von: Rodrigues, E. O., et al.
Veröffentlicht: (2026)
von: Rodrigues, E. O., et al.
Veröffentlicht: (2026)
Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
von: Krasnovsky, Anatoly A.
Veröffentlicht: (2025)
von: Krasnovsky, Anatoly A.
Veröffentlicht: (2025)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
von: Yang, Tong, et al.
Veröffentlicht: (2024)
von: Yang, Tong, et al.
Veröffentlicht: (2024)
Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models
von: Wei, Lai, et al.
Veröffentlicht: (2024)
von: Wei, Lai, et al.
Veröffentlicht: (2024)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
von: Wieser, Frederico, et al.
Veröffentlicht: (2025)
von: Wieser, Frederico, et al.
Veröffentlicht: (2025)
PashtoCorp: A 1.25-Billion-Word Corpus, Evaluation Suite, and Reproducible Pipeline for Low-Resource Language Development
von: Rahman, Hanif
Veröffentlicht: (2026)
von: Rahman, Hanif
Veröffentlicht: (2026)
Lossless and Near-Lossless Compression for Foundation Models
von: Hershcovitch, Moshik, et al.
Veröffentlicht: (2024)
von: Hershcovitch, Moshik, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
von: Sobhani, Mahbub E, et al.
Veröffentlicht: (2025) -
Language Modeling Is Compression
von: Delétang, Grégoire, et al.
Veröffentlicht: (2023) -
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
von: Nagle, Alliot, et al.
Veröffentlicht: (2024) -
Embedding Is (Almost) All You Need: Retrieval-Augmented Inference for Generalizable Genomic Prediction Tasks
von: Datta, Nirjhor, et al.
Veröffentlicht: (2025) -
Efficient Learned Data Compression via Dual-Stream Feature Decoupling
von: Ma, Huidong, et al.
Veröffentlicht: (2026)