Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal
Fuente:
arXiv
Saved in:
| Main Authors: | Lian, Haoran, Xiong, Yizhe, Niu, Jianwei, Mo, Shasha, Su, Zhenpeng, Lin, Zijia, Chen, Hui, Liu, Peng, Han, Jungong, Ding, Guiguang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LBPE: Long-token-first Tokenization to Improve Large Language Models
by: Lian, Haoran, et al.
Published: (2024)
by: Lian, Haoran, et al.
Published: (2024)
Temporal Scaling Law for Large Language Models
by: Xiong, Yizhe, et al.
Published: (2024)
by: Xiong, Yizhe, et al.
Published: (2024)
UniAttn: Reducing Inference Costs via Softmax Unification for Post-Training LLMs
by: Xiong, Yizhe, et al.
Published: (2025)
by: Xiong, Yizhe, et al.
Published: (2025)
MaskMoE: Boosting Token-Level Learning via Routing Mask in Mixture-of-Experts
by: Su, Zhenpeng, et al.
Published: (2024)
by: Su, Zhenpeng, et al.
Published: (2024)
Neutralizing Token Aggregation via Information Augmentation for Efficient Test-Time Adaptation
by: Xiong, Yizhe, et al.
Published: (2025)
by: Xiong, Yizhe, et al.
Published: (2025)
GraphBPE: Molecular Graphs Meet Byte-Pair Encoding
by: Shen, Yuchen, et al.
Published: (2024)
by: Shen, Yuchen, et al.
Published: (2024)
[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
Fast Quiet-STaR: Thinking Without Thought Tokens
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Breaking the Stage Barrier: A Novel Single-Stage Approach to Long Context Extension for Large Language Models
by: Lian, Haoran, et al.
Published: (2024)
by: Lian, Haoran, et al.
Published: (2024)
RepViT-SAM: Towards Real-Time Segmenting Anything
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
LSNet: See Large, Focus Small
by: Wang, Ao, et al.
Published: (2025)
by: Wang, Ao, et al.
Published: (2025)
RepViT: Revisiting Mobile CNN From ViT Perspective
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
CartesianMoE: Boosting Knowledge Sharing among Experts via Cartesian Product Routing in Mixture-of-Experts
by: Su, Zhenpeng, et al.
Published: (2024)
by: Su, Zhenpeng, et al.
Published: (2024)
Finedeep: Mitigating Sparse Activation in Dense LLMs via Multi-Layer Fine-Grained Experts
by: Pan, Leiyu, et al.
Published: (2025)
by: Pan, Leiyu, et al.
Published: (2025)
Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations
by: Liang, Yiwen, et al.
Published: (2025)
by: Liang, Yiwen, et al.
Published: (2025)
Byte BPE Tokenization as an Inverse string Homomorphism
by: Geng, Saibo, et al.
Published: (2024)
by: Geng, Saibo, et al.
Published: (2024)
PYRA: Parallel Yielding Re-Activation for Training-Inference Efficient Task Adaptation
by: Xiong, Yizhe, et al.
Published: (2024)
by: Xiong, Yizhe, et al.
Published: (2024)
DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMs
by: Lv, Minxuan, et al.
Published: (2025)
by: Lv, Minxuan, et al.
Published: (2025)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
by: Kadamba, Venu Gopal, et al.
Published: (2026)
by: Kadamba, Venu Gopal, et al.
Published: (2026)
CAIT: Triple-Win Compression towards High Accuracy, Fast Inference, and Favorable Transferability For ViTs
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
YOLOE: Real-Time Seeing Anything
by: Wang, Ao, et al.
Published: (2025)
by: Wang, Ao, et al.
Published: (2025)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
by: Hu, Yifan, et al.
Published: (2025)
by: Hu, Yifan, et al.
Published: (2025)
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models
by: Sun, Fengyuan, et al.
Published: (2025)
by: Sun, Fengyuan, et al.
Published: (2025)
Learn from the Learnt: Source-Free Active Domain Adaptation via Contrastive Sampling and Visual Persistence
by: Lyu, Mengyao, et al.
Published: (2024)
by: Lyu, Mengyao, et al.
Published: (2024)
YOLOv10: Real-Time End-to-End Object Detection
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
Protein Structure Tokenization via Geometric Byte Pair Encoding
by: Sun, Michael, et al.
Published: (2025)
by: Sun, Michael, et al.
Published: (2025)
MiLe Loss: a New Entropy-Weighed Loss for Mitigating the Bias of Learning Difficulties in Large Language Models
by: Su, Zhenpeng, et al.
Published: (2023)
by: Su, Zhenpeng, et al.
Published: (2023)
LiteToken: Removing Intermediate Merge Residues From BPE Tokenizers
by: Sun, Yike, et al.
Published: (2026)
by: Sun, Yike, et al.
Published: (2026)
BlockBPE: Parallel BPE Tokenization
by: You, Amos
Published: (2025)
by: You, Amos
Published: (2025)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
by: Zhang, Wanpeng, et al.
Published: (2024)
by: Zhang, Wanpeng, et al.
Published: (2024)
Context Enhancement with Reconstruction as Sequence for Unified Unsupervised Anomaly Detection
by: Yang, Hui-Yue, et al.
Published: (2024)
by: Yang, Hui-Yue, et al.
Published: (2024)
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Theoretical Analysis of Byte-Pair Encoding
by: Kozma, László, et al.
Published: (2024)
by: Kozma, László, et al.
Published: (2024)
Tracking and Segmenting Anything in Any Modality
by: Zhang, Tianlu, et al.
Published: (2025)
by: Zhang, Tianlu, et al.
Published: (2025)
Batching BPE Tokenization Merges
by: Morgan, Alexander P.
Published: (2024)
by: Morgan, Alexander P.
Published: (2024)
PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
A Formal Perspective on Byte-Pair Encoding
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
Constructing a BPE Tokenization DFA
by: Berglund, Martin, et al.
Published: (2024)
by: Berglund, Martin, et al.
Published: (2024)
Hybrid Tokenization Strategy for DNA Language Model using Byte Pair Encoding and K-MER Methods
by: Sapkota, Ganesh, et al.
Published: (2025)
by: Sapkota, Ganesh, et al.
Published: (2025)
Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning
by: Yang, Hui-Yue, et al.
Published: (2024)
by: Yang, Hui-Yue, et al.
Published: (2024)
Similar Items
-
LBPE: Long-token-first Tokenization to Improve Large Language Models
by: Lian, Haoran, et al.
Published: (2024) -
Temporal Scaling Law for Large Language Models
by: Xiong, Yizhe, et al.
Published: (2024) -
UniAttn: Reducing Inference Costs via Softmax Unification for Post-Training LLMs
by: Xiong, Yizhe, et al.
Published: (2025) -
MaskMoE: Boosting Token-Level Learning via Routing Mask in Mixture-of-Experts
by: Su, Zhenpeng, et al.
Published: (2024) -
Neutralizing Token Aggregation via Information Augmentation for Efficient Test-Time Adaptation
by: Xiong, Yizhe, et al.
Published: (2025)