Kathleen: Oscillator-Based Byte-Level Text Classification Without Tokenization or Attention
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Fountzoulas, George |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level Tokenizers
par: Jang, Eugene, et autres
Publié: (2024)
par: Jang, Eugene, et autres
Publié: (2024)
Distilling Token-Trained Models into Byte-Level Models
par: Bao, Zishuo, et autres
Publié: (2026)
par: Bao, Zishuo, et autres
Publié: (2026)
Cross-Tokenizer LLM Distillation through a Byte-Level Interface
par: Singh, Avyav Kumar, et autres
Publié: (2026)
par: Singh, Avyav Kumar, et autres
Publié: (2026)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
par: Kadamba, Venu Gopal, et autres
Publié: (2026)
par: Kadamba, Venu Gopal, et autres
Publié: (2026)
An Encoder-Integrated PhoBERT with Graph Attention for Vietnamese Token-Level Classification
par: Nguyen, Ba-Quang
Publié: (2025)
par: Nguyen, Ba-Quang
Publié: (2025)
An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification
par: Rusli, Andre, et autres
Publié: (2024)
par: Rusli, Andre, et autres
Publié: (2024)
Byte BPE Tokenization as an Inverse string Homomorphism
par: Geng, Saibo, et autres
Publié: (2024)
par: Geng, Saibo, et autres
Publié: (2024)
Multi-Level Attention and Contrastive Learning for Enhanced Text Classification with an Optimized Transformer
par: Gao, Jia, et autres
Publié: (2025)
par: Gao, Jia, et autres
Publié: (2025)
Token Masking Improves Transformer-Based Text Classification
par: Xu, Xianglong, et autres
Publié: (2025)
par: Xu, Xianglong, et autres
Publié: (2025)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
par: Hu, Yifan, et autres
Publié: (2025)
par: Hu, Yifan, et autres
Publié: (2025)
Back to Bytes: Revisiting Tokenization Through UTF-8
par: Moryossef, Amit, et autres
Publié: (2025)
par: Moryossef, Amit, et autres
Publié: (2025)
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
par: Phan, Buu, et autres
Publié: (2024)
par: Phan, Buu, et autres
Publié: (2024)
Text Classification Based on Knowledge Graphs and Improved Attention Mechanism
par: Li, Siyu, et autres
Publié: (2024)
par: Li, Siyu, et autres
Publié: (2024)
ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer
par: Deng, Chunyuan, et autres
Publié: (2026)
par: Deng, Chunyuan, et autres
Publié: (2026)
Byte Latent Transformer: Patches Scale Better Than Tokens
par: Pagnoni, Artidoro, et autres
Publié: (2024)
par: Pagnoni, Artidoro, et autres
Publié: (2024)
BanglaByT5: Byte-Level Modelling for Bangla
par: Bhattacharyya, Pramit, et autres
Publié: (2025)
par: Bhattacharyya, Pramit, et autres
Publié: (2025)
Towards Token-Level Text Anomaly Detection
par: Cao, Yang, et autres
Publié: (2026)
par: Cao, Yang, et autres
Publié: (2026)
Token Prediction as Implicit Classification to Identify LLM-Generated Text
par: Chen, Yutian, et autres
Publié: (2023)
par: Chen, Yutian, et autres
Publié: (2023)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
par: Mihaila, George
Publié: (2026)
par: Mihaila, George
Publié: (2026)
MambaByte: Token-free Selective State Space Model
par: Wang, Junxiong, et autres
Publié: (2024)
par: Wang, Junxiong, et autres
Publié: (2024)
Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data
par: Xia, Han, et autres
Publié: (2024)
par: Xia, Han, et autres
Publié: (2024)
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
par: Deng, Difan, et autres
Publié: (2026)
par: Deng, Difan, et autres
Publié: (2026)
Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
par: Gigant, Théo, et autres
Publié: (2026)
par: Gigant, Théo, et autres
Publié: (2026)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
par: Shravan, Rohan
Publié: (2026)
par: Shravan, Rohan
Publié: (2026)
Multi-Level Contextual Token Relation Modeling for Machine-Generated Text Detection
par: Wu, Chenwang, et autres
Publié: (2026)
par: Wu, Chenwang, et autres
Publié: (2026)
Multi-Token Attention
par: Golovneva, Olga, et autres
Publié: (2025)
par: Golovneva, Olga, et autres
Publié: (2025)
Arabic Text Diacritization In The Age Of Transfer Learning: Token Classification Is All You Need
par: Skiredj, Abderrahman, et autres
Publié: (2024)
par: Skiredj, Abderrahman, et autres
Publié: (2024)
Advancing Text Classification with Large Language Models and Neural Attention Mechanisms
par: Lyu, Ning, et autres
Publié: (2025)
par: Lyu, Ning, et autres
Publié: (2025)
Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Sequence-Level Likelihood
par: Lin, Xingyu, et autres
Publié: (2026)
par: Lin, Xingyu, et autres
Publié: (2026)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
par: Zhang, Wanpeng, et autres
Publié: (2024)
par: Zhang, Wanpeng, et autres
Publié: (2024)
Hybrid Tokenization Strategy for DNA Language Model using Byte Pair Encoding and K-MER Methods
par: Sapkota, Ganesh, et autres
Publié: (2025)
par: Sapkota, Ganesh, et autres
Publié: (2025)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
par: Yun, Jungmin, et autres
Publié: (2024)
par: Yun, Jungmin, et autres
Publié: (2024)
Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Markov Likelihood
par: Lin, Xingyu, et autres
Publié: (2025)
par: Lin, Xingyu, et autres
Publié: (2025)
SpaceByte: Towards Deleting Tokenization from Large Language Modeling
par: Slagle, Kevin
Publié: (2024)
par: Slagle, Kevin
Publié: (2024)
UTF-8 Plumbing: Byte-level Tokenizers Unavoidably Enable LLMs to Generate Ill-formed UTF-8
par: Firestone, Preston, et autres
Publié: (2025)
par: Firestone, Preston, et autres
Publié: (2025)
Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal
par: Lian, Haoran, et autres
Publié: (2024)
par: Lian, Haoran, et autres
Publié: (2024)
Vaporetto: Efficient Japanese Tokenization Based on Improved Pointwise Linear Classification
par: Akabe, Koichi, et autres
Publié: (2024)
par: Akabe, Koichi, et autres
Publié: (2024)
Mitigating Bias in Text Classification via Prompt-Based Text Transformation
par: Barker, Charmaine, et autres
Publié: (2023)
par: Barker, Charmaine, et autres
Publié: (2023)
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
par: Foroutan, Negar, et autres
Publié: (2025)
par: Foroutan, Negar, et autres
Publié: (2025)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
par: Kallini, Julie, et autres
Publié: (2024)
par: Kallini, Julie, et autres
Publié: (2024)
Documents similaires
-
Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level Tokenizers
par: Jang, Eugene, et autres
Publié: (2024) -
Distilling Token-Trained Models into Byte-Level Models
par: Bao, Zishuo, et autres
Publié: (2026) -
Cross-Tokenizer LLM Distillation through a Byte-Level Interface
par: Singh, Avyav Kumar, et autres
Publié: (2026) -
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
par: Kadamba, Venu Gopal, et autres
Publié: (2026) -
An Encoder-Integrated PhoBERT with Graph Attention for Vietnamese Token-Level Classification
par: Nguyen, Ba-Quang
Publié: (2025)