HoloByte: Continuous Hyperspherical Distillation for Tokenizer-Free Modeling
Fuente:
arXiv
Saved in:
| Main Author: | Khasia, Vladimer |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Attention: True Adaptive World Models via Spherical Kernel Operator
by: Khasia, Vladimer
Published: (2026)
by: Khasia, Vladimer
Published: (2026)
HAS-VQ: Hessian-Adaptive Sparse Vector Quantization for High-Fidelity LLM Compression
by: Khasia, Vladimer
Published: (2026)
by: Khasia, Vladimer
Published: (2026)
Hybrid Dual-Path Linear Transformations for Efficient Transformer Architectures
by: Khasia, Vladimer
Published: (2026)
by: Khasia, Vladimer
Published: (2026)
BASIS: Balanced Activation Sketching with Invariant Scalars for "Ghost Backpropagation"
by: Khasia, Vladimer
Published: (2026)
by: Khasia, Vladimer
Published: (2026)
Spectral-Window Hybrid (SWH)
by: Khasia, Vladimer
Published: (2026)
by: Khasia, Vladimer
Published: (2026)
The Adaptive Vekua Cascade: A Differentiable Spectral-Analytic Solver for Physics-Informed Representation
by: Khasia, Vladimer
Published: (2025)
by: Khasia, Vladimer
Published: (2025)
DeepVekua: Geometric-Spectral Representation Learning for Physics-Informed Fields
by: Khasia, Vladimer
Published: (2025)
by: Khasia, Vladimer
Published: (2025)
Dynamic Subspace Composition: Efficient Adaptation via Contractive Basis Expansion
by: Khasia, Vladimer
Published: (2025)
by: Khasia, Vladimer
Published: (2025)
The Vekua Layer: Exact Physical Priors for Implicit Neural Representations via Generalized Analytic Functions
by: Khasia, Vladimer
Published: (2025)
by: Khasia, Vladimer
Published: (2025)
Primal: A Unified Deterministic Framework for Quasi-Orthogonal Hashing and Manifold Learning
by: Khasia, Vladimer
Published: (2025)
by: Khasia, Vladimer
Published: (2025)
Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
by: Ke, Guolin, et al.
Published: (2025)
by: Ke, Guolin, et al.
Published: (2025)
ByteGen: A Tokenizer-Free Generative Model for Orderbook Events in Byte Space
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer
by: Deng, Chunyuan, et al.
Published: (2026)
by: Deng, Chunyuan, et al.
Published: (2026)
MambaByte: Token-free Selective State Space Model
by: Wang, Junxiong, et al.
Published: (2024)
by: Wang, Junxiong, et al.
Published: (2024)
HYPO: Hyperspherical Out-of-Distribution Generalization
by: Bai, Haoyue, et al.
Published: (2024)
by: Bai, Haoyue, et al.
Published: (2024)
Rethinking Language Model Scaling under Transferable Hypersphere Optimization
by: Ren, Liliang, et al.
Published: (2026)
by: Ren, Liliang, et al.
Published: (2026)
Uncertainty Estimation via Hyperspherical Confidence Mapping
by: Choi, Eunseo, et al.
Published: (2026)
by: Choi, Eunseo, et al.
Published: (2026)
Hyperspherical Normalization for Scalable Deep Reinforcement Learning
by: Lee, Hojoon, et al.
Published: (2025)
by: Lee, Hojoon, et al.
Published: (2025)
Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning
by: Zhou, Zhongyi, et al.
Published: (2025)
by: Zhou, Zhongyi, et al.
Published: (2025)
Hyperspherical Forward-Forward with Prototypical Representations
by: Sarode, Shalini, et al.
Published: (2026)
by: Sarode, Shalini, et al.
Published: (2026)
SpaceByte: Towards Deleting Tokenization from Large Language Modeling
by: Slagle, Kevin
Published: (2024)
by: Slagle, Kevin
Published: (2024)
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
by: Kallini, Julie, et al.
Published: (2024)
by: Kallini, Julie, et al.
Published: (2024)
Deep Orthogonal Hypersphere Compression for Anomaly Detection
by: Zhang, Yunhe, et al.
Published: (2023)
by: Zhang, Yunhe, et al.
Published: (2023)
Angular Regularization for Positive-Unlabeled Learning on the Hypersphere
by: Sevetlidis, Vasileios, et al.
Published: (2025)
by: Sevetlidis, Vasileios, et al.
Published: (2025)
Constrained Machine Learning Through Hyperspherical Representation
by: Signorelli, Gaetano, et al.
Published: (2025)
by: Signorelli, Gaetano, et al.
Published: (2025)
O$n$ Learning Deep O($n$)-Equivariant Hyperspheres
by: Melnyk, Pavlo, et al.
Published: (2023)
by: Melnyk, Pavlo, et al.
Published: (2023)
Hyper-SET: Designing Transformers via Hyperspherical Energy Minimization
by: Hu, Yunzhe, et al.
Published: (2025)
by: Hu, Yunzhe, et al.
Published: (2025)
HyperCore: Coreset Selection under Noise via Hypersphere Models
by: Moser, Brian B., et al.
Published: (2025)
by: Moser, Brian B., et al.
Published: (2025)
The Efficiency Gap in Byte Modeling
by: Lee, Celine, et al.
Published: (2026)
by: Lee, Celine, et al.
Published: (2026)
HOLa: HoloLens Object Labeling
by: Schwimmbeck, Michael, et al.
Published: (2024)
by: Schwimmbeck, Michael, et al.
Published: (2024)
RAVQ-HoloNet: Rate-Adaptive Vector-Quantized Hologram Compression
by: Rafiei, Shima, et al.
Published: (2025)
by: Rafiei, Shima, et al.
Published: (2025)
Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere
by: Ju, Li, et al.
Published: (2025)
by: Ju, Li, et al.
Published: (2025)
nGPT: Normalized Transformer with Representation Learning on the Hypersphere
by: Loshchilov, Ilya, et al.
Published: (2024)
by: Loshchilov, Ilya, et al.
Published: (2024)
Learning Hyperspherical Time-Frequency Representations for Time-Series Out-of-Distribution Detection
by: Lunardi, Willian T., et al.
Published: (2026)
by: Lunardi, Willian T., et al.
Published: (2026)
Enhancing Diversity in Bayesian Deep Learning via Hyperspherical Energy Minimization of CKA
by: Smerkous, David, et al.
Published: (2024)
by: Smerkous, David, et al.
Published: (2024)
Improving the Generation of VAEs with High Dimensional Latent Spaces by the use of Hyperspherical Coordinates
by: Ascarate, Alejandro, et al.
Published: (2025)
by: Ascarate, Alejandro, et al.
Published: (2025)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2026)
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2026)
Token Distillation: Attention-aware Input Embeddings For New Tokens
by: Dobler, Konstantin, et al.
Published: (2025)
by: Dobler, Konstantin, et al.
Published: (2025)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
by: Kadamba, Venu Gopal, et al.
Published: (2026)
by: Kadamba, Venu Gopal, et al.
Published: (2026)
Similar Items
-
Beyond Attention: True Adaptive World Models via Spherical Kernel Operator
by: Khasia, Vladimer
Published: (2026) -
HAS-VQ: Hessian-Adaptive Sparse Vector Quantization for High-Fidelity LLM Compression
by: Khasia, Vladimer
Published: (2026) -
Hybrid Dual-Path Linear Transformations for Efficient Transformer Architectures
by: Khasia, Vladimer
Published: (2026) -
BASIS: Balanced Activation Sketching with Invariant Scalars for "Ghost Backpropagation"
by: Khasia, Vladimer
Published: (2026) -
Spectral-Window Hybrid (SWH)
by: Khasia, Vladimer
Published: (2026)