Tokenisation via Convex Relaxations
Fuente:
arXiv
Salvato in:
| Autori principali: | Tempus, Jan, Whittington, Philip, Schmidt, Craig W., Komm, Dennis, Pimentel, Tiago |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tokenisation over Bounded Alphabets is Hard
di: Kastreva, Violeta, et al.
Pubblicazione: (2025)
di: Kastreva, Violeta, et al.
Pubblicazione: (2025)
Tokenisation is NP-Complete
di: Whittington, Philip, et al.
Pubblicazione: (2024)
di: Whittington, Philip, et al.
Pubblicazione: (2024)
Causal Estimation of Tokenisation Bias
di: Lesci, Pietro, et al.
Pubblicazione: (2025)
di: Lesci, Pietro, et al.
Pubblicazione: (2025)
Convergence and Divergence of Language Models under Different Random Seeds
di: Fehlauer, Finlay, et al.
Pubblicazione: (2025)
di: Fehlauer, Finlay, et al.
Pubblicazione: (2025)
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
di: Ko, Jongwoo, et al.
Pubblicazione: (2026)
di: Ko, Jongwoo, et al.
Pubblicazione: (2026)
Stabilizing Policy Optimization via Logits Convexity
di: Chen, Hongzhan, et al.
Pubblicazione: (2026)
di: Chen, Hongzhan, et al.
Pubblicazione: (2026)
Do Generalisation Results Generalise?
di: Boglioni, Matteo, et al.
Pubblicazione: (2025)
di: Boglioni, Matteo, et al.
Pubblicazione: (2025)
On the Effect of (Near) Duplicate Subwords in Language Modelling
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
Exploring Anti-Aging Literature via ConvexTopics and Large Language Models
di: Yeganova, Lana E., et al.
Pubblicazione: (2026)
di: Yeganova, Lana E., et al.
Pubblicazione: (2026)
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
di: Xu, Yuzhuang, et al.
Pubblicazione: (2024)
di: Xu, Yuzhuang, et al.
Pubblicazione: (2024)
Stochasticity in Tokenisation Improves Robustness
di: Steger, Sophie, et al.
Pubblicazione: (2026)
di: Steger, Sophie, et al.
Pubblicazione: (2026)
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
di: Bae, Sangmin, et al.
Pubblicazione: (2024)
di: Bae, Sangmin, et al.
Pubblicazione: (2024)
Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping
di: Bystrich, Tobias, et al.
Pubblicazione: (2026)
di: Bystrich, Tobias, et al.
Pubblicazione: (2026)
LLM-OptiRA: LLM-Driven Optimization of Resource Allocation for Non-Convex Problems in Wireless Communications
di: Peng, Xinyue, et al.
Pubblicazione: (2025)
di: Peng, Xinyue, et al.
Pubblicazione: (2025)
Forecasting Events in Soccer Matches Through Language
di: Mendes-Neves, Tiago, et al.
Pubblicazione: (2024)
di: Mendes-Neves, Tiago, et al.
Pubblicazione: (2024)
ByteSpan: Information-Driven Subword Tokenisation
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
Zamba: A Compact 7B SSM Hybrid Model
di: Glorioso, Paolo, et al.
Pubblicazione: (2024)
di: Glorioso, Paolo, et al.
Pubblicazione: (2024)
The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
di: Dickson, Craig
Pubblicazione: (2025)
di: Dickson, Craig
Pubblicazione: (2025)
Deep learning models for representing out-of-vocabulary words
di: Lochter, Johannes V., et al.
Pubblicazione: (2020)
di: Lochter, Johannes V., et al.
Pubblicazione: (2020)
Self-Calibrating Language Models via Test-Time Discriminative Distillation
di: Hedna, Mohamed Rissal, et al.
Pubblicazione: (2026)
di: Hedna, Mohamed Rissal, et al.
Pubblicazione: (2026)
The Zamba2 Suite: Technical Report
di: Glorioso, Paolo, et al.
Pubblicazione: (2024)
di: Glorioso, Paolo, et al.
Pubblicazione: (2024)
Multi-State-Action Tokenisation in Decision Transformers for Multi-Discrete Action Spaces
di: Moodley, Perusha, et al.
Pubblicazione: (2024)
di: Moodley, Perusha, et al.
Pubblicazione: (2024)
Active Learning for Identifying Disaster-Related Tweets: A Comparison with Keyword Filtering and Generic Fine-Tuning
di: Hanny, David, et al.
Pubblicazione: (2024)
di: Hanny, David, et al.
Pubblicazione: (2024)
Large Scale Transfer Learning for Tabular Data via Language Modeling
di: Gardner, Josh, et al.
Pubblicazione: (2024)
di: Gardner, Josh, et al.
Pubblicazione: (2024)
Dissecting Linear Recurrent Models: How Different Gating Strategies Drive Selectivity and Generalization
di: Bouhadjar, Younes, et al.
Pubblicazione: (2026)
di: Bouhadjar, Younes, et al.
Pubblicazione: (2026)
Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
di: Bu, Zhiqi, et al.
Pubblicazione: (2026)
di: Bu, Zhiqi, et al.
Pubblicazione: (2026)
BroRL: Scaling Reinforcement Learning via Broadened Exploration
di: Hu, Jian, et al.
Pubblicazione: (2025)
di: Hu, Jian, et al.
Pubblicazione: (2025)
Resolving Discrepancies in Compute-Optimal Scaling of Language Models
di: Porian, Tomer, et al.
Pubblicazione: (2024)
di: Porian, Tomer, et al.
Pubblicazione: (2024)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
di: Lee, Tzu-Yun, et al.
Pubblicazione: (2025)
di: Lee, Tzu-Yun, et al.
Pubblicazione: (2025)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
di: He, Mutian, et al.
Pubblicazione: (2025)
di: He, Mutian, et al.
Pubblicazione: (2025)
A Bayesian Interpretation of Adaptive Low-Rank Adaptation
di: Chen, Haolin, et al.
Pubblicazione: (2024)
di: Chen, Haolin, et al.
Pubblicazione: (2024)
Disentangling Latent Shifts of In-Context Learning with Weak Supervision
di: Jukić, Josip, et al.
Pubblicazione: (2024)
di: Jukić, Josip, et al.
Pubblicazione: (2024)
DIVERSED: Relaxed Speculative Decoding via Dynamic Ensemble Verification
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
Understanding Addition and Subtraction in Transformers
di: Quirke, Philip, et al.
Pubblicazione: (2024)
di: Quirke, Philip, et al.
Pubblicazione: (2024)
TATA: Stance Detection via Topic-Agnostic and Topic-Aware Embeddings
di: Hanley, Hans W. A., et al.
Pubblicazione: (2023)
di: Hanley, Hans W. A., et al.
Pubblicazione: (2023)
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
di: Petrov, Aleksandar, et al.
Pubblicazione: (2023)
di: Petrov, Aleksandar, et al.
Pubblicazione: (2023)
SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
di: Wu, Jianghao, et al.
Pubblicazione: (2025)
di: Wu, Jianghao, et al.
Pubblicazione: (2025)
Tethered Reasoning: Decoupling Entropy from Hallucination in Quantized LLMs via Manifold Steering
di: Atkinson, Craig
Pubblicazione: (2026)
di: Atkinson, Craig
Pubblicazione: (2026)
Datasets, Documents, and Repetitions: The Practicalities of Unequal Data Quality
di: Fang, Alex, et al.
Pubblicazione: (2025)
di: Fang, Alex, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Tokenisation over Bounded Alphabets is Hard
di: Kastreva, Violeta, et al.
Pubblicazione: (2025) -
Tokenisation is NP-Complete
di: Whittington, Philip, et al.
Pubblicazione: (2024) -
Causal Estimation of Tokenisation Bias
di: Lesci, Pietro, et al.
Pubblicazione: (2025) -
Convergence and Divergence of Language Models under Different Random Seeds
di: Fehlauer, Finlay, et al.
Pubblicazione: (2025) -
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
di: Ko, Jongwoo, et al.
Pubblicazione: (2026)