Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
Fuente:
arXiv
Salvato in:
| Autori principali: | Schmidt, Craig W., Reddy, Varshini, Tanner, Chris, Pinter, Yuval |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tokenization Is More Than Compression
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
di: Fang, Xi, et al.
Pubblicazione: (2025)
di: Fang, Xi, et al.
Pubblicazione: (2025)
Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
di: Li, Shengrui, et al.
Pubblicazione: (2026)
di: Li, Shengrui, et al.
Pubblicazione: (2026)
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
di: Zhang, Zhehao, et al.
Pubblicazione: (2025)
di: Zhang, Zhehao, et al.
Pubblicazione: (2025)
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2024)
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2024)
Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context
di: An, Tao
Pubblicazione: (2025)
di: An, Tao
Pubblicazione: (2025)
One Agent to Serve All: a Lite-Adaptive Stylized AI Assistant for Millions of Multi-Style Official Accounts
di: Fan, Xingyu, et al.
Pubblicazione: (2025)
di: Fan, Xingyu, et al.
Pubblicazione: (2025)
TREX: Tokenizer Regression for Optimal Data Mixture
di: Won, Inho, et al.
Pubblicazione: (2026)
di: Won, Inho, et al.
Pubblicazione: (2026)
Language corpora for the Dutch medical domain
di: van Es, B.
Pubblicazione: (2026)
di: van Es, B.
Pubblicazione: (2026)
ADE: Adaptive Dictionary Embeddings -- Scaling Multi-Anchor Representations to Large Language Models
di: Demirci, Orhan, et al.
Pubblicazione: (2026)
di: Demirci, Orhan, et al.
Pubblicazione: (2026)
MemeLens: Multilingual Multitask VLMs for Memes
di: Shahroor, Ali Ezzat, et al.
Pubblicazione: (2026)
di: Shahroor, Ali Ezzat, et al.
Pubblicazione: (2026)
Rapid Biomedical Research Classification: The Pandemic PACT Advanced Categorisation Engine
di: Rohanian, Omid, et al.
Pubblicazione: (2024)
di: Rohanian, Omid, et al.
Pubblicazione: (2024)
Repairing Regex Vulnerabilities via Localization-Guided Instructions
di: Sung, Sicheol, et al.
Pubblicazione: (2025)
di: Sung, Sicheol, et al.
Pubblicazione: (2025)
Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders
di: Lequeu, Pierre-Antoine, et al.
Pubblicazione: (2026)
di: Lequeu, Pierre-Antoine, et al.
Pubblicazione: (2026)
Dynamic Demonstration Retrieval and Cognitive Understanding for Emotional Support Conversation
di: Xu, Zhe, et al.
Pubblicazione: (2024)
di: Xu, Zhe, et al.
Pubblicazione: (2024)
Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
di: Xu, Beining, et al.
Pubblicazione: (2025)
di: Xu, Beining, et al.
Pubblicazione: (2025)
A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents
di: Soman, Sumit, et al.
Pubblicazione: (2025)
di: Soman, Sumit, et al.
Pubblicazione: (2025)
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
di: Yu, Haeun, et al.
Pubblicazione: (2024)
di: Yu, Haeun, et al.
Pubblicazione: (2024)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
di: Alam, Firoj, et al.
Pubblicazione: (2025)
di: Alam, Firoj, et al.
Pubblicazione: (2025)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
di: Lee, Yejin, et al.
Pubblicazione: (2025)
di: Lee, Yejin, et al.
Pubblicazione: (2025)
LayerTracer: A Joint Task-Particle and Vulnerable-Layer Analysis framework for Arbitrary Large Language Model Architectures
di: Wu, Yuhang, et al.
Pubblicazione: (2026)
di: Wu, Yuhang, et al.
Pubblicazione: (2026)
CATER: Leveraging LLM to Pioneer a Multidimensional, Reference-Independent Paradigm in Translation Quality Evaluation
di: IIDA, Kurando, et al.
Pubblicazione: (2024)
di: IIDA, Kurando, et al.
Pubblicazione: (2024)
Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning
di: Ming, Xiaoyang, et al.
Pubblicazione: (2026)
di: Ming, Xiaoyang, et al.
Pubblicazione: (2026)
A Comprehensive Survey of Compression Algorithms for Language Models
di: Park, Seungcheol, et al.
Pubblicazione: (2024)
di: Park, Seungcheol, et al.
Pubblicazione: (2024)
Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
di: Feucht, Sheridan, et al.
Pubblicazione: (2026)
di: Feucht, Sheridan, et al.
Pubblicazione: (2026)
ProSwitch: Knowledge-Guided Instruction Tuning to Switch Between Professional and Non-Professional Responses
di: Zong, Chang, et al.
Pubblicazione: (2024)
di: Zong, Chang, et al.
Pubblicazione: (2024)
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
di: Lee, Yejin, et al.
Pubblicazione: (2025)
di: Lee, Yejin, et al.
Pubblicazione: (2025)
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
di: Tu, Songjun, et al.
Pubblicazione: (2025)
di: Tu, Songjun, et al.
Pubblicazione: (2025)
Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis
di: Huang, Donghao, et al.
Pubblicazione: (2026)
di: Huang, Donghao, et al.
Pubblicazione: (2026)
Efficient Adaptive Rejection Sampling for Accelerating Speculative Decoding in Large Language Models
di: Sun, Chendong, et al.
Pubblicazione: (2025)
di: Sun, Chendong, et al.
Pubblicazione: (2025)
Topeax -- An Improved Clustering Topic Model with Density Peak Detection and Lexical-Semantic Term Importance
di: Kardos, Márton
Pubblicazione: (2026)
di: Kardos, Márton
Pubblicazione: (2026)
Reducing Hallucinations in Summarization via Reinforcement Learning with Entity Hallucination Index
di: Katwe, Praveenkumar, et al.
Pubblicazione: (2025)
di: Katwe, Praveenkumar, et al.
Pubblicazione: (2025)
S2vNTM: Semi-supervised vMF Neural Topic Modeling
di: Xu, Weijie, et al.
Pubblicazione: (2023)
di: Xu, Weijie, et al.
Pubblicazione: (2023)
Structured Sentiment Analysis as Transition-based Dependency Graph Parsing
di: Fernández-González, Daniel
Pubblicazione: (2023)
di: Fernández-González, Daniel
Pubblicazione: (2023)
Bielik 11B v2 Technical Report
di: Ociepa, Krzysztof, et al.
Pubblicazione: (2025)
di: Ociepa, Krzysztof, et al.
Pubblicazione: (2025)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
di: Marcuzzo, Matteo, et al.
Pubblicazione: (2025)
di: Marcuzzo, Matteo, et al.
Pubblicazione: (2025)
Pun Unintended: LLMs and the Illusion of Humor Understanding
di: Zangari, Alessandro, et al.
Pubblicazione: (2025)
di: Zangari, Alessandro, et al.
Pubblicazione: (2025)
HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent
di: Xu, Weijie, et al.
Pubblicazione: (2024)
di: Xu, Weijie, et al.
Pubblicazione: (2024)
Multi-chain Graph Refinement and Selection for Reliable Reasoning in Large Language Models
di: Yang, Yujiao, et al.
Pubblicazione: (2025)
di: Yang, Yujiao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Tokenization Is More Than Compression
di: Schmidt, Craig W., et al.
Pubblicazione: (2024) -
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
di: Fang, Xi, et al.
Pubblicazione: (2025) -
Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
di: Li, Shengrui, et al.
Pubblicazione: (2026) -
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
di: Zhang, Zhehao, et al.
Pubblicazione: (2025) -
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2024)