CUTE: Measuring LLMs' Understanding of Their Tokens
Fuente:
arXiv
Salvato in:
| Autori principali: | Edman, Lukas, Schmid, Helmut, Fraser, Alexander |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EXECUTE: A Multilingual Benchmark for LLM Token Understanding
di: Edman, Lukas, et al.
Pubblicazione: (2025)
di: Edman, Lukas, et al.
Pubblicazione: (2025)
Mask and You Shall Receive: Optimizing Masked Language Modeling For Pretraining BabyLMs
di: Edman, Lukas, et al.
Pubblicazione: (2025)
di: Edman, Lukas, et al.
Pubblicazione: (2025)
Are BabyLMs Second Language Learners?
di: Edman, Lukas, et al.
Pubblicazione: (2024)
di: Edman, Lukas, et al.
Pubblicazione: (2024)
Beyond Literal Token Overlap: Token Alignability for Multilinguality
di: Hämmerl, Katharina, et al.
Pubblicazione: (2025)
di: Hämmerl, Katharina, et al.
Pubblicazione: (2025)
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
di: Nie, Ercong, et al.
Pubblicazione: (2025)
di: Nie, Ercong, et al.
Pubblicazione: (2025)
CUTE: A Multilingual Dataset for Enhancing Cross-Lingual Knowledge Transfer in Low-Resource Languages
di: Zhuang, Wenhao, et al.
Pubblicazione: (2025)
di: Zhuang, Wenhao, et al.
Pubblicazione: (2025)
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
di: Ghorbanpour, Faeze, et al.
Pubblicazione: (2025)
di: Ghorbanpour, Faeze, et al.
Pubblicazione: (2025)
LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback
di: Lai, Wen, et al.
Pubblicazione: (2024)
di: Lai, Wen, et al.
Pubblicazione: (2024)
Understanding Cross-Lingual Alignment -- A Survey
di: Hämmerl, Katharina, et al.
Pubblicazione: (2024)
di: Hämmerl, Katharina, et al.
Pubblicazione: (2024)
Style-Specific Neurons for Steering LLMs in Text Style Transfer
di: Lai, Wen, et al.
Pubblicazione: (2024)
di: Lai, Wen, et al.
Pubblicazione: (2024)
Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine Translation
di: Edman, Lukas, et al.
Pubblicazione: (2023)
di: Edman, Lukas, et al.
Pubblicazione: (2023)
PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions
di: Niazi, Ruhallah, et al.
Pubblicazione: (2026)
di: Niazi, Ruhallah, et al.
Pubblicazione: (2026)
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks
di: Ma, Bolei, et al.
Pubblicazione: (2024)
di: Ma, Bolei, et al.
Pubblicazione: (2024)
Enhancing Character-Level Understanding in LLMs through Token Internal Structure Learning
di: Xu, Zhu, et al.
Pubblicazione: (2024)
di: Xu, Zhu, et al.
Pubblicazione: (2024)
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
di: Ghorbanpour, Faeze, et al.
Pubblicazione: (2025)
di: Ghorbanpour, Faeze, et al.
Pubblicazione: (2025)
Hate Personified: Investigating the role of LLMs in content moderation
di: Masud, Sarah, et al.
Pubblicazione: (2024)
di: Masud, Sarah, et al.
Pubblicazione: (2024)
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have
di: Hangya, Viktor, et al.
Pubblicazione: (2023)
di: Hangya, Viktor, et al.
Pubblicazione: (2023)
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
di: Hiraoka, Tatsuya, et al.
Pubblicazione: (2025)
di: Hiraoka, Tatsuya, et al.
Pubblicazione: (2025)
Behavior-Equivalent Token: Single-Token Replacement for Long Prompts in LLMs
di: Dong, Jiancheng, et al.
Pubblicazione: (2025)
di: Dong, Jiancheng, et al.
Pubblicazione: (2025)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
di: Yuan, Shuzhou, et al.
Pubblicazione: (2025)
di: Yuan, Shuzhou, et al.
Pubblicazione: (2025)
Why Are We Lonely? Leveraging LLMs to Measure and Understand Loneliness in Caregivers and Non-caregivers
di: Kim, Michelle Damin, et al.
Pubblicazione: (2026)
di: Kim, Michelle Damin, et al.
Pubblicazione: (2026)
Measuring Scalar Constructs in Social Science with LLMs
di: Licht, Hauke, et al.
Pubblicazione: (2025)
di: Licht, Hauke, et al.
Pubblicazione: (2025)
WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
di: Thede, Lukas, et al.
Pubblicazione: (2025)
di: Thede, Lukas, et al.
Pubblicazione: (2025)
Accelerating Production LLMs with Combined Token/Embedding Speculators
di: Wertheimer, Davis, et al.
Pubblicazione: (2024)
di: Wertheimer, Davis, et al.
Pubblicazione: (2024)
Optimizing Korean-Centric LLMs via Token Pruning
di: Kim, Hoyeol, et al.
Pubblicazione: (2026)
di: Kim, Hoyeol, et al.
Pubblicazione: (2026)
Beyond Tokens: Concept-Level Training Objectives for LLMs
di: Iyer, Laya, et al.
Pubblicazione: (2026)
di: Iyer, Laya, et al.
Pubblicazione: (2026)
CrossNews-UA: A Cross-lingual News Semantic Similarity Benchmark for Ukrainian, Polish, Russian, and English
di: Dementieva, Daryna, et al.
Pubblicazione: (2025)
di: Dementieva, Daryna, et al.
Pubblicazione: (2025)
EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainian
di: Dementieva, Daryna, et al.
Pubblicazione: (2025)
di: Dementieva, Daryna, et al.
Pubblicazione: (2025)
Language Model Re-rankers are Fooled by Lexical Similarities
di: Hagström, Lovisa, et al.
Pubblicazione: (2025)
di: Hagström, Lovisa, et al.
Pubblicazione: (2025)
Toward a Theory of Tokenization in LLMs
di: Rajaraman, Nived, et al.
Pubblicazione: (2024)
di: Rajaraman, Nived, et al.
Pubblicazione: (2024)
LLMs are Not Just Next Token Predictors
di: Downes, Stephen M., et al.
Pubblicazione: (2024)
di: Downes, Stephen M., et al.
Pubblicazione: (2024)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
di: Song, Yuhan, et al.
Pubblicazione: (2025)
di: Song, Yuhan, et al.
Pubblicazione: (2025)
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
di: Goyal, Agam, et al.
Pubblicazione: (2025)
di: Goyal, Agam, et al.
Pubblicazione: (2025)
How Language Directions Align with Token Geometry in Multilingual LLMs
di: Kim, JaeSeong, et al.
Pubblicazione: (2025)
di: Kim, JaeSeong, et al.
Pubblicazione: (2025)
Speculating LLMs' Chinese Training Data Pollution from Their Tokens
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
How does a Language-Specific Tokenizer affect LLMs?
di: Seo, Jean, et al.
Pubblicazione: (2025)
di: Seo, Jean, et al.
Pubblicazione: (2025)
Broken Words, Broken Performance: Effect of Tokenization on Performance of LLMs
di: Pawar, Sachin, et al.
Pubblicazione: (2025)
di: Pawar, Sachin, et al.
Pubblicazione: (2025)
Measuring Intrinsic Dimension of Token Embeddings
di: Kataiwa, Takuya, et al.
Pubblicazione: (2025)
di: Kataiwa, Takuya, et al.
Pubblicazione: (2025)
Do LLMs Understand Why We Write Diaries? A Method for Purpose Extraction and Clustering
di: Goloviznina, Valeriya, et al.
Pubblicazione: (2025)
di: Goloviznina, Valeriya, et al.
Pubblicazione: (2025)
Documenti analoghi
-
EXECUTE: A Multilingual Benchmark for LLM Token Understanding
di: Edman, Lukas, et al.
Pubblicazione: (2025) -
Mask and You Shall Receive: Optimizing Masked Language Modeling For Pretraining BabyLMs
di: Edman, Lukas, et al.
Pubblicazione: (2025) -
Are BabyLMs Second Language Learners?
di: Edman, Lukas, et al.
Pubblicazione: (2024) -
Beyond Literal Token Overlap: Token Alignability for Multilinguality
di: Hämmerl, Katharina, et al.
Pubblicazione: (2025) -
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
di: Nie, Ercong, et al.
Pubblicazione: (2025)