Tokenization Preference for Human and Machine Learning Model: An Annotation Study
Fuente:
arXiv
Saved in:
| Main Authors: | Hiraoka, Tatsuya, Iwakura, Tomoya |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization
by: Tsuji, Kohei, et al.
Published: (2024)
by: Tsuji, Kohei, et al.
Published: (2024)
Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errors
by: Tsuji, Kohei, et al.
Published: (2025)
by: Tsuji, Kohei, et al.
Published: (2025)
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
by: Atuhurra, Jesse, et al.
Published: (2025)
by: Atuhurra, Jesse, et al.
Published: (2025)
Constructing Multilingual Visual-Text Datasets Revealing Visual Multilingual Ability of Vision Language Models
by: Atuhurra, Jesse, et al.
Published: (2024)
by: Atuhurra, Jesse, et al.
Published: (2024)
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
by: Hiraoka, Tatsuya, et al.
Published: (2025)
by: Hiraoka, Tatsuya, et al.
Published: (2025)
Repetition Neurons: How Do Language Models Produce Repetitions?
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
by: Doan, Nhi Hoai, et al.
Published: (2025)
by: Doan, Nhi Hoai, et al.
Published: (2025)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Bit-level BPE: Below the byte boundary
by: Moon, Sangwhan, et al.
Published: (2025)
by: Moon, Sangwhan, et al.
Published: (2025)
Aligning Large Language Model Behavior with Human Citation Preferences
by: Ando, Kenichiro, et al.
Published: (2026)
by: Ando, Kenichiro, et al.
Published: (2026)
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
by: Chen, Beiduo, et al.
Published: (2026)
by: Chen, Beiduo, et al.
Published: (2026)
Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling
by: Tan, Shaomu, et al.
Published: (2025)
by: Tan, Shaomu, et al.
Published: (2025)
Improving Named Entity Extraction Accuracy using Unlabeled Data and Several Extractors
by: Tomoya Iwakura
Published: (2009)
by: Tomoya Iwakura
Published: (2009)
Augmenting Dialog with Think-Aloud Utterances for Modeling Individual Personality Traits by LLM
by: Ishikura, Seiya, et al.
Published: (2025)
by: Ishikura, Seiya, et al.
Published: (2025)
The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning
by: Yang, Sen, et al.
Published: (2024)
by: Yang, Sen, et al.
Published: (2024)
Minority-Aware Satisfaction Estimation in Dialogue Systems via Preference-Adaptive Reinforcement Learning
by: Fu, Yahui, et al.
Published: (2025)
by: Fu, Yahui, et al.
Published: (2025)
Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments
by: Daynauth, Roland, et al.
Published: (2024)
by: Daynauth, Roland, et al.
Published: (2024)
A Survey on Human Preference Learning for Large Language Models
by: Jiang, Ruili, et al.
Published: (2024)
by: Jiang, Ruili, et al.
Published: (2024)
CTPD: Cross Tokenizer Preference Distillation
by: Nguyen, Truong, et al.
Published: (2026)
by: Nguyen, Truong, et al.
Published: (2026)
Diverging Preferences: When do Annotators Disagree and do Models Know?
by: Zhang, Michael JQ, et al.
Published: (2024)
by: Zhang, Michael JQ, et al.
Published: (2024)
Users as Annotators: LLM Preference Learning from Comparison Mode
by: Cai, Zhongze, et al.
Published: (2025)
by: Cai, Zhongze, et al.
Published: (2025)
Preference-grounded Token-level Guidance for Language Model Fine-tuning
by: Yang, Shentao, et al.
Published: (2023)
by: Yang, Shentao, et al.
Published: (2023)
Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data
by: Xia, Han, et al.
Published: (2024)
by: Xia, Han, et al.
Published: (2024)
Number Representations in LLMs: A Computational Parallel to Human Perception
by: AlquBoj, H. V., et al.
Published: (2025)
by: AlquBoj, H. V., et al.
Published: (2025)
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
Token-weighted Direct Preference Optimization with Attention
by: Huang, Chengyu, et al.
Published: (2026)
by: Huang, Chengyu, et al.
Published: (2026)
Reverse Engineering Human Preferences with Reinforcement Learning
by: Alazraki, Lisa, et al.
Published: (2025)
by: Alazraki, Lisa, et al.
Published: (2025)
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
by: Kasner, Zdeněk, et al.
Published: (2025)
by: Kasner, Zdeněk, et al.
Published: (2025)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
Enhancing the Preference Extractor in Multi-turn Dialogues: From Annotating Disasters to Accurate Preference Extraction
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
WorldPM: Scaling Human Preference Modeling
by: Wang, Binghai, et al.
Published: (2025)
by: Wang, Binghai, et al.
Published: (2025)
LRHP: Learning Representations for Human Preferences via Preference Pairs
by: Wang, Chenglong, et al.
Published: (2024)
by: Wang, Chenglong, et al.
Published: (2024)
SimulPL: Aligning Human Preferences in Simultaneous Machine Translation
by: Yu, Donglei, et al.
Published: (2025)
by: Yu, Donglei, et al.
Published: (2025)
Token-level Direct Preference Optimization
by: Zeng, Yongcheng, et al.
Published: (2024)
by: Zeng, Yongcheng, et al.
Published: (2024)
CARE: Multilingual Human Preference Learning for Cultural Awareness
by: Guo, Geyang, et al.
Published: (2025)
by: Guo, Geyang, et al.
Published: (2025)
Large Language Models as Annotators for Machine Translation Quality Estimation
by: Wang, Sidi, et al.
Published: (2026)
by: Wang, Sidi, et al.
Published: (2026)
Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay Perspective
by: Shao, Ruichen, et al.
Published: (2025)
by: Shao, Ruichen, et al.
Published: (2025)
Similar Items
-
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization
by: Tsuji, Kohei, et al.
Published: (2024) -
Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errors
by: Tsuji, Kohei, et al.
Published: (2025) -
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
by: Atuhurra, Jesse, et al.
Published: (2025) -
Constructing Multilingual Visual-Text Datasets Revealing Visual Multilingual Ability of Vision Language Models
by: Atuhurra, Jesse, et al.
Published: (2024) -
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024)