Distributional Properties of Subword Regularization
Fuente:
arXiv
Saved in:
| Main Authors: | Cognetta, Marco, Zouhar, Vilém, Okazaki, Naoaki |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Two Counterexamples to Tokenization and the Noiseless Channel
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
by: Zouhar, Vilém
Published: (2024)
by: Zouhar, Vilém
Published: (2024)
Tokenization as Finite-State Transduction
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Decoding-Free Sampling Strategies for LLM Marginalization
by: Pohl, David, et al.
Published: (2025)
by: Pohl, David, et al.
Published: (2025)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Quality and Quantity of Machine Translation References for Automatic Metrics
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
Pearmut: Human Evaluation of Translation Made Trivial
by: Zouhar, Vilém, et al.
Published: (2026)
by: Zouhar, Vilém, et al.
Published: (2026)
AI-Assisted Human Evaluation of Machine Translation
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
Multimodal Shannon Game with Images
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
Tutorial: $φ$-Transductions in OpenFst via the Gallic Semiring
by: Cognetta, Marco, et al.
Published: (2025)
by: Cognetta, Marco, et al.
Published: (2025)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
A Bayesian Optimization Approach to Machine Translation Reranking
by: Cheng, Julius, et al.
Published: (2024)
by: Cheng, Julius, et al.
Published: (2024)
Biased Tales: Cultural and Topic Bias in Generating Children's Stories
by: Rooein, Donya, et al.
Published: (2025)
by: Rooein, Donya, et al.
Published: (2025)
Evaluating Optimal Reference Translations
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
by: Ma, Youmi, et al.
Published: (2026)
by: Ma, Youmi, et al.
Published: (2026)
How to Engage Your Readers? Generating Guiding Questions to Promote Active Reading
by: Cui, Peng, et al.
Published: (2024)
by: Cui, Peng, et al.
Published: (2024)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
by: Sarti, Gabriele, et al.
Published: (2025)
by: Sarti, Gabriele, et al.
Published: (2025)
Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer
by: Ma, Youmi, et al.
Published: (2024)
by: Ma, Youmi, et al.
Published: (2024)
Estimating Machine Translation Difficulty
by: Proietti, Lorenzo, et al.
Published: (2025)
by: Proietti, Lorenzo, et al.
Published: (2025)
Social Bias Evaluation for Large Language Models Requires Prompt Variations
by: Hida, Rem, et al.
Published: (2024)
by: Hida, Rem, et al.
Published: (2024)
Bit-level BPE: Below the byte boundary
by: Moon, Sangwhan, et al.
Published: (2025)
by: Moon, Sangwhan, et al.
Published: (2025)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
by: Shiotani, Taihei, et al.
Published: (2026)
by: Shiotani, Taihei, et al.
Published: (2026)
Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach
by: Kaneko, Masahiro, et al.
Published: (2023)
by: Kaneko, Masahiro, et al.
Published: (2023)
SAIE Framework: Support Alone Isn't Enough -- Advancing LLM Training with Adversarial Remarks
by: Loem, Mengsay, et al.
Published: (2023)
by: Loem, Mengsay, et al.
Published: (2023)
Pitfalls and Outlooks in Using COMET
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
Early-Exit and Instant Confidence Translation Quality Estimation
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
Drifting Objectives for Refining Discrete Diffusion Language Models
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
Intent-Aware Self-Correction for Mitigating Social Biases in Large Language Models
by: Anantaprayoon, Panatchakorn, et al.
Published: (2025)
by: Anantaprayoon, Panatchakorn, et al.
Published: (2025)
Diffusion-State Policy Optimization for Masked Diffusion Language Models
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
Multilingual Performance Biases of Large Language Models in Education
by: Gupta, Vansh, et al.
Published: (2025)
by: Gupta, Vansh, et al.
Published: (2025)
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization
by: Tsuji, Kohei, et al.
Published: (2024)
by: Tsuji, Kohei, et al.
Published: (2024)
Sampling-based Pseudo-Likelihood for Membership Inference Attacks
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
Similar Items
-
Two Counterexamples to Tokenization and the Noiseless Channel
by: Cognetta, Marco, et al.
Published: (2024) -
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
by: Zouhar, Vilém
Published: (2024) -
Tokenization as Finite-State Transduction
by: Cognetta, Marco, et al.
Published: (2024) -
Decoding-Free Sampling Strategies for LLM Marginalization
by: Pohl, David, et al.
Published: (2025) -
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
by: Cognetta, Marco, et al.
Published: (2024)