Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ovalle, Anaelia, Mehrabi, Ninareh, Goyal, Palash, Dhamala, Jwala, Chang, Kai-Wei, Zemel, Richard, Galstyan, Aram, Pinter, Yuval, Gupta, Rahul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tree-of-Traversals: A Zero-Shot Reasoning Algorithm for Augmenting Black-box Language Models with Knowledge Graphs
von: Markowitz, Elan, et al.
Veröffentlicht: (2024)
von: Markowitz, Elan, et al.
Veröffentlicht: (2024)
On the steerability of large language models toward data-driven personas
von: Li, Junyi, et al.
Veröffentlicht: (2023)
von: Li, Junyi, et al.
Veröffentlicht: (2023)
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
Attribute Controlled Fine-tuning for Large Language Models: A Case Study on Detoxification
von: Meng, Tao, et al.
Veröffentlicht: (2024)
von: Meng, Tao, et al.
Veröffentlicht: (2024)
FLIRT: Feedback Loop In-context Red Teaming
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2023)
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2023)
Kaleidoscopic Teaming in Multi Agent Simulations
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2025)
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2025)
Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
K-Edit: Language Model Editing with Contextual Knowledge Awareness
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
DECOR: Auditing LLM Deception via Information Manipulation Theory
von: Cai, Linyue, et al.
Veröffentlicht: (2026)
von: Cai, Linyue, et al.
Veröffentlicht: (2026)
Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a Time
von: Li, Huihan, et al.
Veröffentlicht: (2025)
von: Li, Huihan, et al.
Veröffentlicht: (2025)
SWAN: Semantic Watermarking with Abstract Meaning Representation
von: Ye, Ziping, et al.
Veröffentlicht: (2026)
von: Ye, Ziping, et al.
Veröffentlicht: (2026)
Which Pieces Does Unigram Tokenization Really Need?
von: Land, Sander, et al.
Veröffentlicht: (2025)
von: Land, Sander, et al.
Veröffentlicht: (2025)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
von: Uzan, Omri, et al.
Veröffentlicht: (2025)
von: Uzan, Omri, et al.
Veröffentlicht: (2025)
Faster Superword Tokenization
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026)
Protecting Privacy in Classifiers by Token Manipulation
von: Harel, Re'em, et al.
Veröffentlicht: (2024)
von: Harel, Re'em, et al.
Veröffentlicht: (2024)
Prompt Perturbation Consistency Learning for Robust Language Models
von: Qiang, Yao, et al.
Veröffentlicht: (2024)
von: Qiang, Yao, et al.
Veröffentlicht: (2024)
Token-Level Privacy in Large Language Models
von: Harel, Re'em, et al.
Veröffentlicht: (2025)
von: Harel, Re'em, et al.
Veröffentlicht: (2025)
Splintering Nonconcatenative Languages for Better Tokenization
von: Gazit, Bar, et al.
Veröffentlicht: (2025)
von: Gazit, Bar, et al.
Veröffentlicht: (2025)
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Faithful Model Evaluation for Model-Based Metrics
von: Goyal, Palash, et al.
Veröffentlicht: (2023)
von: Goyal, Palash, et al.
Veröffentlicht: (2023)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
von: Reddy, Varshini, et al.
Veröffentlicht: (2025)
von: Reddy, Varshini, et al.
Veröffentlicht: (2025)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
von: Uzan, Omri, et al.
Veröffentlicht: (2024)
von: Uzan, Omri, et al.
Veröffentlicht: (2024)
Tokenization with Split Trees
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2026)
Tokenization Is More Than Compression
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
Strategize Globally, Adapt Locally: A Multi-Turn Red Teaming Agent with Dual-Level Learning
von: Chen, Si, et al.
Veröffentlicht: (2025)
von: Chen, Si, et al.
Veröffentlicht: (2025)
DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
von: Atil, Berk, et al.
Veröffentlicht: (2026)
von: Atil, Berk, et al.
Veröffentlicht: (2026)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
von: Batsuren, Khuyagbaatar, et al.
Veröffentlicht: (2024)
von: Batsuren, Khuyagbaatar, et al.
Veröffentlicht: (2024)
Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
von: Du, Yufeng, et al.
Veröffentlicht: (2026)
von: Du, Yufeng, et al.
Veröffentlicht: (2026)
Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token
von: Zychlinski, Shaked, et al.
Veröffentlicht: (2025)
von: Zychlinski, Shaked, et al.
Veröffentlicht: (2025)
The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2024)
von: Ovalle, Anaelia, et al.
Veröffentlicht: (2024)
ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
von: Liang, Jiacheng, et al.
Veröffentlicht: (2026)
von: Liang, Jiacheng, et al.
Veröffentlicht: (2026)
Asymmetric Phase Coding Audio Watermarking
von: Yang, Guang, et al.
Veröffentlicht: (2026)
von: Yang, Guang, et al.
Veröffentlicht: (2026)
FERRET: Framework for Expansion Reliant Red Teaming
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2026)
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2026)
Think Clearly: Improving Reasoning via Redundant Token Pruning
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Not Every Token Needs Forgetting: Selective Unlearning to Limit Change in Utility in Large Language Model Unlearning
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
Probing Subphonemes in Morphology Models
von: Astrach, Gal, et al.
Veröffentlicht: (2025)
von: Astrach, Gal, et al.
Veröffentlicht: (2025)
Hebrew Diacritics Restoration using Visual Representation
von: Elboher, Yair, et al.
Veröffentlicht: (2025)
von: Elboher, Yair, et al.
Veröffentlicht: (2025)
Don't Touch My Diacritics
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
von: Gorman, Kyle, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Tree-of-Traversals: A Zero-Shot Reasoning Algorithm for Augmenting Black-box Language Models with Knowledge Graphs
von: Markowitz, Elan, et al.
Veröffentlicht: (2024) -
On the steerability of large language models toward data-driven personas
von: Li, Junyi, et al.
Veröffentlicht: (2023) -
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024) -
Attribute Controlled Fine-tuning for Large Language Models: A Case Study on Detoxification
von: Meng, Tao, et al.
Veröffentlicht: (2024) -
FLIRT: Feedback Loop In-context Red Teaming
von: Mehrabi, Ninareh, et al.
Veröffentlicht: (2023)