Probing Subphonemes in Morphology Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Astrach, Gal, Pinter, Yuval |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hebrew Diacritics Restoration using Visual Representation
por: Elboher, Yair, et al.
Publicado: (2025)
por: Elboher, Yair, et al.
Publicado: (2025)
Information Types in Product Reviews
por: Shapira, Ori, et al.
Publicado: (2025)
por: Shapira, Ori, et al.
Publicado: (2025)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
por: Uzan, Omri, et al.
Publicado: (2025)
por: Uzan, Omri, et al.
Publicado: (2025)
Don't Touch My Diacritics
por: Gorman, Kyle, et al.
Publicado: (2024)
por: Gorman, Kyle, et al.
Publicado: (2024)
BiVert: Bidirectional Vocabulary Evaluation using Relations for Machine Translation
por: Cherf, Carinne, et al.
Publicado: (2024)
por: Cherf, Carinne, et al.
Publicado: (2024)
The Degree of Language Diacriticity and Its Effect on Tasks
por: Cohen, Adi, et al.
Publicado: (2026)
por: Cohen, Adi, et al.
Publicado: (2026)
Which Pieces Does Unigram Tokenization Really Need?
por: Land, Sander, et al.
Publicado: (2025)
por: Land, Sander, et al.
Publicado: (2025)
Token-Level Privacy in Large Language Models
por: Harel, Re'em, et al.
Publicado: (2025)
por: Harel, Re'em, et al.
Publicado: (2025)
Faster Superword Tokenization
por: Schmidt, Craig W., et al.
Publicado: (2026)
por: Schmidt, Craig W., et al.
Publicado: (2026)
Splintering Nonconcatenative Languages for Better Tokenization
por: Gazit, Bar, et al.
Publicado: (2025)
por: Gazit, Bar, et al.
Publicado: (2025)
Protecting Privacy in Classifiers by Token Manipulation
por: Harel, Re'em, et al.
Publicado: (2024)
por: Harel, Re'em, et al.
Publicado: (2024)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
por: Uzan, Omri, et al.
Publicado: (2024)
por: Uzan, Omri, et al.
Publicado: (2024)
OMPar: Automatic Parallelization with AI-Driven Source-to-Source Compilation
por: Kadosh, Tal, et al.
Publicado: (2024)
por: Kadosh, Tal, et al.
Publicado: (2024)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
por: Cognetta, Marco, et al.
Publicado: (2024)
por: Cognetta, Marco, et al.
Publicado: (2024)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
por: Reddy, Varshini, et al.
Publicado: (2025)
por: Reddy, Varshini, et al.
Publicado: (2025)
The Effect of Scripts and Formats on LLM Numeracy
por: Reddy, Varshini, et al.
Publicado: (2026)
por: Reddy, Varshini, et al.
Publicado: (2026)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
por: Schmidt, Craig W., et al.
Publicado: (2025)
por: Schmidt, Craig W., et al.
Publicado: (2025)
Sensitivity to Subphonemic Differences in First Language Predicts Vocabulary Size in a Foreign Language
por: Efthymia C. Kapnoula, et al.
Publicado: (2024)
por: Efthymia C. Kapnoula, et al.
Publicado: (2024)
Tokenization with Split Trees
por: Schmidt, Craig W., et al.
Publicado: (2026)
por: Schmidt, Craig W., et al.
Publicado: (2026)
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
por: Gal, Rinon, et al.
Publicado: (2024)
por: Gal, Rinon, et al.
Publicado: (2024)
Correction to “Sensitivity to Subphonemic Differences in First Language Predicts Vocabulary Size in a Foreign Language”
Publicado: (2026)
Publicado: (2026)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
por: Batsuren, Khuyagbaatar, et al.
Publicado: (2024)
por: Batsuren, Khuyagbaatar, et al.
Publicado: (2024)
IMPACT: Inflectional Morphology Probes Across Complex Typologies
por: Saeed, Mohammed J., et al.
Publicado: (2025)
por: Saeed, Mohammed J., et al.
Publicado: (2025)
MPIrigen: MPI Code Generation through Domain-Specific Language Models
por: Schneider, Nadav, et al.
Publicado: (2024)
por: Schneider, Nadav, et al.
Publicado: (2024)
Leveraging NTPs for Efficient Hallucination Detection in VLMs
por: Azachi, Ofir, et al.
Publicado: (2025)
por: Azachi, Ofir, et al.
Publicado: (2025)
Where Vision Becomes Text: Locating the OCR Routing Bottleneck in Vision-Language Models
por: Steinberg, Jonathan, et al.
Publicado: (2026)
por: Steinberg, Jonathan, et al.
Publicado: (2026)
LR-DWM: Efficient Watermarking for Diffusion Language Models
por: Raban, Ofek, et al.
Publicado: (2026)
por: Raban, Ofek, et al.
Publicado: (2026)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
por: Yona, Gal, et al.
Publicado: (2024)
por: Yona, Gal, et al.
Publicado: (2024)
Beyond Performance: Quantifying and Mitigating Label Bias in LLMs
por: Reif, Yuval, et al.
Publicado: (2024)
por: Reif, Yuval, et al.
Publicado: (2024)
Tokenization Is More Than Compression
por: Schmidt, Craig W., et al.
Publicado: (2024)
por: Schmidt, Craig W., et al.
Publicado: (2024)
Knowledge Graphs are Implicit Reward Models: Path-Derived Signals Enable Compositional Reasoning
por: Kansal, Yuval, et al.
Publicado: (2026)
por: Kansal, Yuval, et al.
Publicado: (2026)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
por: Kossen, Jannik, et al.
Publicado: (2024)
por: Kossen, Jannik, et al.
Publicado: (2024)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
por: Ovalle, Anaelia, et al.
Publicado: (2023)
por: Ovalle, Anaelia, et al.
Publicado: (2023)
The Roots of Performance Disparity in Multilingual Language Models: Intrinsic Modeling Difficulty or Design Choices?
por: Shani, Chen, et al.
Publicado: (2026)
por: Shani, Chen, et al.
Publicado: (2026)
Linear Relational Decoding of Morphology in Language Models
por: Xia, Eric, et al.
Publicado: (2025)
por: Xia, Eric, et al.
Publicado: (2025)
Confounding Factors in Relating Model Performance to Morphology
por: Poelman, Wessel, et al.
Publicado: (2025)
por: Poelman, Wessel, et al.
Publicado: (2025)
Labeled Morphological Segmentation with Semi-Markov Models
por: Cotterell, Ryan, et al.
Publicado: (2024)
por: Cotterell, Ryan, et al.
Publicado: (2024)
Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark
por: Blevins, Terra, et al.
Publicado: (2026)
por: Blevins, Terra, et al.
Publicado: (2026)
Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark
por: Mayhew, Stephen, et al.
Publicado: (2023)
por: Mayhew, Stephen, et al.
Publicado: (2023)
Language Models Change Facts Based on the Way You Talk
por: Kearney, Matthew, et al.
Publicado: (2025)
por: Kearney, Matthew, et al.
Publicado: (2025)
Ejemplares similares
-
Hebrew Diacritics Restoration using Visual Representation
por: Elboher, Yair, et al.
Publicado: (2025) -
Information Types in Product Reviews
por: Shapira, Ori, et al.
Publicado: (2025) -
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
por: Uzan, Omri, et al.
Publicado: (2025) -
Don't Touch My Diacritics
por: Gorman, Kyle, et al.
Publicado: (2024) -
BiVert: Bidirectional Vocabulary Evaluation using Relations for Machine Translation
por: Cherf, Carinne, et al.
Publicado: (2024)