Gespeichert in:
| 1. Verfasser: | Ehrmanntraut, Anton |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.02841 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
New Encoders for German Trained from Scratch: Comparing ModernGBERT with Converted LLM2Vec Models
von: Wunderle, Julia, et al.
Veröffentlicht: (2025)
von: Wunderle, Julia, et al.
Veröffentlicht: (2025)
The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
von: Gienapp, Lukas, et al.
Veröffentlicht: (2025)
von: Gienapp, Lukas, et al.
Veröffentlicht: (2025)
Improving OCR for Historical Texts of Multiple Languages
von: Westerdijk, Hylke, et al.
Veröffentlicht: (2025)
von: Westerdijk, Hylke, et al.
Veröffentlicht: (2025)
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
von: Su, DiJia, et al.
Veröffentlicht: (2025)
von: Su, DiJia, et al.
Veröffentlicht: (2025)
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
Classifying German Language Proficiency Levels Using Large Language Models
von: Ahlers, Elias-Leander, et al.
Veröffentlicht: (2025)
von: Ahlers, Elias-Leander, et al.
Veröffentlicht: (2025)
Normalization of Lithuanian Text Using Regular Expressions
von: Kasparaitis, Pijus
Veröffentlicht: (2023)
von: Kasparaitis, Pijus
Veröffentlicht: (2023)
Named Entity Recognition of Historical Texts via Large Language Model
von: Zhang, Shibingfeng, et al.
Veröffentlicht: (2025)
von: Zhang, Shibingfeng, et al.
Veröffentlicht: (2025)
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding
von: Idrissi-Yaghir, Ahmad, et al.
Veröffentlicht: (2024)
von: Idrissi-Yaghir, Ahmad, et al.
Veröffentlicht: (2024)
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
von: Klöser, Lars, et al.
Veröffentlicht: (2024)
von: Klöser, Lars, et al.
Veröffentlicht: (2024)
AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3
von: Kashirskiy, Mark, et al.
Veröffentlicht: (2025)
von: Kashirskiy, Mark, et al.
Veröffentlicht: (2025)
Two Spelling Normalization Approaches Based on Large Language Models
von: Domingo, Miguel, et al.
Veröffentlicht: (2025)
von: Domingo, Miguel, et al.
Veröffentlicht: (2025)
Ideology Prediction of German Political Texts
von: Schneider, Sinclair, et al.
Veröffentlicht: (2026)
von: Schneider, Sinclair, et al.
Veröffentlicht: (2026)
The Lou Dataset -- Exploring the Impact of Gender-Fair Language in German Text Classification
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
LSTM-Based Text Generation: A Study on Historical Datasets
von: Hussein, Mustafa Abbas Hussein, et al.
Veröffentlicht: (2024)
von: Hussein, Mustafa Abbas Hussein, et al.
Veröffentlicht: (2024)
An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification
von: Rusli, Andre, et al.
Veröffentlicht: (2024)
von: Rusli, Andre, et al.
Veröffentlicht: (2024)
Problematic Tokens: Tokenizer Bias in Large Language Models
von: Yang, Jin, et al.
Veröffentlicht: (2024)
von: Yang, Jin, et al.
Veröffentlicht: (2024)
Research on Graph-Retrieval Augmented Generation Based on Historical Text Knowledge Graphs
von: Fan, Yang, et al.
Veröffentlicht: (2025)
von: Fan, Yang, et al.
Veröffentlicht: (2025)
What Kinds of Tokens Benefit from Distant Text? An Analysis on Long Context Language Modeling
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
Decoding the Past: Explainable Machine Learning Models for Dating Historical Texts
von: Pinto, Paulo J. N., et al.
Veröffentlicht: (2025)
von: Pinto, Paulo J. N., et al.
Veröffentlicht: (2025)
PolyNorm: Few-Shot LLM-Based Text Normalization for Text-to-Speech
von: Wong, Michel, et al.
Veröffentlicht: (2025)
von: Wong, Michel, et al.
Veröffentlicht: (2025)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
Token Masking Improves Transformer-Based Text Classification
von: Xu, Xianglong, et al.
Veröffentlicht: (2025)
von: Xu, Xianglong, et al.
Veröffentlicht: (2025)
Explainability-Based Token Replacement on LLM-Generated Text
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
Rethinking Tokenization: Crafting Better Tokenizers for Large Language Models
von: Yang, Jinbiao
Veröffentlicht: (2024)
von: Yang, Jinbiao
Veröffentlicht: (2024)
Interpretable Recognition of Cognitive Distortions in Natural Language Texts
von: Kolonin, Anton, et al.
Veröffentlicht: (2025)
von: Kolonin, Anton, et al.
Veröffentlicht: (2025)
Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models
von: Yan, Ruiyi, et al.
Veröffentlicht: (2025)
von: Yan, Ruiyi, et al.
Veröffentlicht: (2025)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
von: Mao, Yu, et al.
Veröffentlicht: (2025)
von: Mao, Yu, et al.
Veröffentlicht: (2025)
TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization
von: Ho, Luong, et al.
Veröffentlicht: (2025)
von: Ho, Luong, et al.
Veröffentlicht: (2025)
Adversarial Attacks on AI-Generated Text Detection Models: A Token Probability-Based Approach Using Embeddings
von: Kadhim, Ahmed K., et al.
Veröffentlicht: (2025)
von: Kadhim, Ahmed K., et al.
Veröffentlicht: (2025)
Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
von: An, Ruize, et al.
Veröffentlicht: (2025)
von: An, Ruize, et al.
Veröffentlicht: (2025)
Token and Span Classification for Entity Recognition in French Historical Encyclopedias
von: Moncla, Ludovic, et al.
Veröffentlicht: (2025)
von: Moncla, Ludovic, et al.
Veröffentlicht: (2025)
German Text Embedding Clustering Benchmark
von: Wehrli, Silvan, et al.
Veröffentlicht: (2024)
von: Wehrli, Silvan, et al.
Veröffentlicht: (2024)
Disentangling Reasoning Tokens and Boilerplate Tokens For Language Model Fine-tuning
von: Ye, Ziang, et al.
Veröffentlicht: (2024)
von: Ye, Ziang, et al.
Veröffentlicht: (2024)
UniMoT: Unified Molecule-Text Language Model with Discrete Token Representation
von: Guo, Shuhan, et al.
Veröffentlicht: (2024)
von: Guo, Shuhan, et al.
Veröffentlicht: (2024)
Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
von: Benito-Rodriguez, Éloïse, et al.
Veröffentlicht: (2025)
von: Benito-Rodriguez, Éloïse, et al.
Veröffentlicht: (2025)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
Optimizing Token Usage on Large Language Model Conversations Using the Design Structure Matrix
von: Alarcia, Ramon Maria Garcia, et al.
Veröffentlicht: (2024)
von: Alarcia, Ramon Maria Garcia, et al.
Veröffentlicht: (2024)
Kathleen: Oscillator-Based Byte-Level Text Classification Without Tokenization or Attention
von: Fountzoulas, George
Veröffentlicht: (2026)
von: Fountzoulas, George
Veröffentlicht: (2026)
Ähnliche Einträge
-
New Encoders for German Trained from Scratch: Comparing ModernGBERT with Converted LLM2Vec Models
von: Wunderle, Julia, et al.
Veröffentlicht: (2025) -
The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
von: Gienapp, Lukas, et al.
Veröffentlicht: (2025) -
Improving OCR for Historical Texts of Multiple Languages
von: Westerdijk, Hylke, et al.
Veröffentlicht: (2025) -
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
von: Su, DiJia, et al.
Veröffentlicht: (2025) -
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)