Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Land, Sander, Bartolo, Max |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Which Pieces Does Unigram Tokenization Really Need?
di: Land, Sander, et al.
Pubblicazione: (2025)
di: Land, Sander, et al.
Pubblicazione: (2025)
Knowledge Editing for Large Language Model with Knowledge Neuronal Ensemble
di: Li, Yongchang, et al.
Pubblicazione: (2024)
di: Li, Yongchang, et al.
Pubblicazione: (2024)
Towards Effective and Efficient Continual Pre-training of Large Language Models
di: Chen, Jie, et al.
Pubblicazione: (2024)
di: Chen, Jie, et al.
Pubblicazione: (2024)
A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models
di: Alkan, Atilla Kaan, et al.
Pubblicazione: (2025)
di: Alkan, Atilla Kaan, et al.
Pubblicazione: (2025)
Review GIDE -- Restaurant Review Gastrointestinal Illness Detection and Extraction with Large Language Models
di: Laurence, Timothy, et al.
Pubblicazione: (2025)
di: Laurence, Timothy, et al.
Pubblicazione: (2025)
HInter: Exposing Hidden Intersectional Bias in Large Language Models
di: Souani, Badr, et al.
Pubblicazione: (2025)
di: Souani, Badr, et al.
Pubblicazione: (2025)
A Primer on Large Language Models and their Limitations
di: Johnson, Sandra, et al.
Pubblicazione: (2024)
di: Johnson, Sandra, et al.
Pubblicazione: (2024)
Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings
di: Goldin, Gili, et al.
Pubblicazione: (2024)
di: Goldin, Gili, et al.
Pubblicazione: (2024)
Evaluating Large Language Models for Public Health Classification and Extraction Tasks
di: Harris, Joshua, et al.
Pubblicazione: (2024)
di: Harris, Joshua, et al.
Pubblicazione: (2024)
Optimization Strategies for Enhancing Resource Efficiency in Transformers & Large Language Models
di: Wallace, Tom, et al.
Pubblicazione: (2025)
di: Wallace, Tom, et al.
Pubblicazione: (2025)
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
di: Andrylie, Lyzander Marciano, et al.
Pubblicazione: (2025)
di: Andrylie, Lyzander Marciano, et al.
Pubblicazione: (2025)
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models
di: Truong, Sang T., et al.
Pubblicazione: (2024)
di: Truong, Sang T., et al.
Pubblicazione: (2024)
LangMARL: Natural Language Multi-Agent Reinforcement Learning
di: Yao, Huaiyuan, et al.
Pubblicazione: (2026)
di: Yao, Huaiyuan, et al.
Pubblicazione: (2026)
The Superalignment of Superhuman Intelligence with Large Language Models
di: Huang, Minlie, et al.
Pubblicazione: (2024)
di: Huang, Minlie, et al.
Pubblicazione: (2024)
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2023)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2023)
An Unforgeable Publicly Verifiable Watermark for Large Language Models
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
Exploring State Tracking Capabilities of Large Language Models
di: Rezaee, Kiamehr, et al.
Pubblicazione: (2025)
di: Rezaee, Kiamehr, et al.
Pubblicazione: (2025)
Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
di: Li, Shengrui, et al.
Pubblicazione: (2026)
di: Li, Shengrui, et al.
Pubblicazione: (2026)
Distilling Large Language Models for Efficient Clinical Information Extraction
di: Vedula, Karthik S., et al.
Pubblicazione: (2024)
di: Vedula, Karthik S., et al.
Pubblicazione: (2024)
A Survey of Text Watermarking in the Era of Large Language Models
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
di: Liu, Aiwei, et al.
Pubblicazione: (2023)
Large Language Models Report Subjective Experience Under Self-Referential Processing
di: Berg, Cameron, et al.
Pubblicazione: (2025)
di: Berg, Cameron, et al.
Pubblicazione: (2025)
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
di: Bayram, M. Ali, et al.
Pubblicazione: (2024)
di: Bayram, M. Ali, et al.
Pubblicazione: (2024)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding -- A Survey
di: Fang, Xi, et al.
Pubblicazione: (2024)
di: Fang, Xi, et al.
Pubblicazione: (2024)
Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models
di: Safavi-Naini, Seyed Amir Ahmad, et al.
Pubblicazione: (2024)
di: Safavi-Naini, Seyed Amir Ahmad, et al.
Pubblicazione: (2024)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
di: Pan, Leyi, et al.
Pubblicazione: (2025)
di: Pan, Leyi, et al.
Pubblicazione: (2025)
The Paradox of Poetic Intent in Back-Translation: Evaluating the Quality of Large Language Models in Chinese Translation
di: Weigang, Li, et al.
Pubblicazione: (2025)
di: Weigang, Li, et al.
Pubblicazione: (2025)
Unveiling Attractor Cycles in Large Language Models: A Dynamical Systems View of Successive Paraphrasing
di: Wang, Zhilin, et al.
Pubblicazione: (2025)
di: Wang, Zhilin, et al.
Pubblicazione: (2025)
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
di: Yao, Ben, et al.
Pubblicazione: (2025)
di: Yao, Ben, et al.
Pubblicazione: (2025)
Transforming and Combining Rewards for Aligning Large Language Models
di: Wang, Zihao, et al.
Pubblicazione: (2024)
di: Wang, Zihao, et al.
Pubblicazione: (2024)
Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal
di: Oehri, Markus, et al.
Pubblicazione: (2025)
di: Oehri, Markus, et al.
Pubblicazione: (2025)
"As Eastern Powers, I will veto." : An Investigation of Nation-level Bias of Large Language Models in International Relations
di: Choi, Jonghyeon, et al.
Pubblicazione: (2025)
di: Choi, Jonghyeon, et al.
Pubblicazione: (2025)
ExpliCa: Evaluating Explicit Causal Reasoning in Large Language Models
di: Miliani, Martina, et al.
Pubblicazione: (2025)
di: Miliani, Martina, et al.
Pubblicazione: (2025)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
di: Fu, Tianyu, et al.
Pubblicazione: (2025)
di: Fu, Tianyu, et al.
Pubblicazione: (2025)
A Legal Framework for Natural Language Processing Model Training in Portugal
di: Almeida, Rúben, et al.
Pubblicazione: (2024)
di: Almeida, Rúben, et al.
Pubblicazione: (2024)
Fast Quiet-STaR: Thinking Without Thought Tokens
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
WaterSeeker: Pioneering Efficient Detection of Watermarked Segments in Large Documents
di: Pan, Leyi, et al.
Pubblicazione: (2024)
di: Pan, Leyi, et al.
Pubblicazione: (2024)
A Flexible Large Language Models Guardrail Development Methodology Applied to Off-Topic Prompt Detection
di: Chua, Gabriel, et al.
Pubblicazione: (2024)
di: Chua, Gabriel, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Which Pieces Does Unigram Tokenization Really Need?
di: Land, Sander, et al.
Pubblicazione: (2025) -
Knowledge Editing for Large Language Model with Knowledge Neuronal Ensemble
di: Li, Yongchang, et al.
Pubblicazione: (2024) -
Towards Effective and Efficient Continual Pre-training of Large Language Models
di: Chen, Jie, et al.
Pubblicazione: (2024) -
A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models
di: Alkan, Atilla Kaan, et al.
Pubblicazione: (2025) -
Review GIDE -- Restaurant Review Gastrointestinal Illness Detection and Extraction with Large Language Models
di: Laurence, Timothy, et al.
Pubblicazione: (2025)