Guardado en:
| Autores principales: | Mehra, Somesh, Garcia, Javier Alonso, Mauch, Lukas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2502.09419 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The pitfalls of next-token prediction
por: Bachmann, Gregor, et al.
Publicado: (2024)
por: Bachmann, Gregor, et al.
Publicado: (2024)
Language models are better than humans at next-token prediction
por: Shlegeris, Buck, et al.
Publicado: (2022)
por: Shlegeris, Buck, et al.
Publicado: (2022)
GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
por: Choudhary, Anand, et al.
Publicado: (2025)
por: Choudhary, Anand, et al.
Publicado: (2025)
Where is the signal in tokenization space?
por: Geh, Renato Lui, et al.
Publicado: (2024)
por: Geh, Renato Lui, et al.
Publicado: (2024)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
por: Kim, Taesu, et al.
Publicado: (2024)
por: Kim, Taesu, et al.
Publicado: (2024)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
por: Kumar, Anshul
Publicado: (2026)
por: Kumar, Anshul
Publicado: (2026)
Looking beyond the next token
por: Thankaraj, Abitha, et al.
Publicado: (2025)
por: Thankaraj, Abitha, et al.
Publicado: (2025)
Visualizing token importance for black-box language models
por: Rauba, Paulius, et al.
Publicado: (2025)
por: Rauba, Paulius, et al.
Publicado: (2025)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
por: Singh, Aaditya K., et al.
Publicado: (2024)
por: Singh, Aaditya K., et al.
Publicado: (2024)
Do language models plan ahead for future tokens?
por: Wu, Wilson, et al.
Publicado: (2024)
por: Wu, Wilson, et al.
Publicado: (2024)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
por: Nagarajan, Vaishnavh, et al.
Publicado: (2025)
por: Nagarajan, Vaishnavh, et al.
Publicado: (2025)
Statistical multi-metric evaluation and visualization of LLM system predictive performance
por: Ackerman, Samuel, et al.
Publicado: (2025)
por: Ackerman, Samuel, et al.
Publicado: (2025)
Global-Order GFlowNets
por: Pastor-Pérez, Lluís, et al.
Publicado: (2025)
por: Pastor-Pérez, Lluís, et al.
Publicado: (2025)
Byte-token Enhanced Language Models for Temporal Point Processes Analysis
por: Kong, Quyu, et al.
Publicado: (2025)
por: Kong, Quyu, et al.
Publicado: (2025)
Shaping capabilities with token-level data filtering
por: Rathi, Neil, et al.
Publicado: (2026)
por: Rathi, Neil, et al.
Publicado: (2026)
Scaling Transformer to 1M tokens and beyond with RMT
por: Bulatov, Aydar, et al.
Publicado: (2023)
por: Bulatov, Aydar, et al.
Publicado: (2023)
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
por: Zhao, Yize, et al.
Publicado: (2024)
por: Zhao, Yize, et al.
Publicado: (2024)
Interpretable Next-token Prediction via the Generalized Induction Head
por: Kim, Eunji, et al.
Publicado: (2024)
por: Kim, Eunji, et al.
Publicado: (2024)
Sample-efficient LLM Optimization with Reset Replay
por: Liu, Zichuan, et al.
Publicado: (2025)
por: Liu, Zichuan, et al.
Publicado: (2025)
AtteSTNet -- An attention and subword tokenization based approach for code-switched text hate speech detection
por: Shingi, Geet, et al.
Publicado: (2021)
por: Shingi, Geet, et al.
Publicado: (2021)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
por: Kwek, Eugene, et al.
Publicado: (2025)
por: Kwek, Eugene, et al.
Publicado: (2025)
PreFT: Prefill-only finetuning for efficient inference
por: Lanpouthakoun, Andrew, et al.
Publicado: (2026)
por: Lanpouthakoun, Andrew, et al.
Publicado: (2026)
Essential-Web v1.0: 24T tokens of organized web data
por: AI, Essential, et al.
Publicado: (2025)
por: AI, Essential, et al.
Publicado: (2025)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
por: Marconato, Emanuele, et al.
Publicado: (2024)
por: Marconato, Emanuele, et al.
Publicado: (2024)
Assessing LLM Text Detection in Educational Contexts: Does Human Contribution Affect Detection?
por: Gehring, Lukas, et al.
Publicado: (2025)
por: Gehring, Lukas, et al.
Publicado: (2025)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
por: Kramár, János, et al.
Publicado: (2024)
por: Kramár, János, et al.
Publicado: (2024)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
por: Xu, Yijie, et al.
Publicado: (2025)
por: Xu, Yijie, et al.
Publicado: (2025)
Exploring space efficiency in a tree-based linear model for extreme multi-label classification
por: Lin, He-Zhe, et al.
Publicado: (2024)
por: Lin, He-Zhe, et al.
Publicado: (2024)
Transformers for molecular property prediction: Domain adaptation efficiently improves performance
por: Sultan, Afnan, et al.
Publicado: (2025)
por: Sultan, Afnan, et al.
Publicado: (2025)
Publicly-Detectable Watermarking for Language Models
por: Fairoze, Jaiden, et al.
Publicado: (2023)
por: Fairoze, Jaiden, et al.
Publicado: (2023)
LLM-based feature generation from text for interpretable machine learning
por: Balek, Vojtěch, et al.
Publicado: (2024)
por: Balek, Vojtěch, et al.
Publicado: (2024)
Are LLM-based methods good enough for detecting unfair terms of service?
por: Frasheri, Mirgita, et al.
Publicado: (2024)
por: Frasheri, Mirgita, et al.
Publicado: (2024)
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
por: Shin, Seungjun, et al.
Publicado: (2025)
por: Shin, Seungjun, et al.
Publicado: (2025)
Amortizing intractable inference in large language models
por: Hu, Edward J., et al.
Publicado: (2023)
por: Hu, Edward J., et al.
Publicado: (2023)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
por: Qiu, Zeju, et al.
Publicado: (2026)
por: Qiu, Zeju, et al.
Publicado: (2026)
Entropy trajectory shape predicts LLM reasoning reliability: A diagnostic study of uncertainty dynamics in chain-of-thought
por: Zhao, Xinghao
Publicado: (2026)
por: Zhao, Xinghao
Publicado: (2026)
Order-Preserving GFlowNets
por: Chen, Yihang, et al.
Publicado: (2023)
por: Chen, Yihang, et al.
Publicado: (2023)
MALADE: Orchestration of LLM-powered Agents with Retrieval Augmented Generation for Pharmacovigilance
por: Choi, Jihye, et al.
Publicado: (2024)
por: Choi, Jihye, et al.
Publicado: (2024)
Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuning
por: Tan, Qitao, et al.
Publicado: (2025)
por: Tan, Qitao, et al.
Publicado: (2025)
Kanana: Compute-efficient Bilingual Language Models
por: Kanana LLM Team, et al.
Publicado: (2025)
por: Kanana LLM Team, et al.
Publicado: (2025)
Ejemplares similares
-
The pitfalls of next-token prediction
por: Bachmann, Gregor, et al.
Publicado: (2024) -
Language models are better than humans at next-token prediction
por: Shlegeris, Buck, et al.
Publicado: (2022) -
GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
por: Choudhary, Anand, et al.
Publicado: (2025) -
Where is the signal in tokenization space?
por: Geh, Renato Lui, et al.
Publicado: (2024) -
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
por: Kim, Taesu, et al.
Publicado: (2024)