All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
Fuente:
arXiv
Guardado en:
| Autores principales: | Mamidanna, Siddarth, Rai, Daking, Yao, Ziyu, Zhou, Yilun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
por: Rai, Daking, et al.
Publicado: (2024)
por: Rai, Daking, et al.
Publicado: (2024)
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
por: Rai, Daking, et al.
Publicado: (2024)
por: Rai, Daking, et al.
Publicado: (2024)
Mechanistic Understanding of Language Models in Syntactic Code Completion
por: Miller, Samuel, et al.
Publicado: (2025)
por: Miller, Samuel, et al.
Publicado: (2025)
Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
por: Rai, Daking, et al.
Publicado: (2025)
por: Rai, Daking, et al.
Publicado: (2025)
Data-driven Circuit Discovery for Interpretability of Language Models
por: Rai, Daking, et al.
Publicado: (2026)
por: Rai, Daking, et al.
Publicado: (2026)
Vocabulary Transfer for Biomedical Texts: Add Tokens if You Can Not Add Data
por: Singh, Priyanka, et al.
Publicado: (2022)
por: Singh, Priyanka, et al.
Publicado: (2022)
LaTIM: Measuring Latent Token-to-Token Interactions in Mamba Models
por: Pitorro, Hugo, et al.
Publicado: (2025)
por: Pitorro, Hugo, et al.
Publicado: (2025)
Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER
por: Ewais, Ahmed, et al.
Publicado: (2026)
por: Ewais, Ahmed, et al.
Publicado: (2026)
Egalitarian Language Representation in Language Models: It All Begins with Tokenizers
por: Velayuthan, Menan, et al.
Publicado: (2024)
por: Velayuthan, Menan, et al.
Publicado: (2024)
Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs
por: Feucht, Sheridan, et al.
Publicado: (2024)
por: Feucht, Sheridan, et al.
Publicado: (2024)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
por: Walker, Nicholas
Publicado: (2024)
por: Walker, Nicholas
Publicado: (2024)
Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect
por: Klerings, Alina, et al.
Publicado: (2025)
por: Klerings, Alina, et al.
Publicado: (2025)
Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5
por: Dang, Thao Anh, et al.
Publicado: (2024)
por: Dang, Thao Anh, et al.
Publicado: (2024)
AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3
por: Kashirskiy, Mark, et al.
Publicado: (2025)
por: Kashirskiy, Mark, et al.
Publicado: (2025)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
por: Nikankin, Yaniv, et al.
Publicado: (2024)
por: Nikankin, Yaniv, et al.
Publicado: (2024)
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
por: Doda, Shravan
Publicado: (2026)
por: Doda, Shravan
Publicado: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Fin-ExBERT: User Intent based Text Extraction in Financial Context using Graph-Augmented BERT and trainable Plugin
por: Sarker, Soumick, et al.
Publicado: (2025)
por: Sarker, Soumick, et al.
Publicado: (2025)
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
por: Khanna, Danush, et al.
Publicado: (2025)
por: Khanna, Danush, et al.
Publicado: (2025)
UIPress: Bringing Optical Token Compression to UI-to-Code Generation
por: Dai, Dasen, et al.
Publicado: (2026)
por: Dai, Dasen, et al.
Publicado: (2026)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
por: Ashuach, Tomer, et al.
Publicado: (2025)
por: Ashuach, Tomer, et al.
Publicado: (2025)
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
por: Bayram, M. Ali, et al.
Publicado: (2025)
por: Bayram, M. Ali, et al.
Publicado: (2025)
Fast Quiet-STaR: Thinking Without Thought Tokens
por: Huang, Wei, et al.
Publicado: (2025)
por: Huang, Wei, et al.
Publicado: (2025)
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
por: Fagnou, Erwan, et al.
Publicado: (2026)
por: Fagnou, Erwan, et al.
Publicado: (2026)
SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM
por: Wang, Quandong, et al.
Publicado: (2024)
por: Wang, Quandong, et al.
Publicado: (2024)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2026)
por: Cho, Seonglae, et al.
Publicado: (2026)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
por: He, Yanjin, et al.
Publicado: (2025)
por: He, Yanjin, et al.
Publicado: (2025)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
por: Shravan, Rohan
Publicado: (2026)
por: Shravan, Rohan
Publicado: (2026)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
por: Liu, Aiwei, et al.
Publicado: (2024)
por: Liu, Aiwei, et al.
Publicado: (2024)
Beyond Token Length: Step Pruner for Efficient and Accurate Reasoning in Large Language Models
por: Wu, Canhui, et al.
Publicado: (2025)
por: Wu, Canhui, et al.
Publicado: (2025)
Tokenization Is More Than Compression
por: Schmidt, Craig W., et al.
Publicado: (2024)
por: Schmidt, Craig W., et al.
Publicado: (2024)
BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base
por: Shravan, Rohan
Publicado: (2026)
por: Shravan, Rohan
Publicado: (2026)
From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition
por: Aslam, Nazia, et al.
Publicado: (2026)
por: Aslam, Nazia, et al.
Publicado: (2026)
MIX : a Multi-task Learning Approach to Solve Open-Domain Question Answering
por: Chaybouti, Sofian, et al.
Publicado: (2020)
por: Chaybouti, Sofian, et al.
Publicado: (2020)
Math Natural Language Inference: this should be easy!
por: de Paiva, Valeria, et al.
Publicado: (2025)
por: de Paiva, Valeria, et al.
Publicado: (2025)
Next Token Prediction Is a Dead End for Creativity
por: Olatunji, Ibukun, et al.
Publicado: (2025)
por: Olatunji, Ibukun, et al.
Publicado: (2025)
COMET-poly: Machine Translation Metric Grounded in Other Candidates
por: Züfle, Maike, et al.
Publicado: (2025)
por: Züfle, Maike, et al.
Publicado: (2025)
Universal-2-TF: Robust All-Neural Text Formatting for ASR
por: Khare, Yash, et al.
Publicado: (2025)
por: Khare, Yash, et al.
Publicado: (2025)
Advancing Polish Language Modeling through Tokenizer Optimization in the Bielik v3 7B and 11B Series
por: Ociepa, Krzysztof, et al.
Publicado: (2026)
por: Ociepa, Krzysztof, et al.
Publicado: (2026)
Ejemplares similares
-
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
por: Rai, Daking, et al.
Publicado: (2024) -
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
por: Rai, Daking, et al.
Publicado: (2024) -
Mechanistic Understanding of Language Models in Syntactic Code Completion
por: Miller, Samuel, et al.
Publicado: (2025) -
Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
por: Rai, Daking, et al.
Publicado: (2025) -
Data-driven Circuit Discovery for Interpretability of Language Models
por: Rai, Daking, et al.
Publicado: (2026)