Saved in:
| Main Author: | Yadavalli, Bharadwaj |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.20683 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STU-PID: Steering Token Usage via PID Controller for Efficient Large Language Model Reasoning
by: Bharadwaj, Aryasomayajula Ram
Published: (2025)
by: Bharadwaj, Aryasomayajula Ram
Published: (2025)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
by: Goel, Raghavv, et al.
Published: (2025)
by: Goel, Raghavv, et al.
Published: (2025)
Dynamic Order Template Prediction for Generative Aspect-Based Sentiment Analysis
by: Jun, Yonghyun, et al.
Published: (2024)
by: Jun, Yonghyun, et al.
Published: (2024)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
by: Tian, Yuandong, et al.
Published: (2023)
by: Tian, Yuandong, et al.
Published: (2023)
Half the Nonlinearity Is Wasted: Measuring and Reallocating the Transformer's MLP Budget
by: Balogh, Peter
Published: (2026)
by: Balogh, Peter
Published: (2026)
Lost in Space: Finding the Right Tokens for Structured Output
by: Hamilton, Sil, et al.
Published: (2025)
by: Hamilton, Sil, et al.
Published: (2025)
Single-Position Intervention Fails: Distributed Output Templates Drive In-Context Learning
by: Cheng, Bryan, et al.
Published: (2026)
by: Cheng, Bryan, et al.
Published: (2026)
Topic Modelling: Going Beyond Token Outputs
by: Williams, Lowri, et al.
Published: (2024)
by: Williams, Lowri, et al.
Published: (2024)
SparseGrad: A Selective Method for Efficient Fine-tuning of MLP Layers
by: Chekalina, Viktoriia, et al.
Published: (2024)
by: Chekalina, Viktoriia, et al.
Published: (2024)
Weight Tying Biases Token Embeddings Towards the Output Space
by: Lopardo, Antonio, et al.
Published: (2026)
by: Lopardo, Antonio, et al.
Published: (2026)
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Understanding Token Probability Encoding in Output Embeddings
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
Memory-Efficient Fine-Tuning of Transformers via Token Selection
by: Simoulin, Antoine, et al.
Published: (2025)
by: Simoulin, Antoine, et al.
Published: (2025)
SENTRA: Selected-Next-Token Transformer for LLM Text Detection
by: Plyler, Mitchell, et al.
Published: (2025)
by: Plyler, Mitchell, et al.
Published: (2025)
Disentangling MLP Neuron Weights in Vocabulary Space
by: Avrahamy, Asaf, et al.
Published: (2026)
by: Avrahamy, Asaf, et al.
Published: (2026)
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
GPT-DETOX: An In-Context Learning-Based Paraphraser for Text Detoxification
by: Pesaranghader, Ali, et al.
Published: (2024)
by: Pesaranghader, Ali, et al.
Published: (2024)
HeceTokenizer: A Syllable-Based Tokenization Approach for Turkish Retrieval
by: Gulgonul, Senol
Published: (2026)
by: Gulgonul, Senol
Published: (2026)
What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels
by: Yadavalli, Aditya, et al.
Published: (2025)
by: Yadavalli, Aditya, et al.
Published: (2025)
Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation
by: Zhang, Ruiqi, et al.
Published: (2026)
by: Zhang, Ruiqi, et al.
Published: (2026)
Certain but not Probable? Differentiating Certainty from Probability in LLM Token Outputs for Probabilistic Scenarios
by: Toney-Wails, Autumn, et al.
Published: (2025)
by: Toney-Wails, Autumn, et al.
Published: (2025)
Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs
by: Venkatesh, Sohan
Published: (2026)
by: Venkatesh, Sohan
Published: (2026)
Detection and Measurement of Syntactic Templates in Generated Text
by: Shaib, Chantal, et al.
Published: (2024)
by: Shaib, Chantal, et al.
Published: (2024)
In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
by: Zhang, Ruiqi, et al.
Published: (2024)
by: Zhang, Ruiqi, et al.
Published: (2024)
DiffuMask: Diffusion Language Model for Token-level Prompt Pruning
by: Zheng, Caleb, et al.
Published: (2026)
by: Zheng, Caleb, et al.
Published: (2026)
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling
by: Huang, Hongzhi, et al.
Published: (2025)
by: Huang, Hongzhi, et al.
Published: (2025)
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
by: Yue, Ling, et al.
Published: (2026)
by: Yue, Ling, et al.
Published: (2026)
Token Masking Improves Transformer-Based Text Classification
by: Xu, Xianglong, et al.
Published: (2025)
by: Xu, Xianglong, et al.
Published: (2025)
A Better LLM Evaluator for Text Generation: The Impact of Prompt Output Sequencing and Optimization
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
TAB-PO: Preference Optimization with a Token-Level Adaptive Barrier for Token-Critical Structured Generation
by: Fodeh, Samah, et al.
Published: (2026)
by: Fodeh, Samah, et al.
Published: (2026)
Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
by: Cappelletti, Silvia, et al.
Published: (2025)
by: Cappelletti, Silvia, et al.
Published: (2025)
PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
by: Watts, Ishaan, et al.
Published: (2024)
by: Watts, Ishaan, et al.
Published: (2024)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
by: Sharma, Aman, et al.
Published: (2025)
by: Sharma, Aman, et al.
Published: (2025)
The Hidden Space of Safety: Understanding Preference-Tuned LLMs in Multilingual context
by: Verma, Nikhil, et al.
Published: (2025)
by: Verma, Nikhil, et al.
Published: (2025)
S3D: A Simple and Cost-Effective Self-Speculative Decoding Scheme for Low-Memory GPUs
by: Zhong, Wei, et al.
Published: (2024)
by: Zhong, Wei, et al.
Published: (2024)
Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation
by: Maye-Lasserre, Tom, et al.
Published: (2026)
by: Maye-Lasserre, Tom, et al.
Published: (2026)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
MLP-KAN: Unifying Deep Representation and Function Learning
by: He, Yunhong, et al.
Published: (2024)
by: He, Yunhong, et al.
Published: (2024)
MLP Memory: A Retriever-Pretrained Memory for Large Language Models
by: Wei, Rubin, et al.
Published: (2025)
by: Wei, Rubin, et al.
Published: (2025)
Similar Items
-
STU-PID: Steering Token Usage via PID Controller for Efficient Large Language Model Reasoning
by: Bharadwaj, Aryasomayajula Ram
Published: (2025) -
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
by: Goel, Raghavv, et al.
Published: (2025) -
Dynamic Order Template Prediction for Generative Aspect-Based Sentiment Analysis
by: Jun, Yonghyun, et al.
Published: (2024) -
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
by: Tian, Yuandong, et al.
Published: (2023) -
Half the Nonlinearity Is Wasted: Measuring and Reallocating the Transformer's MLP Budget
by: Balogh, Peter
Published: (2026)