When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Dadfar, Zachary Pedram |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
von: Wang, Qianli, et al.
Veröffentlicht: (2026)
von: Wang, Qianli, et al.
Veröffentlicht: (2026)
Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
von: Deng, Yihe, et al.
Veröffentlicht: (2023)
von: Deng, Yihe, et al.
Veröffentlicht: (2023)
Reference-based Metrics Disprove Themselves in Question Generation
von: Nguyen, Bang, et al.
Veröffentlicht: (2024)
von: Nguyen, Bang, et al.
Veröffentlicht: (2024)
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
von: Zelikman, Eric, et al.
Veröffentlicht: (2024)
von: Zelikman, Eric, et al.
Veröffentlicht: (2024)
LaMDA: Large Model Fine-Tuning via Spectrally Decomposed Low-Dimensional Adaptation
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2024)
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2024)
Fast Vocabulary Transfer for Language Model Compression
von: Gee, Leonidas, et al.
Veröffentlicht: (2024)
von: Gee, Leonidas, et al.
Veröffentlicht: (2024)
LLMCheckup: Conversational Examination of Large Language Models via Interpretability Tools and Self-Explanations
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
Lossless Vocabulary Reduction for Auto-Regressive Language Models
von: Chijiwa, Daiki, et al.
Veröffentlicht: (2025)
von: Chijiwa, Daiki, et al.
Veröffentlicht: (2025)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
Think When You Need: Self-Adaptive Chain-of-Thought Learning
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
von: Zhang, Jinbin, et al.
Veröffentlicht: (2025)
von: Zhang, Jinbin, et al.
Veröffentlicht: (2025)
Dynamic Vocabulary Pruning in Early-Exit LLMs
von: Vincenti, Jort, et al.
Veröffentlicht: (2024)
von: Vincenti, Jort, et al.
Veröffentlicht: (2024)
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
von: Khullar, Dipika, et al.
Veröffentlicht: (2026)
von: Khullar, Dipika, et al.
Veröffentlicht: (2026)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
von: Imai, Saki, et al.
Veröffentlicht: (2026)
von: Imai, Saki, et al.
Veröffentlicht: (2026)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
von: Roytburg, Dani, et al.
Veröffentlicht: (2025)
von: Roytburg, Dani, et al.
Veröffentlicht: (2025)
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
A Family of LLMs Liberated from Static Vocabularies
von: Alpha, Aleph, et al.
Veröffentlicht: (2026)
von: Alpha, Aleph, et al.
Veröffentlicht: (2026)
Rule by Rule: Learning with Confidence through Vocabulary Expansion
von: Nössig, Albert, et al.
Veröffentlicht: (2024)
von: Nössig, Albert, et al.
Veröffentlicht: (2024)
An Examination on the Effectiveness of Divide-and-Conquer Prompting in Large Language Models
von: Zhang, Yizhou, et al.
Veröffentlicht: (2024)
von: Zhang, Yizhou, et al.
Veröffentlicht: (2024)
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
von: Chen, Daiwei, et al.
Veröffentlicht: (2026)
von: Chen, Daiwei, et al.
Veröffentlicht: (2026)
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
von: Hayou, Soufiane, et al.
Veröffentlicht: (2025)
von: Hayou, Soufiane, et al.
Veröffentlicht: (2025)
When Bad Data Leads to Good Models
von: Li, Kenneth, et al.
Veröffentlicht: (2025)
von: Li, Kenneth, et al.
Veröffentlicht: (2025)
Post-Training Probability Manifold Correction via Structured SVD Pruning and Self-Referential Distillation
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
von: Fei, Wu, et al.
Veröffentlicht: (2025)
von: Fei, Wu, et al.
Veröffentlicht: (2025)
Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2025)
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2025)
Endogenous Resistance to Activation Steering in Language Models
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
Communicating Activations Between Language Model Agents
von: Ramesh, Vignav, et al.
Veröffentlicht: (2025)
von: Ramesh, Vignav, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
von: Zhang, Hugh, et al.
Veröffentlicht: (2024)
von: Zhang, Hugh, et al.
Veröffentlicht: (2024)
MrBERT: Modern Multilingual Encoders via Vocabulary, Domain, and Dimensional Adaptation
von: Tamayo, Daniel, et al.
Veröffentlicht: (2026)
von: Tamayo, Daniel, et al.
Veröffentlicht: (2026)
Vocabulary shapes cross-lingual variation of word-order learnability in language models
von: Martins, Jonas Mayer, et al.
Veröffentlicht: (2026)
von: Martins, Jonas Mayer, et al.
Veröffentlicht: (2026)
Learning a Generative Meta-Model of LLM Activations
von: Luo, Grace, et al.
Veröffentlicht: (2026)
von: Luo, Grace, et al.
Veröffentlicht: (2026)
Spectral Editing of Activations for Large Language Model Alignment
von: Qiu, Yifu, et al.
Veröffentlicht: (2024)
von: Qiu, Yifu, et al.
Veröffentlicht: (2024)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
When Attention Sink Emerges in Language Models: An Empirical View
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
AdaptThink: Reasoning Models Can Learn When to Think
von: Zhang, Jiajie, et al.
Veröffentlicht: (2025)
von: Zhang, Jiajie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
von: Wang, Qianli, et al.
Veröffentlicht: (2026) -
Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
von: Deng, Yihe, et al.
Veröffentlicht: (2023) -
Reference-based Metrics Disprove Themselves in Question Generation
von: Nguyen, Bang, et al.
Veröffentlicht: (2024) -
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
von: Zelikman, Eric, et al.
Veröffentlicht: (2024) -
LaMDA: Large Model Fine-Tuning via Spectrally Decomposed Low-Dimensional Adaptation
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2024)