Saved in:
| Main Author: | Nadaf, Mohammed Suhail B |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.02608 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
reward-lens: A Mechanistic Interpretability Library for Reward Models
by: Nadaf, Mohammed Suhail B
Published: (2026)
by: Nadaf, Mohammed Suhail B
Published: (2026)
Through a Steerable Lens: Magnifying Neural Network Interpretability via Phase-Based Extrapolation
by: Mahdisoltani, Farzaneh, et al.
Published: (2025)
by: Mahdisoltani, Farzaneh, et al.
Published: (2025)
Contextual Multinomial Logit Bandits with General Value Functions
by: Zhang, Mengxiao, et al.
Published: (2024)
by: Zhang, Mengxiao, et al.
Published: (2024)
Spectral Image Tokenizer
by: Esteves, Carlos, et al.
Published: (2024)
by: Esteves, Carlos, et al.
Published: (2024)
From Projection to Prediction: Beyond Logits for Scalable Language Models
by: Dong, Jianbing, et al.
Published: (2025)
by: Dong, Jianbing, et al.
Published: (2025)
Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
Beyond Logit Adjustment: A Residual Decomposition Framework for Long-Tailed Reranking
by: Wang, Zhanliang, et al.
Published: (2026)
by: Wang, Zhanliang, et al.
Published: (2026)
Logits-Based Finetuning
by: Li, Jingyao, et al.
Published: (2025)
by: Li, Jingyao, et al.
Published: (2025)
SLED: Self Logits Evolution Decoding for Improving Factuality in Large Language Models
by: Zhang, Jianyi, et al.
Published: (2024)
by: Zhang, Jianyi, et al.
Published: (2024)
Stable and Steerable Sparse Autoencoders with Weight Regularization
by: Jedryszek, Piotr, et al.
Published: (2026)
by: Jedryszek, Piotr, et al.
Published: (2026)
GPT-2 Through the Lens of Vector Symbolic Architectures
by: Knittel, Johannes, et al.
Published: (2024)
by: Knittel, Johannes, et al.
Published: (2024)
Beyond Hidden-Layer Manipulation: Semantically-Aware Logit Interventions for Debiasing LLMs
by: Xia, Wei
Published: (2025)
by: Xia, Wei
Published: (2025)
Steerable Neural ODEs on Homogeneous Spaces
by: Andersdotter, Emma, et al.
Published: (2026)
by: Andersdotter, Emma, et al.
Published: (2026)
Variance-Adaptive Optimal Algorithm for Reinforcement Learning with Multinomial Logit Function Approximation
by: Kim, Wonyoung, et al.
Published: (2026)
by: Kim, Wonyoung, et al.
Published: (2026)
Clifford-Steerable Convolutional Neural Networks
by: Zhdanov, Maksim, et al.
Published: (2024)
by: Zhdanov, Maksim, et al.
Published: (2024)
Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection
by: Ma, Mingyu Derek, et al.
Published: (2025)
by: Ma, Mingyu Derek, et al.
Published: (2025)
The Implicit Bias of Logit Regularization
by: Beck, Alon, et al.
Published: (2026)
by: Beck, Alon, et al.
Published: (2026)
Learning Steerable Clarification Policies with Collaborative Self-play
by: Berant, Jonathan, et al.
Published: (2025)
by: Berant, Jonathan, et al.
Published: (2025)
DeLTa: A Decoding Strategy based on Logit Trajectory Prediction Improves Factuality and Reasoning Ability
by: He, Yunzhen, et al.
Published: (2025)
by: He, Yunzhen, et al.
Published: (2025)
Learning Policy Representations for Steerable Behavior Synthesis
by: Li, Beiming, et al.
Published: (2026)
by: Li, Beiming, et al.
Published: (2026)
Generative Molecular Design with Steerable and Granular Synthesizability Control
by: Guo, Jeff, et al.
Published: (2025)
by: Guo, Jeff, et al.
Published: (2025)
A Probabilistic Approach to Learning the Degree of Equivariance in Steerable CNNs
by: Veefkind, Lars, et al.
Published: (2024)
by: Veefkind, Lars, et al.
Published: (2024)
Logits Poisoning Attack in Federated Distillation
by: Tang, Yuhan, et al.
Published: (2024)
by: Tang, Yuhan, et al.
Published: (2024)
DALD: Improving Logits-based Detector without Logits from Black-box LLMs
by: Zeng, Cong, et al.
Published: (2024)
by: Zeng, Cong, et al.
Published: (2024)
Beyond Real Data: Synthetic Data through the Lens of Regularization
by: Shidani, Amitis, et al.
Published: (2025)
by: Shidani, Amitis, et al.
Published: (2025)
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
by: Zhang, Yunfan, et al.
Published: (2025)
by: Zhang, Yunfan, et al.
Published: (2025)
Improving Actor-Critic Training with Steerable Action-Value Approximation Errors
by: Tasdighi, Bahareh, et al.
Published: (2024)
by: Tasdighi, Bahareh, et al.
Published: (2024)
Top-$nσ$: Not All Logits Are You Need
by: Tang, Chenxia, et al.
Published: (2024)
by: Tang, Chenxia, et al.
Published: (2024)
Exploiting Features and Logits in Heterogeneous Federated Learning
by: Chan, Yun-Hin, et al.
Published: (2022)
by: Chan, Yun-Hin, et al.
Published: (2022)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
by: Kulkarni, Akshay, et al.
Published: (2025)
by: Kulkarni, Akshay, et al.
Published: (2025)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
by: Li, Jin, et al.
Published: (2025)
by: Li, Jin, et al.
Published: (2025)
Bayesian Preference Learning for Test-Time Steerable Reward Models
by: Hong, Jiwoo, et al.
Published: (2026)
by: Hong, Jiwoo, et al.
Published: (2026)
Building a Scalable, Effective, and Steerable Search and Ranking Platform
by: Celikik, Marjan, et al.
Published: (2024)
by: Celikik, Marjan, et al.
Published: (2024)
Steerable Scene Generation with Post Training and Inference-Time Search
by: Pfaff, Nicholas, et al.
Published: (2025)
by: Pfaff, Nicholas, et al.
Published: (2025)
Logit Distance Bounds Representational Similarity
by: Nielsen, Beatrix M. G., et al.
Published: (2026)
by: Nielsen, Beatrix M. G., et al.
Published: (2026)
Logit Reweighting for Topic-Focused Summarization
by: Braun, Joschka, et al.
Published: (2025)
by: Braun, Joschka, et al.
Published: (2025)
Network Inversion of Binarised Neural Nets
by: Suhail, Pirzada, et al.
Published: (2024)
by: Suhail, Pirzada, et al.
Published: (2024)
EXP-CAM: Explanation Generation and Circuit Discovery Using Classifier Activation Matching
by: Suhail, Pirzada, et al.
Published: (2025)
by: Suhail, Pirzada, et al.
Published: (2025)
Compositional Generalization in Autoregressive Models via Logit Composition
by: Kumar, Aakash, et al.
Published: (2026)
by: Kumar, Aakash, et al.
Published: (2026)
Online Continual Learning via Logit Adjusted Softmax
by: Huang, Zhehao, et al.
Published: (2023)
by: Huang, Zhehao, et al.
Published: (2023)
Similar Items
-
reward-lens: A Mechanistic Interpretability Library for Reward Models
by: Nadaf, Mohammed Suhail B
Published: (2026) -
Through a Steerable Lens: Magnifying Neural Network Interpretability via Phase-Based Extrapolation
by: Mahdisoltani, Farzaneh, et al.
Published: (2025) -
Contextual Multinomial Logit Bandits with General Value Functions
by: Zhang, Mengxiao, et al.
Published: (2024) -
Spectral Image Tokenizer
by: Esteves, Carlos, et al.
Published: (2024) -
From Projection to Prediction: Beyond Logits for Scalable Language Models
by: Dong, Jianbing, et al.
Published: (2025)