Beyond Hidden-Layer Manipulation: Semantically-Aware Logit Interventions for Debiasing LLMs
Fuente:
arXiv
Guardado en:
| Autor principal: | Xia, Wei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Projection to Prediction: Beyond Logits for Scalable Language Models
por: Dong, Jianbing, et al.
Publicado: (2025)
por: Dong, Jianbing, et al.
Publicado: (2025)
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models
por: Zhang, Yuheng, et al.
Publicado: (2026)
por: Zhang, Yuheng, et al.
Publicado: (2026)
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
por: Emadi, Seyed Morteza
Publicado: (2026)
por: Emadi, Seyed Morteza
Publicado: (2026)
Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization
por: Luo, Haocheng, et al.
Publicado: (2026)
por: Luo, Haocheng, et al.
Publicado: (2026)
Debiasing Kernel-Based Generative Models
por: Qin, Tian, et al.
Publicado: (2025)
por: Qin, Tian, et al.
Publicado: (2025)
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
por: Guan, Zhong, et al.
Publicado: (2026)
por: Guan, Zhong, et al.
Publicado: (2026)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
por: Schiekiera, Louis, et al.
Publicado: (2026)
por: Schiekiera, Louis, et al.
Publicado: (2026)
Debiasing Multilingual LLMs in Cross-lingual Latent Space
por: Peng, Qiwei, et al.
Publicado: (2025)
por: Peng, Qiwei, et al.
Publicado: (2025)
Isotonic Layer: A Unified Framework for Recommendation Calibration and Debiasing
por: Cheng, Hailing, et al.
Publicado: (2026)
por: Cheng, Hailing, et al.
Publicado: (2026)
Logit Distance Bounds Representational Similarity
por: Nielsen, Beatrix M. G., et al.
Publicado: (2026)
por: Nielsen, Beatrix M. G., et al.
Publicado: (2026)
Layer by Layer: Uncovering Hidden Representations in Language Models
por: Skean, Oscar, et al.
Publicado: (2025)
por: Skean, Oscar, et al.
Publicado: (2025)
Logit Dynamics in Softmax Policy Gradient Methods
por: Li, Yingru
Publicado: (2025)
por: Li, Yingru
Publicado: (2025)
The Hidden Power of Normalization Layers in Neural Networks: Exponential Capacity Control
por: Than, Khoat
Publicado: (2025)
por: Than, Khoat
Publicado: (2025)
Beyond Predictions in Neural ODEs: Identification and Interventions
por: Aliee, Hananeh, et al.
Publicado: (2021)
por: Aliee, Hananeh, et al.
Publicado: (2021)
SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion
por: Hu, Zizhao, et al.
Publicado: (2026)
por: Hu, Zizhao, et al.
Publicado: (2026)
Peak-Controlled Logits Poisoning Attack in Federated Distillation
por: Tang, Yuhan, et al.
Publicado: (2024)
por: Tang, Yuhan, et al.
Publicado: (2024)
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
por: Sturman, Olivia, et al.
Publicado: (2024)
por: Sturman, Olivia, et al.
Publicado: (2024)
Debiasing Machine Unlearning with Counterfactual Examples
por: Chen, Ziheng, et al.
Publicado: (2024)
por: Chen, Ziheng, et al.
Publicado: (2024)
Logit Distillation on Manifolds: Mapping by Learning
por: Yang, Yiru, et al.
Publicado: (2026)
por: Yang, Yiru, et al.
Publicado: (2026)
On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
por: Fan, Chenghao, et al.
Publicado: (2024)
por: Fan, Chenghao, et al.
Publicado: (2024)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
por: Li, Jin, et al.
Publicado: (2025)
por: Li, Jin, et al.
Publicado: (2025)
SCALA: Split Federated Learning with Concatenated Activations and Logit Adjustments
por: Yang, Jiarong, et al.
Publicado: (2024)
por: Yang, Jiarong, et al.
Publicado: (2024)
Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
por: Chen, Xinzhu, et al.
Publicado: (2025)
por: Chen, Xinzhu, et al.
Publicado: (2025)
Debiasing Reward Models by Representation Learning with Guarantees
por: Ng, Ignavier, et al.
Publicado: (2025)
por: Ng, Ignavier, et al.
Publicado: (2025)
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
por: Lee, Jihoon, et al.
Publicado: (2025)
por: Lee, Jihoon, et al.
Publicado: (2025)
Transformer Normalisation Layers and the Independence of Semantic Subspaces
por: Menary, Stephen, et al.
Publicado: (2024)
por: Menary, Stephen, et al.
Publicado: (2024)
Q-function Decomposition with Intervention Semantics with Factored Action Spaces
por: Lee, Junkyu, et al.
Publicado: (2025)
por: Lee, Junkyu, et al.
Publicado: (2025)
ReSAE: Residualized Sparse Autoencoders for Multi-Layer Transformer Interventions
por: Poduval, Prathyush, et al.
Publicado: (2026)
por: Poduval, Prathyush, et al.
Publicado: (2026)
time2time: Causal Intervention in Hidden States to Simulate Rare Events in Time Series Foundation Models
por: Sanyal, Debdeep, et al.
Publicado: (2025)
por: Sanyal, Debdeep, et al.
Publicado: (2025)
Formalising the Logit Shift Induced by LoRA: A Technical Note
por: Shi, Xiang, et al.
Publicado: (2026)
por: Shi, Xiang, et al.
Publicado: (2026)
Discovering Hidden Algebraic Structures via Transformers with Rank-Aware Beam GRPO
por: Lee, Jaeha, et al.
Publicado: (2025)
por: Lee, Jaeha, et al.
Publicado: (2025)
Fredformer: Frequency Debiased Transformer for Time Series Forecasting
por: Piao, Xihao, et al.
Publicado: (2024)
por: Piao, Xihao, et al.
Publicado: (2024)
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
por: Williams, Marcus, et al.
Publicado: (2024)
por: Williams, Marcus, et al.
Publicado: (2024)
CRAFT: Forgetting-Aware Intervention-Based Adaptation for Continual Learning
por: Hossen, Md Anwar, et al.
Publicado: (2026)
por: Hossen, Md Anwar, et al.
Publicado: (2026)
Learning to Receive Help: Intervention-Aware Concept Embedding Models
por: Zarlenga, Mateo Espinosa, et al.
Publicado: (2023)
por: Zarlenga, Mateo Espinosa, et al.
Publicado: (2023)
Model-Level GNN Explanations via Rule-to-Graph Readout for Logit Reconstruction
por: Lu, Shengyao, et al.
Publicado: (2025)
por: Lu, Shengyao, et al.
Publicado: (2025)
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
por: Janiak, Jett, et al.
Publicado: (2023)
por: Janiak, Jett, et al.
Publicado: (2023)
Neural Breadcrumbs: Membership Inference Attacks on LLMs Through Hidden State and Attention Pattern Analysis
por: Makhija, Disha, et al.
Publicado: (2025)
por: Makhija, Disha, et al.
Publicado: (2025)
Data-Free Pruning of Self-Attention Layers in LLMs
por: Saikumar, Dhananjay, et al.
Publicado: (2025)
por: Saikumar, Dhananjay, et al.
Publicado: (2025)
GRID: Graph-based Reasoning for Intervention and Discovery in Built Environments
por: Ehsan, Taqiya, et al.
Publicado: (2025)
por: Ehsan, Taqiya, et al.
Publicado: (2025)
Ejemplares similares
-
From Projection to Prediction: Beyond Logits for Scalable Language Models
por: Dong, Jianbing, et al.
Publicado: (2025) -
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models
por: Zhang, Yuheng, et al.
Publicado: (2026) -
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
por: Emadi, Seyed Morteza
Publicado: (2026) -
Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization
por: Luo, Haocheng, et al.
Publicado: (2026) -
Debiasing Kernel-Based Generative Models
por: Qin, Tian, et al.
Publicado: (2025)