Partially Rewriting a Transformer in Natural Language
Fuente:
arXiv
Saved in:
| Main Authors: | Paulo, Gonçalo, Belrose, Nora |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Mechanistic Anomaly Detection for "Quirky" Language Models
by: Johnston, David O., et al.
Published: (2025)
by: Johnston, David O., et al.
Published: (2025)
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024)
by: Marshall, Thomas, et al.
Published: (2024)
Transcoders Beat Sparse Autoencoders for Interpretability
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Eliciting Latent Knowledge from Quirky Language Models
by: Mallen, Alex, et al.
Published: (2023)
by: Mallen, Alex, et al.
Published: (2023)
LEACE: Perfect linear concept erasure in closed form
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
Small Language Models Improve Giants by Rewriting Their Outputs
by: Vernikos, Giorgos, et al.
Published: (2023)
by: Vernikos, Giorgos, et al.
Published: (2023)
Sample, Don't Search: Rethinking Test-Time Alignment for Language Models
by: Faria, Gonçalo, et al.
Published: (2025)
by: Faria, Gonçalo, et al.
Published: (2025)
Estimating the Probability of Sampling a Trained Neural Network at Random
by: Scherlis, Adam, et al.
Published: (2025)
by: Scherlis, Adam, et al.
Published: (2025)
Slowing Learning by Erasing Simple Features
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Converting MLPs into Polynomials in Closed Form
by: Belrose, Nora, et al.
Published: (2025)
by: Belrose, Nora, et al.
Published: (2025)
Balancing Label Quantity and Quality for Scalable Elicitation
by: Mallen, Alex, et al.
Published: (2024)
by: Mallen, Alex, et al.
Published: (2024)
Understanding Gradient Descent through the Training Jacobian
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
RAZOR: Sharpening Knowledge by Cutting Bias with Unsupervised Text Rewriting
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Transformer-based Model for ASR N-Best Rescoring and Rewriting
by: Kang, Iwen E., et al.
Published: (2024)
by: Kang, Iwen E., et al.
Published: (2024)
Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
PRewrite: Prompt Rewriting with Reinforcement Learning
by: Kong, Weize, et al.
Published: (2024)
by: Kong, Weize, et al.
Published: (2024)
Examining Two Hop Reasoning Through Information Content Scaling
by: Johnston, David, et al.
Published: (2025)
by: Johnston, David, et al.
Published: (2025)
Mind the Gap: Data Rewriting for Stable Off-Policy Supervised Fine-Tuning
by: Zhao, Shiwan, et al.
Published: (2025)
by: Zhao, Shiwan, et al.
Published: (2025)
Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT
by: Wang, Jiacheng, et al.
Published: (2026)
by: Wang, Jiacheng, et al.
Published: (2026)
Structured Style-Rewrite with Chain-of-Thought Planning for Low-Resource Character Dialogue
by: Zhu, Chanhui
Published: (2026)
by: Zhu, Chanhui
Published: (2026)
EntmaxKV: Support-Aware Decoding for Entmax Attention
by: Duarte, Gonçalo, et al.
Published: (2026)
by: Duarte, Gonçalo, et al.
Published: (2026)
ylmmcl at Multilingual Text Detoxification 2025: Lexicon-Guided Detoxification and Classifier-Gated Rewriting
by: Lai-Lopez, Nicole, et al.
Published: (2025)
by: Lai-Lopez, Nicole, et al.
Published: (2025)
Short-form Text Rewriting with Phi Silica
by: Tadimeti, Divya, et al.
Published: (2026)
by: Tadimeti, Divya, et al.
Published: (2026)
Reasoning-Grounded Natural Language Explanations for Language Models
by: Cahlik, Vojtech, et al.
Published: (2025)
by: Cahlik, Vojtech, et al.
Published: (2025)
Towards More Accurate Prediction of Human Empathy and Emotion in Text and Multi-turn Conversations by Combining Advanced NLP, Transformers-based Networks, and Linguistic Methodologies
by: Singh, Manisha, et al.
Published: (2024)
by: Singh, Manisha, et al.
Published: (2024)
Hypothesis-Conditioned Query Rewriting for Decision-Useful Retrieval
by: Chang, Hangeol, et al.
Published: (2026)
by: Chang, Hangeol, et al.
Published: (2026)
Multi-Target Cross-Lingual Summarization: a novel task and a language-neutral approach
by: Pernes, Diogo, et al.
Published: (2024)
by: Pernes, Diogo, et al.
Published: (2024)
Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering
by: Li, Yangyi, et al.
Published: (2025)
by: Li, Yangyi, et al.
Published: (2025)
Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
by: Hong, Joey, et al.
Published: (2025)
by: Hong, Joey, et al.
Published: (2025)
Conformal Prediction for Natural Language Processing: A Survey
by: Campos, Margarida M., et al.
Published: (2024)
by: Campos, Margarida M., et al.
Published: (2024)
Natural Language Processing and Multimodal Stock Price Prediction
by: Taylor, Kevin, et al.
Published: (2024)
by: Taylor, Kevin, et al.
Published: (2024)
Anatomical Heterogeneity in Transformer Language Models
by: Wietrzykowski, Tomasz
Published: (2026)
by: Wietrzykowski, Tomasz
Published: (2026)
QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
by: Cho, Nicole, et al.
Published: (2025)
by: Cho, Nicole, et al.
Published: (2025)
Binary Sparse Coding for Interpretability
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Similar Items
-
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024) -
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024) -
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025) -
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025) -
Mechanistic Anomaly Detection for "Quirky" Language Models
by: Johnston, David O., et al.
Published: (2025)