The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Yu, Tian, Yang, Ravfogel, Shauli, Sachan, Mrinmaya, Ash, Elliott, Hoyle, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Linear Adversarial Concept Erasure
by: Ravfogel, Shauli, et al.
Published: (2022)
by: Ravfogel, Shauli, et al.
Published: (2022)
Kernelized Concept Erasure
by: Ravfogel, Shauli, et al.
Published: (2022)
by: Ravfogel, Shauli, et al.
Published: (2022)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
by: Ni, Jingwei, et al.
Published: (2025)
by: Ni, Jingwei, et al.
Published: (2025)
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation
by: Ni, Jingwei, et al.
Published: (2024)
by: Ni, Jingwei, et al.
Published: (2024)
Emergence of Linear Truth Encodings in Language Models
by: Ravfogel, Shauli, et al.
Published: (2025)
by: Ravfogel, Shauli, et al.
Published: (2025)
Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
by: Chen, Hung-Ting, et al.
Published: (2025)
by: Chen, Hung-Ting, et al.
Published: (2025)
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
by: Ni, Jingwei, et al.
Published: (2024)
by: Ni, Jingwei, et al.
Published: (2024)
Log-linear Guardedness and its Implications
by: Ravfogel, Shauli, et al.
Published: (2022)
by: Ravfogel, Shauli, et al.
Published: (2022)
Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
by: Xiong, Chenfei, et al.
Published: (2025)
by: Xiong, Chenfei, et al.
Published: (2025)
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
by: Ben-Zaken, Elad, et al.
Published: (2021)
by: Ben-Zaken, Elad, et al.
Published: (2021)
Geometric Factual Recall in Transformers
by: Ravfogel, Shauli, et al.
Published: (2026)
by: Ravfogel, Shauli, et al.
Published: (2026)
Investigating the Zone of Proximal Development of Language Models for In-Context Learning
by: Cui, Peng, et al.
Published: (2025)
by: Cui, Peng, et al.
Published: (2025)
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
by: Hong, Yihuai, et al.
Published: (2024)
by: Hong, Yihuai, et al.
Published: (2024)
A Practical Method for Generating String Counterfactuals
by: Avitan, Matan, et al.
Published: (2024)
by: Avitan, Matan, et al.
Published: (2024)
State over Tokens: Characterizing the Role of Reasoning Tokens
by: Levy, Mosh, et al.
Published: (2025)
by: Levy, Mosh, et al.
Published: (2025)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
by: Shafran, Or, et al.
Published: (2026)
by: Shafran, Or, et al.
Published: (2026)
Gumbel Counterfactual Generation From Language Models
by: Ravfogel, Shauli, et al.
Published: (2024)
by: Ravfogel, Shauli, et al.
Published: (2024)
On Affine Homotopy between Language Encoders
by: Chan, Robin SM, et al.
Published: (2024)
by: Chan, Robin SM, et al.
Published: (2024)
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
by: Cohen, Amir DN, et al.
Published: (2024)
by: Cohen, Amir DN, et al.
Published: (2024)
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
by: Wu, Tianyi, et al.
Published: (2025)
by: Wu, Tianyi, et al.
Published: (2025)
Preserving Task-Relevant Information Under Linear Concept Removal
by: Holstege, Floris, et al.
Published: (2025)
by: Holstege, Floris, et al.
Published: (2025)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
AI-Assisted Human Evaluation of Machine Translation
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
by: Petty, Jackson, et al.
Published: (2025)
by: Petty, Jackson, et al.
Published: (2025)
IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
by: Maimon, Aviya, et al.
Published: (2025)
by: Maimon, Aviya, et al.
Published: (2025)
Probing for Arithmetic Errors in Language Models
by: Sun, Yucheng, et al.
Published: (2025)
by: Sun, Yucheng, et al.
Published: (2025)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
by: Do, Heejin, et al.
Published: (2026)
by: Do, Heejin, et al.
Published: (2026)
Measuring Scalar Constructs in Social Science with LLMs
by: Licht, Hauke, et al.
Published: (2025)
by: Licht, Hauke, et al.
Published: (2025)
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
by: Schäfer, Anton, et al.
Published: (2024)
by: Schäfer, Anton, et al.
Published: (2024)
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
by: Rassin, Royi, et al.
Published: (2023)
by: Rassin, Royi, et al.
Published: (2023)
Representation Surgery: Theory and Practice of Affine Steering
by: Singh, Shashwat, et al.
Published: (2024)
by: Singh, Shashwat, et al.
Published: (2024)
Grammar Control in Dialogue Response Generation for Language Learning Chatbots
by: Glandorf, Dominik, et al.
Published: (2025)
by: Glandorf, Dominik, et al.
Published: (2025)
Efficiently Computing Susceptibility to Context in Language Models
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
World Models for Math Story Problems
by: Opedal, Andreas, et al.
Published: (2023)
by: Opedal, Andreas, et al.
Published: (2023)
Description-Based Text Similarity
by: Ravfogel, Shauli, et al.
Published: (2023)
by: Ravfogel, Shauli, et al.
Published: (2023)
LEACE: Perfect linear concept erasure in closed form
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
Do Vision-Language Models Really Understand Visual Language?
by: Hou, Yifan, et al.
Published: (2024)
by: Hou, Yifan, et al.
Published: (2024)
Similar Items
-
Linear Adversarial Concept Erasure
by: Ravfogel, Shauli, et al.
Published: (2022) -
Kernelized Concept Erasure
by: Ravfogel, Shauli, et al.
Published: (2022) -
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
by: Ni, Jingwei, et al.
Published: (2025) -
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation
by: Ni, Jingwei, et al.
Published: (2024) -
Emergence of Linear Truth Encodings in Language Models
by: Ravfogel, Shauli, et al.
Published: (2025)