Gespeichert in:
| Hauptverfasser: | Shafran, Or, Ronen, Shaked, Fahn, Omri, Ravfogel, Shauli, Geiger, Atticus, Geva, Mor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.02464 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Constructing Interpretable Features from Compositional Neuron Groups
von: Shafran, Or, et al.
Veröffentlicht: (2025)
von: Shafran, Or, et al.
Veröffentlicht: (2025)
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
von: Hong, Yihuai, et al.
Veröffentlicht: (2024)
von: Hong, Yihuai, et al.
Veröffentlicht: (2024)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
von: Huang, Jing, et al.
Veröffentlicht: (2024)
von: Huang, Jing, et al.
Veröffentlicht: (2024)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
Gumbel Counterfactual Generation From Language Models
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2024)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2024)
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
von: Ben-Zaken, Elad, et al.
Veröffentlicht: (2021)
von: Ben-Zaken, Elad, et al.
Veröffentlicht: (2021)
Log-linear Guardedness and its Implications
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
Emergence of Linear Truth Encodings in Language Models
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2025)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2025)
Detecting (Un)answerability in Large Language Models with Linear Directions
von: Lavi, Maor Juliet, et al.
Veröffentlicht: (2025)
von: Lavi, Maor Juliet, et al.
Veröffentlicht: (2025)
Estimating Knowledge in Large Language Models Without Generating a Single Token
von: Gottesman, Daniela, et al.
Veröffentlicht: (2024)
von: Gottesman, Daniela, et al.
Veröffentlicht: (2024)
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
von: Cohen, Amir DN, et al.
Veröffentlicht: (2024)
von: Cohen, Amir DN, et al.
Veröffentlicht: (2024)
Geometric Factual Recall in Transformers
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2026)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2026)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
von: Yona, Itay, et al.
Veröffentlicht: (2026)
von: Yona, Itay, et al.
Veröffentlicht: (2026)
A Practical Method for Generating String Counterfactuals
von: Avitan, Matan, et al.
Veröffentlicht: (2024)
von: Avitan, Matan, et al.
Veröffentlicht: (2024)
RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context
von: Petty, Jackson, et al.
Veröffentlicht: (2025)
von: Petty, Jackson, et al.
Veröffentlicht: (2025)
Kernelized Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
State over Tokens: Characterizing the Role of Reasoning Tokens
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
Linear Adversarial Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
Activation Steering via Generative Causal Mediation
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
von: Ahrac, Sagi, et al.
Veröffentlicht: (2026)
von: Ahrac, Sagi, et al.
Veröffentlicht: (2026)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
von: Yona, Gal, et al.
Veröffentlicht: (2024)
von: Yona, Gal, et al.
Veröffentlicht: (2024)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
von: Ivgi, Maor, et al.
Veröffentlicht: (2024)
von: Ivgi, Maor, et al.
Veröffentlicht: (2024)
Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2025)
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2025)
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure
von: Fan, Yu, et al.
Veröffentlicht: (2025)
von: Fan, Yu, et al.
Veröffentlicht: (2025)
Pretrained LLMs Learn Multiple Types of Uncertainty
von: Cohen, Roi, et al.
Veröffentlicht: (2025)
von: Cohen, Roi, et al.
Veröffentlicht: (2025)
Inferring Functionality of Attention Heads from their Parameters
von: Elhelo, Amit, et al.
Veröffentlicht: (2024)
von: Elhelo, Amit, et al.
Veröffentlicht: (2024)
IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
von: Maimon, Aviya, et al.
Veröffentlicht: (2025)
von: Maimon, Aviya, et al.
Veröffentlicht: (2025)
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
von: Cohen, Ido, et al.
Veröffentlicht: (2024)
von: Cohen, Ido, et al.
Veröffentlicht: (2024)
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
von: Rassin, Royi, et al.
Veröffentlicht: (2023)
von: Rassin, Royi, et al.
Veröffentlicht: (2023)
Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
von: Yalon, Noam Steinmetz, et al.
Veröffentlicht: (2026)
von: Yalon, Noam Steinmetz, et al.
Veröffentlicht: (2026)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
Representation Surgery: Theory and Practice of Affine Steering
von: Singh, Shashwat, et al.
Veröffentlicht: (2024)
von: Singh, Shashwat, et al.
Veröffentlicht: (2024)
Description-Based Text Similarity
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2023)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2023)
Discrete Diffusion Models Exploit Asymmetry to Solve Lookahead Planning Tasks
von: Trainin, Itamar, et al.
Veröffentlicht: (2026)
von: Trainin, Itamar, et al.
Veröffentlicht: (2026)
Hallucinations Undermine Trust; Metacognition is a Way Forward
von: Yona, Gal, et al.
Veröffentlicht: (2026)
von: Yona, Gal, et al.
Veröffentlicht: (2026)
Eliciting Textual Descriptions from Representations of Continuous Prompts
von: Ramati, Dana, et al.
Veröffentlicht: (2024)
von: Ramati, Dana, et al.
Veröffentlicht: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
von: Yona, Gal, et al.
Veröffentlicht: (2024)
von: Yona, Gal, et al.
Veröffentlicht: (2024)
Do Large Language Models Latently Perform Multi-Hop Reasoning?
von: Yang, Sohee, et al.
Veröffentlicht: (2024)
von: Yang, Sohee, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Constructing Interpretable Features from Compositional Neuron Groups
von: Shafran, Or, et al.
Veröffentlicht: (2025) -
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025) -
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
von: Hong, Yihuai, et al.
Veröffentlicht: (2024) -
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
von: Huang, Jing, et al.
Veröffentlicht: (2024) -
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)