Confidence Regulation Neurons in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Stolfo, Alessandro, Wu, Ben, Gurnee, Wes, Belinkov, Yonatan, Song, Xingyi, Sachan, Mrinmaya, Nanda, Neel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Universal Neurons in GPT2 Language Models
von: Gurnee, Wes, et al.
Veröffentlicht: (2024)
von: Gurnee, Wes, et al.
Veröffentlicht: (2024)
Language Models Represent Space and Time
von: Gurnee, Wes, et al.
Veröffentlicht: (2023)
von: Gurnee, Wes, et al.
Veröffentlicht: (2023)
Probing for Arithmetic Errors in Language Models
von: Sun, Yucheng, et al.
Veröffentlicht: (2025)
von: Sun, Yucheng, et al.
Veröffentlicht: (2025)
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
Dense SAE Latents Are Features, Not Bugs
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
Explorations of Self-Repair in Language Models
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
Improving Large Language Model Safety with Contrastive Representation Learning
von: Simko, Samuel, et al.
Veröffentlicht: (2025)
von: Simko, Samuel, et al.
Veröffentlicht: (2025)
The Remarkable Robustness of LLMs: Stages of Inference?
von: Lad, Vedang, et al.
Veröffentlicht: (2024)
von: Lad, Vedang, et al.
Veröffentlicht: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
Towards Aligning Language Models with Textual Feedback
von: Lloret, Saüc Abadal, et al.
Veröffentlicht: (2024)
von: Lloret, Saüc Abadal, et al.
Veröffentlicht: (2024)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
von: Daheim, Nico, et al.
Veröffentlicht: (2024)
von: Daheim, Nico, et al.
Veröffentlicht: (2024)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
von: Do, Heejin, et al.
Veröffentlicht: (2026)
von: Do, Heejin, et al.
Veröffentlicht: (2026)
SAEs Are Good for Steering -- If You Select the Right Features
von: Arad, Dana, et al.
Veröffentlicht: (2025)
von: Arad, Dana, et al.
Veröffentlicht: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
von: Itzhak, Itay, et al.
Veröffentlicht: (2025)
von: Itzhak, Itay, et al.
Veröffentlicht: (2025)
Improving Instruction-Following in Language Models through Activation Steering
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
On the Emergence of Induction Heads for In-Context Learning
von: Musat, Tiberiu, et al.
Veröffentlicht: (2025)
von: Musat, Tiberiu, et al.
Veröffentlicht: (2025)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
von: Marks, Samuel, et al.
Veröffentlicht: (2024)
von: Marks, Samuel, et al.
Veröffentlicht: (2024)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
von: Itzhak, Itay, et al.
Veröffentlicht: (2026)
von: Itzhak, Itay, et al.
Veröffentlicht: (2026)
Can Large Language Models Infer Causation from Correlation?
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
von: Macina, Jakub, et al.
Veröffentlicht: (2025)
von: Macina, Jakub, et al.
Veröffentlicht: (2025)
Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented Generation
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
Autoformalizing Natural Language to First-Order Logic: A Case Study in Logical Fallacy Detection
von: Lalwani, Abhinav, et al.
Veröffentlicht: (2024)
von: Lalwani, Abhinav, et al.
Veröffentlicht: (2024)
Learning to Reason Efficiently with A* Post-Training
von: Opedal, Andreas, et al.
Veröffentlicht: (2026)
von: Opedal, Andreas, et al.
Veröffentlicht: (2026)
Are Language Models Efficient Reasoners? A Perspective from Logic Programming
von: Opedal, Andreas, et al.
Veröffentlicht: (2025)
von: Opedal, Andreas, et al.
Veröffentlicht: (2025)
CLadder: Assessing Causal Reasoning in Language Models
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
Thought Anchors: Which LLM Reasoning Steps Matter?
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
Implicit Personalization in Language Models: A Systematic Study
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
Groundedness in Retrieval-augmented Long-form Generation: An Empirical Study
von: Stolfo, Alessandro
Veröffentlicht: (2024)
von: Stolfo, Alessandro
Veröffentlicht: (2024)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
von: Maar, Jim, et al.
Veröffentlicht: (2026)
von: Maar, Jim, et al.
Veröffentlicht: (2026)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
von: Minder, Julian, et al.
Veröffentlicht: (2025)
von: Minder, Julian, et al.
Veröffentlicht: (2025)
Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
von: Itzhak, Itay, et al.
Veröffentlicht: (2023)
von: Itzhak, Itay, et al.
Veröffentlicht: (2023)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
von: Kumar, Abhishek, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Universal Neurons in GPT2 Language Models
von: Gurnee, Wes, et al.
Veröffentlicht: (2024) -
Language Models Represent Space and Time
von: Gurnee, Wes, et al.
Veröffentlicht: (2023) -
Probing for Arithmetic Errors in Language Models
von: Sun, Yucheng, et al.
Veröffentlicht: (2025) -
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024) -
Dense SAE Latents Are Features, Not Bugs
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)