Building Production-Ready Probes For Gemini
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kramár, János, Engels, Joshua, Wang, Zheng, Chughtai, Bilal, Shah, Rohin, Nanda, Neel, Conmy, Arthur |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
von: Kramár, János, et al.
Veröffentlicht: (2024)
von: Kramár, János, et al.
Veröffentlicht: (2024)
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Thought Anchors: Which LLM Reasoning Steps Matter?
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
von: Chughtai, Bilal, et al.
Veröffentlicht: (2024)
von: Chughtai, Bilal, et al.
Veröffentlicht: (2024)
Transformer Circuit Faithfulness Metrics are not Robust
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
Explorations of Self-Repair in Language Models
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
von: Minder, Julian, et al.
Veröffentlicht: (2025)
von: Minder, Julian, et al.
Veröffentlicht: (2025)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Understanding Reasoning in Thinking Language Models via Steering Vectors
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
Base Models Know How to Reason, Thinking Models Learn When
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
Real-Time Detection of Hallucinated Entities in Long-Form Generation
von: Obeso, Oscar, et al.
Veröffentlicht: (2025)
von: Obeso, Oscar, et al.
Veröffentlicht: (2025)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
von: Maar, Jim, et al.
Veröffentlicht: (2026)
von: Maar, Jim, et al.
Veröffentlicht: (2026)
How do LLMs Compute Verbal Confidence
von: Kumaran, Dharshan, et al.
Veröffentlicht: (2026)
von: Kumaran, Dharshan, et al.
Veröffentlicht: (2026)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
The Impact of Inference Acceleration on Bias of LLMs
von: Kirsten, Elisabeth, et al.
Veröffentlicht: (2024)
von: Kirsten, Elisabeth, et al.
Veröffentlicht: (2024)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
von: Casademunt, Helena, et al.
Veröffentlicht: (2026)
von: Casademunt, Helena, et al.
Veröffentlicht: (2026)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
von: Wang, Atticus, et al.
Veröffentlicht: (2025)
von: Wang, Atticus, et al.
Veröffentlicht: (2025)
Evaluating the Goal-Directedness of Large Language Models
von: Everitt, Tom, et al.
Veröffentlicht: (2025)
von: Everitt, Tom, et al.
Veröffentlicht: (2025)
Confidence Regulation Neurons in Language Models
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
von: Laine, Rudolf, et al.
Veröffentlicht: (2024)
von: Laine, Rudolf, et al.
Veröffentlicht: (2024)
Difficulties with Evaluating a Deception Detector for AIs
von: Smith, Lewis, et al.
Veröffentlicht: (2025)
von: Smith, Lewis, et al.
Veröffentlicht: (2025)
Universal Neurons in GPT2 Language Models
von: Gurnee, Wes, et al.
Veröffentlicht: (2024)
von: Gurnee, Wes, et al.
Veröffentlicht: (2024)
Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models
von: Patel, Laksh, et al.
Veröffentlicht: (2025)
von: Patel, Laksh, et al.
Veröffentlicht: (2025)
The Illusion of Readiness in Health AI
von: Gu, Yu, et al.
Veröffentlicht: (2025)
von: Gu, Yu, et al.
Veröffentlicht: (2025)
Zero-Overhead Introspection for Adaptive Test-Time Compute
von: Manvi, Rohin, et al.
Veröffentlicht: (2025)
von: Manvi, Rohin, et al.
Veröffentlicht: (2025)
Dense SAE Latents Are Features, Not Bugs
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
Large Language Models are Geographically Biased
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
Capabilities of Gemini Models in Medicine
von: Saab, Khaled, et al.
Veröffentlicht: (2024)
von: Saab, Khaled, et al.
Veröffentlicht: (2024)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
von: Kumar, Anshul
Veröffentlicht: (2026)
von: Kumar, Anshul
Veröffentlicht: (2026)
Advancing Multimodal Medical Capabilities of Gemini
von: Yang, Lin, et al.
Veröffentlicht: (2024)
von: Yang, Lin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
von: Lieberum, Tom, et al.
Veröffentlicht: (2024) -
AtP*: An efficient and scalable method for localizing LLM behaviour to components
von: Kramár, János, et al.
Veröffentlicht: (2024) -
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024) -
Thought Anchors: Which LLM Reasoning Steps Matter?
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025) -
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)