Enregistré dans:
| Auteurs principaux: | Chughtai, Bilal, Cooney, Alan, Nanda, Neel |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2402.07321 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Building Production-Ready Probes For Gemini
par: Kramár, János, et autres
Publié: (2026)
par: Kramár, János, et autres
Publié: (2026)
Difficulties with Evaluating a Deception Detector for AIs
par: Smith, Lewis, et autres
Publié: (2025)
par: Smith, Lewis, et autres
Publié: (2025)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
par: Yuan, Jiaqing, et autres
Publié: (2024)
par: Yuan, Jiaqing, et autres
Publié: (2024)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
par: Lv, Ang, et autres
Publié: (2024)
par: Lv, Ang, et autres
Publié: (2024)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
par: Smith, Matthew L., et autres
Publié: (2026)
par: Smith, Matthew L., et autres
Publié: (2026)
Transformer Circuit Faithfulness Metrics are not Robust
par: Miller, Joseph, et autres
Publié: (2024)
par: Miller, Joseph, et autres
Publié: (2024)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
par: Minder, Julian, et autres
Publié: (2025)
par: Minder, Julian, et autres
Publié: (2025)
Explorations of Self-Repair in Language Models
par: Rushing, Cody, et autres
Publié: (2024)
par: Rushing, Cody, et autres
Publié: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
par: Zhang, Fred, et autres
Publié: (2023)
par: Zhang, Fred, et autres
Publié: (2023)
Transcoders Find Interpretable LLM Feature Circuits
par: Dunefsky, Jacob, et autres
Publié: (2024)
par: Dunefsky, Jacob, et autres
Publié: (2024)
Understanding Factual Recall in Transformers via Associative Memories
par: Nichani, Eshaan, et autres
Publié: (2024)
par: Nichani, Eshaan, et autres
Publié: (2024)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
par: Wang, Qianli, et autres
Publié: (2025)
par: Wang, Qianli, et autres
Publié: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
par: Maar, Jim, et autres
Publié: (2026)
par: Maar, Jim, et autres
Publié: (2026)
The Impact of Inference Acceleration on Bias of LLMs
par: Kirsten, Elisabeth, et autres
Publié: (2024)
par: Kirsten, Elisabeth, et autres
Publié: (2024)
Locate-then-edit for Multi-hop Factual Recall under Knowledge Editing
par: Zhang, Zhuoran, et autres
Publié: (2024)
par: Zhang, Zhuoran, et autres
Publié: (2024)
Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts
par: Mujahid, Zain Muhammad, et autres
Publié: (2025)
par: Mujahid, Zain Muhammad, et autres
Publié: (2025)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
par: Karvonen, Adam, et autres
Publié: (2024)
par: Karvonen, Adam, et autres
Publié: (2024)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
par: Kramár, János, et autres
Publié: (2024)
par: Kramár, János, et autres
Publié: (2024)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
par: Casademunt, Helena, et autres
Publié: (2026)
par: Casademunt, Helena, et autres
Publié: (2026)
Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction
par: Wu, Qinyuan, et autres
Publié: (2024)
par: Wu, Qinyuan, et autres
Publié: (2024)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
par: Laine, Rudolf, et autres
Publié: (2024)
par: Laine, Rudolf, et autres
Publié: (2024)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
par: Ferrando, Javier, et autres
Publié: (2024)
par: Ferrando, Javier, et autres
Publié: (2024)
Exploring Precision and Recall to assess the quality and diversity of LLMs
par: Bronnec, Florian Le, et autres
Publié: (2024)
par: Bronnec, Florian Le, et autres
Publié: (2024)
Persuasion Tokens for Editing Factual Knowledge in LLMs
par: Youssef, Paul, et autres
Publié: (2026)
par: Youssef, Paul, et autres
Publié: (2026)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
par: Wang, Atticus, et autres
Publié: (2025)
par: Wang, Atticus, et autres
Publié: (2025)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
par: Macar, Uzay, et autres
Publié: (2025)
par: Macar, Uzay, et autres
Publié: (2025)
Thought Anchors: Which LLM Reasoning Steps Matter?
par: Bogdan, Paul C., et autres
Publié: (2025)
par: Bogdan, Paul C., et autres
Publié: (2025)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
par: Mahaut, Matéo, et autres
Publié: (2024)
par: Mahaut, Matéo, et autres
Publié: (2024)
FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Models
par: Wang, Yanling, et autres
Publié: (2024)
par: Wang, Yanling, et autres
Publié: (2024)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
par: Lei, Ge, et autres
Publié: (2025)
par: Lei, Ge, et autres
Publié: (2025)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
par: Rahman, Subhey Sadi, et autres
Publié: (2025)
par: Rahman, Subhey Sadi, et autres
Publié: (2025)
Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
par: Li, Junliang, et autres
Publié: (2025)
par: Li, Junliang, et autres
Publié: (2025)
Scaling sparse feature circuit finding for in-context learning
par: Kharlapenko, Dmitrii, et autres
Publié: (2025)
par: Kharlapenko, Dmitrii, et autres
Publié: (2025)
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees
par: Nie, Fan, et autres
Publié: (2024)
par: Nie, Fan, et autres
Publié: (2024)
Patent Language Model Pretraining with ModernBERT
par: Yousefiramandi, Amirhossein, et autres
Publié: (2025)
par: Yousefiramandi, Amirhossein, et autres
Publié: (2025)
Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks
par: Pink, Mathis, et autres
Publié: (2024)
par: Pink, Mathis, et autres
Publié: (2024)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
par: Casademunt, Helena, et autres
Publié: (2025)
par: Casademunt, Helena, et autres
Publié: (2025)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
par: Arcuschin, Iván, et autres
Publié: (2025)
par: Arcuschin, Iván, et autres
Publié: (2025)
Real-Time Detection of Hallucinated Entities in Long-Form Generation
par: Obeso, Oscar, et autres
Publié: (2025)
par: Obeso, Oscar, et autres
Publié: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
par: Sawczyn, Albert, et autres
Publié: (2025)
par: Sawczyn, Albert, et autres
Publié: (2025)
Documents similaires
-
Building Production-Ready Probes For Gemini
par: Kramár, János, et autres
Publié: (2026) -
Difficulties with Evaluating a Deception Detector for AIs
par: Smith, Lewis, et autres
Publié: (2025) -
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
par: Yuan, Jiaqing, et autres
Publié: (2024) -
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
par: Lv, Ang, et autres
Publié: (2024) -
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
par: Smith, Matthew L., et autres
Publié: (2026)