Guardado en:
| Autores principales: | Chughtai, Bilal, Cooney, Alan, Nanda, Neel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2402.07321 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Building Production-Ready Probes For Gemini
por: Kramár, János, et al.
Publicado: (2026)
por: Kramár, János, et al.
Publicado: (2026)
Difficulties with Evaluating a Deception Detector for AIs
por: Smith, Lewis, et al.
Publicado: (2025)
por: Smith, Lewis, et al.
Publicado: (2025)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
por: Yuan, Jiaqing, et al.
Publicado: (2024)
por: Yuan, Jiaqing, et al.
Publicado: (2024)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
por: Lv, Ang, et al.
Publicado: (2024)
por: Lv, Ang, et al.
Publicado: (2024)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
por: Smith, Matthew L., et al.
Publicado: (2026)
por: Smith, Matthew L., et al.
Publicado: (2026)
Transformer Circuit Faithfulness Metrics are not Robust
por: Miller, Joseph, et al.
Publicado: (2024)
por: Miller, Joseph, et al.
Publicado: (2024)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
por: Minder, Julian, et al.
Publicado: (2025)
por: Minder, Julian, et al.
Publicado: (2025)
Explorations of Self-Repair in Language Models
por: Rushing, Cody, et al.
Publicado: (2024)
por: Rushing, Cody, et al.
Publicado: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
por: Zhang, Fred, et al.
Publicado: (2023)
por: Zhang, Fred, et al.
Publicado: (2023)
Transcoders Find Interpretable LLM Feature Circuits
por: Dunefsky, Jacob, et al.
Publicado: (2024)
por: Dunefsky, Jacob, et al.
Publicado: (2024)
Understanding Factual Recall in Transformers via Associative Memories
por: Nichani, Eshaan, et al.
Publicado: (2024)
por: Nichani, Eshaan, et al.
Publicado: (2024)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
por: Wang, Qianli, et al.
Publicado: (2025)
por: Wang, Qianli, et al.
Publicado: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
por: Maar, Jim, et al.
Publicado: (2026)
por: Maar, Jim, et al.
Publicado: (2026)
The Impact of Inference Acceleration on Bias of LLMs
por: Kirsten, Elisabeth, et al.
Publicado: (2024)
por: Kirsten, Elisabeth, et al.
Publicado: (2024)
Locate-then-edit for Multi-hop Factual Recall under Knowledge Editing
por: Zhang, Zhuoran, et al.
Publicado: (2024)
por: Zhang, Zhuoran, et al.
Publicado: (2024)
Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts
por: Mujahid, Zain Muhammad, et al.
Publicado: (2025)
por: Mujahid, Zain Muhammad, et al.
Publicado: (2025)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
por: Karvonen, Adam, et al.
Publicado: (2024)
por: Karvonen, Adam, et al.
Publicado: (2024)
AtP*: An efficient and scalable method for localizing LLM behaviour to components
por: Kramár, János, et al.
Publicado: (2024)
por: Kramár, János, et al.
Publicado: (2024)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
por: Casademunt, Helena, et al.
Publicado: (2026)
por: Casademunt, Helena, et al.
Publicado: (2026)
Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction
por: Wu, Qinyuan, et al.
Publicado: (2024)
por: Wu, Qinyuan, et al.
Publicado: (2024)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
por: Laine, Rudolf, et al.
Publicado: (2024)
por: Laine, Rudolf, et al.
Publicado: (2024)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
por: Ferrando, Javier, et al.
Publicado: (2024)
por: Ferrando, Javier, et al.
Publicado: (2024)
Exploring Precision and Recall to assess the quality and diversity of LLMs
por: Bronnec, Florian Le, et al.
Publicado: (2024)
por: Bronnec, Florian Le, et al.
Publicado: (2024)
Persuasion Tokens for Editing Factual Knowledge in LLMs
por: Youssef, Paul, et al.
Publicado: (2026)
por: Youssef, Paul, et al.
Publicado: (2026)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
por: Wang, Atticus, et al.
Publicado: (2025)
por: Wang, Atticus, et al.
Publicado: (2025)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
por: Macar, Uzay, et al.
Publicado: (2025)
por: Macar, Uzay, et al.
Publicado: (2025)
Thought Anchors: Which LLM Reasoning Steps Matter?
por: Bogdan, Paul C., et al.
Publicado: (2025)
por: Bogdan, Paul C., et al.
Publicado: (2025)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
por: Mahaut, Matéo, et al.
Publicado: (2024)
por: Mahaut, Matéo, et al.
Publicado: (2024)
FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Models
por: Wang, Yanling, et al.
Publicado: (2024)
por: Wang, Yanling, et al.
Publicado: (2024)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
por: Lei, Ge, et al.
Publicado: (2025)
por: Lei, Ge, et al.
Publicado: (2025)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
por: Rahman, Subhey Sadi, et al.
Publicado: (2025)
por: Rahman, Subhey Sadi, et al.
Publicado: (2025)
Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
por: Li, Junliang, et al.
Publicado: (2025)
por: Li, Junliang, et al.
Publicado: (2025)
Scaling sparse feature circuit finding for in-context learning
por: Kharlapenko, Dmitrii, et al.
Publicado: (2025)
por: Kharlapenko, Dmitrii, et al.
Publicado: (2025)
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees
por: Nie, Fan, et al.
Publicado: (2024)
por: Nie, Fan, et al.
Publicado: (2024)
Patent Language Model Pretraining with ModernBERT
por: Yousefiramandi, Amirhossein, et al.
Publicado: (2025)
por: Yousefiramandi, Amirhossein, et al.
Publicado: (2025)
Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks
por: Pink, Mathis, et al.
Publicado: (2024)
por: Pink, Mathis, et al.
Publicado: (2024)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
por: Casademunt, Helena, et al.
Publicado: (2025)
por: Casademunt, Helena, et al.
Publicado: (2025)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
por: Arcuschin, Iván, et al.
Publicado: (2025)
por: Arcuschin, Iván, et al.
Publicado: (2025)
Real-Time Detection of Hallucinated Entities in Long-Form Generation
por: Obeso, Oscar, et al.
Publicado: (2025)
por: Obeso, Oscar, et al.
Publicado: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
por: Sawczyn, Albert, et al.
Publicado: (2025)
por: Sawczyn, Albert, et al.
Publicado: (2025)
Ejemplares similares
-
Building Production-Ready Probes For Gemini
por: Kramár, János, et al.
Publicado: (2026) -
Difficulties with Evaluating a Deception Detector for AIs
por: Smith, Lewis, et al.
Publicado: (2025) -
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
por: Yuan, Jiaqing, et al.
Publicado: (2024) -
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
por: Lv, Ang, et al.
Publicado: (2024) -
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
por: Smith, Matthew L., et al.
Publicado: (2026)