Estimating Knowledge in Large Language Models Without Generating a Single Token
Fuente:
arXiv
Guardado en:
| Autores principales: | Gottesman, Daniela, Geva, Mor |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
por: Cohen, Ido, et al.
Publicado: (2024)
por: Cohen, Ido, et al.
Publicado: (2024)
Eliciting Textual Descriptions from Representations of Continuous Prompts
por: Ramati, Dana, et al.
Publicado: (2024)
por: Ramati, Dana, et al.
Publicado: (2024)
Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries
por: Biran, Eden, et al.
Publicado: (2024)
por: Biran, Eden, et al.
Publicado: (2024)
LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
por: Gottesman, Daniela, et al.
Publicado: (2025)
por: Gottesman, Daniela, et al.
Publicado: (2025)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
por: Yona, Gal, et al.
Publicado: (2024)
por: Yona, Gal, et al.
Publicado: (2024)
How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?
por: Yang, Sohee, et al.
Publicado: (2025)
por: Yang, Sohee, et al.
Publicado: (2025)
Detecting (Un)answerability in Large Language Models with Linear Directions
por: Lavi, Maor Juliet, et al.
Publicado: (2025)
por: Lavi, Maor Juliet, et al.
Publicado: (2025)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
por: Yona, Gal, et al.
Publicado: (2024)
por: Yona, Gal, et al.
Publicado: (2024)
Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
por: Yalon, Noam Steinmetz, et al.
Publicado: (2026)
por: Yalon, Noam Steinmetz, et al.
Publicado: (2026)
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
por: Gur-Arieh, Yoav, et al.
Publicado: (2025)
por: Gur-Arieh, Yoav, et al.
Publicado: (2025)
Do Large Language Models Latently Perform Multi-Hop Reasoning?
por: Yang, Sohee, et al.
Publicado: (2024)
por: Yang, Sohee, et al.
Publicado: (2024)
Inferring Functionality of Attention Heads from their Parameters
por: Elhelo, Amit, et al.
Publicado: (2024)
por: Elhelo, Amit, et al.
Publicado: (2024)
Precise In-Parameter Concept Erasure in Large Language Models
por: Gur-Arieh, Yoav, et al.
Publicado: (2025)
por: Gur-Arieh, Yoav, et al.
Publicado: (2025)
Rethinking Selective Knowledge Distillation
por: Tavor, Almog, et al.
Publicado: (2026)
por: Tavor, Almog, et al.
Publicado: (2026)
Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?
por: Yang, Sohee, et al.
Publicado: (2024)
por: Yang, Sohee, et al.
Publicado: (2024)
Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models
por: Yona, Itay, et al.
Publicado: (2026)
por: Yona, Itay, et al.
Publicado: (2026)
Hallucinations Undermine Trust; Metacognition is a Way Forward
por: Yona, Gal, et al.
Publicado: (2026)
por: Yona, Gal, et al.
Publicado: (2026)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
por: Ivgi, Maor, et al.
Publicado: (2024)
por: Ivgi, Maor, et al.
Publicado: (2024)
Constructing Interpretable Features from Compositional Neuron Groups
por: Shafran, Or, et al.
Publicado: (2025)
por: Shafran, Or, et al.
Publicado: (2025)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
por: Katz, Shahar, et al.
Publicado: (2024)
por: Katz, Shahar, et al.
Publicado: (2024)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
por: Shafran, Or, et al.
Publicado: (2026)
por: Shafran, Or, et al.
Publicado: (2026)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
por: Huang, Jing, et al.
Publicado: (2024)
por: Huang, Jing, et al.
Publicado: (2024)
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
por: Hong, Yihuai, et al.
Publicado: (2024)
por: Hong, Yihuai, et al.
Publicado: (2024)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
por: Gur-Arieh, Yoav, et al.
Publicado: (2026)
por: Gur-Arieh, Yoav, et al.
Publicado: (2026)
Disentangling MLP Neuron Weights in Vocabulary Space
por: Avrahamy, Asaf, et al.
Publicado: (2026)
por: Avrahamy, Asaf, et al.
Publicado: (2026)
Preventing Rogue Agents Improves Multi-Agent Collaboration
por: Barbi, Ohav, et al.
Publicado: (2025)
por: Barbi, Ohav, et al.
Publicado: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
por: Ahrac, Sagi, et al.
Publicado: (2026)
por: Ahrac, Sagi, et al.
Publicado: (2026)
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
por: Gekhman, Zorik, et al.
Publicado: (2026)
por: Gekhman, Zorik, et al.
Publicado: (2026)
The Hidden Space of Transformer Language Adapters
por: Alabi, Jesujoba O., et al.
Publicado: (2024)
por: Alabi, Jesujoba O., et al.
Publicado: (2024)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
por: Ghandeharioun, Asma, et al.
Publicado: (2024)
por: Ghandeharioun, Asma, et al.
Publicado: (2024)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
por: Din, Alexander Yom, et al.
Publicado: (2023)
por: Din, Alexander Yom, et al.
Publicado: (2023)
Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data
por: Xia, Han, et al.
Publicado: (2024)
por: Xia, Han, et al.
Publicado: (2024)
From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP
por: Mosbach, Marius, et al.
Publicado: (2024)
por: Mosbach, Marius, et al.
Publicado: (2024)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
por: Gur-Arieh, Yoav, et al.
Publicado: (2025)
por: Gur-Arieh, Yoav, et al.
Publicado: (2025)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
por: Parmar, Mihir, et al.
Publicado: (2022)
por: Parmar, Mihir, et al.
Publicado: (2022)
Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language Models
por: Wang, Huazheng, et al.
Publicado: (2025)
por: Wang, Huazheng, et al.
Publicado: (2025)
ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training
por: Han, Feijiang, et al.
Publicado: (2025)
por: Han, Feijiang, et al.
Publicado: (2025)
Problematic Tokens: Tokenizer Bias in Large Language Models
por: Yang, Jin, et al.
Publicado: (2024)
por: Yang, Jin, et al.
Publicado: (2024)
TokenSHAP: Interpreting Large Language Models with Monte Carlo Shapley Value Estimation
por: Goldshmidt, Roni, et al.
Publicado: (2024)
por: Goldshmidt, Roni, et al.
Publicado: (2024)
Rethinking Tokenization: Crafting Better Tokenizers for Large Language Models
por: Yang, Jinbiao
Publicado: (2024)
por: Yang, Jinbiao
Publicado: (2024)
Ejemplares similares
-
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
por: Cohen, Ido, et al.
Publicado: (2024) -
Eliciting Textual Descriptions from Representations of Continuous Prompts
por: Ramati, Dana, et al.
Publicado: (2024) -
Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries
por: Biran, Eden, et al.
Publicado: (2024) -
LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations
por: Gottesman, Daniela, et al.
Publicado: (2025) -
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
por: Yona, Gal, et al.
Publicado: (2024)