ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Zhengxuan, Manning, Christopher D., Potts, Christopher |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
di: Zhong, Zexuan, et al.
Pubblicazione: (2023)
di: Zhong, Zexuan, et al.
Pubblicazione: (2023)
Improved Representation Steering for Language Models
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
ReFT: Representation Finetuning for Language Models
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
di: Huang, Jing, et al.
Pubblicazione: (2024)
di: Huang, Jing, et al.
Pubblicazione: (2024)
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
di: Wu, Zhengxuan, et al.
Pubblicazione: (2023)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2023)
Exploring Compositional Generalization (in COGS/ReCOGS_pos) by Transformers using Restricted Access Sequence Processing (RASP)
di: Bruns, William
Pubblicazione: (2025)
di: Bruns, William
Pubblicazione: (2025)
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
di: Kallini, Julie, et al.
Pubblicazione: (2024)
di: Kallini, Julie, et al.
Pubblicazione: (2024)
HyperSteer: Activation Steering at Scale with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
Humans and transformer LMs: Abstraction drives language learning
di: Jian, Jasper, et al.
Pubblicazione: (2026)
di: Jian, Jasper, et al.
Pubblicazione: (2026)
Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP
di: Rozner, Josh, et al.
Pubblicazione: (2021)
di: Rozner, Josh, et al.
Pubblicazione: (2021)
SLFNet: Generating Semantic Logic Forms from Natural Language Using Semantic Probability Graphs
di: Wu, Hao, et al.
Pubblicazione: (2024)
di: Wu, Hao, et al.
Pubblicazione: (2024)
Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluations
di: Agarwal, Ananth, et al.
Pubblicazione: (2025)
di: Agarwal, Ananth, et al.
Pubblicazione: (2025)
Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language Models
di: Zhang, Yuji, et al.
Pubblicazione: (2024)
di: Zhang, Yuji, et al.
Pubblicazione: (2024)
A paradox of AI fluency
di: Potts, Christopher, et al.
Pubblicazione: (2026)
di: Potts, Christopher, et al.
Pubblicazione: (2026)
Invisible failures in human-AI interactions
di: Potts, Christopher, et al.
Pubblicazione: (2026)
di: Potts, Christopher, et al.
Pubblicazione: (2026)
Base Models Beat Aligned Models at Randomness and Creativity
di: West, Peter, et al.
Pubblicazione: (2025)
di: West, Peter, et al.
Pubblicazione: (2025)
Language models as tools for investigating the distinction between possible and impossible natural languages
di: Kallini, Julie, et al.
Pubblicazione: (2025)
di: Kallini, Julie, et al.
Pubblicazione: (2025)
PreFT: Prefill-only finetuning for efficient inference
di: Lanpouthakoun, Andrew, et al.
Pubblicazione: (2026)
di: Lanpouthakoun, Andrew, et al.
Pubblicazione: (2026)
Osiris: A Lightweight Open-Source Hallucination Detection System
di: Shan, Alex, et al.
Pubblicazione: (2025)
di: Shan, Alex, et al.
Pubblicazione: (2025)
Sneaking Syntax into Transformer Language Models with Tree Regularization
di: Nandi, Ananjan, et al.
Pubblicazione: (2024)
di: Nandi, Ananjan, et al.
Pubblicazione: (2024)
The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
di: Zhang, Yuji, et al.
Pubblicazione: (2025)
di: Zhang, Yuji, et al.
Pubblicazione: (2025)
Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models
di: Laitenberger, Alex, et al.
Pubblicazione: (2025)
di: Laitenberger, Alex, et al.
Pubblicazione: (2025)
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
di: Hase, Peter, et al.
Pubblicazione: (2026)
di: Hase, Peter, et al.
Pubblicazione: (2026)
A New Pair of GloVes
di: Carlson, Riley, et al.
Pubblicazione: (2025)
di: Carlson, Riley, et al.
Pubblicazione: (2025)
Drop Dropout on Single-Epoch Language Model Pretraining
di: Liu, Houjun, et al.
Pubblicazione: (2025)
di: Liu, Houjun, et al.
Pubblicazione: (2025)
Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis
di: Huang, Haoming, et al.
Pubblicazione: (2025)
di: Huang, Haoming, et al.
Pubblicazione: (2025)
Oolong: Investigating What Makes Transfer Learning Hard with Controlled Studies
di: Wu, Zhengxuan, et al.
Pubblicazione: (2022)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2022)
Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models
di: Ògúnrèmí, Tolúlopé, et al.
Pubblicazione: (2025)
di: Ògúnrèmí, Tolúlopé, et al.
Pubblicazione: (2025)
NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild
di: Murty, Shikhar, et al.
Pubblicazione: (2024)
di: Murty, Shikhar, et al.
Pubblicazione: (2024)
ActiShade: Activating Overshadowed Knowledge to Guide Multi-Hop Reasoning in Large Language Models
di: Ma, Huipeng, et al.
Pubblicazione: (2026)
di: Ma, Huipeng, et al.
Pubblicazione: (2026)
Do Language Models Use Their Depth Efficiently?
di: Csordás, Róbert, et al.
Pubblicazione: (2025)
di: Csordás, Róbert, et al.
Pubblicazione: (2025)
Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
di: Peter, Jan-Thorsten, et al.
Pubblicazione: (2025)
di: Peter, Jan-Thorsten, et al.
Pubblicazione: (2025)
Instruction Following without Instruction Tuning
di: Hewitt, John, et al.
Pubblicazione: (2024)
di: Hewitt, John, et al.
Pubblicazione: (2024)
Mapping the Increasing Use of LLMs in Scientific Papers
di: Liang, Weixin, et al.
Pubblicazione: (2024)
di: Liang, Weixin, et al.
Pubblicazione: (2024)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
di: Zhang, Jiayi, et al.
Pubblicazione: (2025)
di: Zhang, Jiayi, et al.
Pubblicazione: (2025)
Improving Pretraining Data Using Perplexity Correlations
di: Thrush, Tristan, et al.
Pubblicazione: (2024)
di: Thrush, Tristan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
di: Zhong, Zexuan, et al.
Pubblicazione: (2023) -
Improved Representation Steering for Language Models
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025) -
ReFT: Representation Finetuning for Language Models
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024) -
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
di: Huang, Jing, et al.
Pubblicazione: (2024) -
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
di: Wu, Zhengxuan, et al.
Pubblicazione: (2023)