Identifying Linear Relational Concepts in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Chanin, David, Hunter, Anthony, Camburu, Oana-Maria |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
di: Siegel, Noah Y., et al.
Pubblicazione: (2024)
di: Siegel, Noah Y., et al.
Pubblicazione: (2024)
SPARSEFIT: Few-shot Prompting with Sparse Fine-tuning for Jointly Generating Predictions and Natural Language Explanations
di: Solano, Jesus, et al.
Pubblicazione: (2023)
di: Solano, Jesus, et al.
Pubblicazione: (2023)
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
di: Arcuschin, Iván, et al.
Pubblicazione: (2026)
di: Arcuschin, Iván, et al.
Pubblicazione: (2026)
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
di: Siegel, Noah Y., et al.
Pubblicazione: (2025)
di: Siegel, Noah Y., et al.
Pubblicazione: (2025)
Identification of Entailment and Contradiction Relations between Natural Language Sentences: A Neurosymbolic Approach
di: Feng, Xuyao, et al.
Pubblicazione: (2024)
di: Feng, Xuyao, et al.
Pubblicazione: (2024)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
AnnoCaseLaw: A Richly-Annotated Dataset For Benchmarking Explainable Legal Judgment Prediction
di: Sesodia, Magnus, et al.
Pubblicazione: (2025)
di: Sesodia, Magnus, et al.
Pubblicazione: (2025)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
di: Joshi, Shruti, et al.
Pubblicazione: (2025)
di: Joshi, Shruti, et al.
Pubblicazione: (2025)
Model Tuning or Prompt Tuning? A Study of Large Language Models for Clinical Concept and Relation Extraction
di: Peng, Cheng, et al.
Pubblicazione: (2023)
di: Peng, Cheng, et al.
Pubblicazione: (2023)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
Biomedical Relation Extraction via Adaptive Document-Relation Cross-Mapping and Concept Unique Identifier
di: Shang, Yufei, et al.
Pubblicazione: (2025)
di: Shang, Yufei, et al.
Pubblicazione: (2025)
Large Language Model for Patent Concept Generation
di: Ren, Runtao, et al.
Pubblicazione: (2024)
di: Ren, Runtao, et al.
Pubblicazione: (2024)
The Structure of Relation Decoding Linear Operators in Large Language Models
di: Christ, Miranda Anna, et al.
Pubblicazione: (2025)
di: Christ, Miranda Anna, et al.
Pubblicazione: (2025)
Emotion Concepts and their Function in a Large Language Model
di: Sofroniew, Nicholas, et al.
Pubblicazione: (2026)
di: Sofroniew, Nicholas, et al.
Pubblicazione: (2026)
ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language Models
di: Li, Haoxuan, et al.
Pubblicazione: (2025)
di: Li, Haoxuan, et al.
Pubblicazione: (2025)
Identifying Knowledge Editing Types in Large Language Models
di: Li, Xiaopeng, et al.
Pubblicazione: (2024)
di: Li, Xiaopeng, et al.
Pubblicazione: (2024)
LLMLagBench: Identifying Temporal Training Boundaries in Large Language Models
di: Pęzik, Piotr, et al.
Pubblicazione: (2025)
di: Pęzik, Piotr, et al.
Pubblicazione: (2025)
Meta-Judging with Large Language Models: Concepts, Methods, and Challenges
di: Silva, Hugo, et al.
Pubblicazione: (2026)
di: Silva, Hugo, et al.
Pubblicazione: (2026)
CoLLEGe: Concept Embedding Generation for Large Language Models
di: Teehan, Ryan, et al.
Pubblicazione: (2024)
di: Teehan, Ryan, et al.
Pubblicazione: (2024)
Making Implicit Premises Explicit in Logical Understanding of Enthymemes
di: Feng, Xuyao, et al.
Pubblicazione: (2026)
di: Feng, Xuyao, et al.
Pubblicazione: (2026)
Identifying and Extracting Rare Disease Phenotypes with Large Language Models
di: Shyr, Cathy, et al.
Pubblicazione: (2023)
di: Shyr, Cathy, et al.
Pubblicazione: (2023)
Identifying Multiple Personalities in Large Language Models with External Evaluation
di: Song, Xiaoyang, et al.
Pubblicazione: (2024)
di: Song, Xiaoyang, et al.
Pubblicazione: (2024)
Evaluating the Efficacy of Large Language Models in Identifying Phishing Attempts
di: Patel, Het, et al.
Pubblicazione: (2024)
di: Patel, Het, et al.
Pubblicazione: (2024)
ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
di: Wu, Yanan, et al.
Pubblicazione: (2024)
di: Wu, Yanan, et al.
Pubblicazione: (2024)
Understanding Enthymemes in Argument Maps: Bridging Argument Mining and Logic-based Argumentation
di: Ben-Naim, Jonathan, et al.
Pubblicazione: (2024)
di: Ben-Naim, Jonathan, et al.
Pubblicazione: (2024)
Linearly-Interpretable Concept Embedding Models for Text Analysis
di: De Santis, Francesco, et al.
Pubblicazione: (2024)
di: De Santis, Francesco, et al.
Pubblicazione: (2024)
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2024)
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2024)
Uncovering Implicit Bias in Large Language Models with Concept Learning Dataset
di: Wang, Leroy Z.
Pubblicazione: (2025)
di: Wang, Leroy Z.
Pubblicazione: (2025)
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2024)
di: Chanin, David, et al.
Pubblicazione: (2024)
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
di: Huang, Youcheng, et al.
Pubblicazione: (2025)
di: Huang, Youcheng, et al.
Pubblicazione: (2025)
Large Language Models to Identify Social Determinants of Health in Electronic Health Records
di: Guevara, Marco, et al.
Pubblicazione: (2023)
di: Guevara, Marco, et al.
Pubblicazione: (2023)
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
di: Sankaranarayanan, Aruna, et al.
Pubblicazione: (2025)
di: Sankaranarayanan, Aruna, et al.
Pubblicazione: (2025)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
di: Röttger, Paul, et al.
Pubblicazione: (2023)
di: Röttger, Paul, et al.
Pubblicazione: (2023)
Navigating the Concept Space of Language Models
di: Marcílio-Jr, Wilson E., et al.
Pubblicazione: (2026)
di: Marcílio-Jr, Wilson E., et al.
Pubblicazione: (2026)
The Geometry of Categorical and Hierarchical Concepts in Large Language Models
di: Park, Kiho, et al.
Pubblicazione: (2024)
di: Park, Kiho, et al.
Pubblicazione: (2024)
Empirical Analysis of Dialogue Relation Extraction with Large Language Models
di: Li, Guozheng, et al.
Pubblicazione: (2024)
di: Li, Guozheng, et al.
Pubblicazione: (2024)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
di: Marconato, Emanuele, et al.
Pubblicazione: (2024)
di: Marconato, Emanuele, et al.
Pubblicazione: (2024)
Mimir: Large-scale Multilingual Concept Modeling
di: Musacchio, Elio, et al.
Pubblicazione: (2026)
di: Musacchio, Elio, et al.
Pubblicazione: (2026)
From Knowledge to Treatment: Large Language Model Assisted Biomedical Concept Representation for Drug Repurposing
di: Xiang, Chengrui, et al.
Pubblicazione: (2025)
di: Xiang, Chengrui, et al.
Pubblicazione: (2025)
Promoting Equality in Large Language Models: Identifying and Mitigating the Implicit Bias based on Bayesian Theory
di: Deng, Yongxin, et al.
Pubblicazione: (2024)
di: Deng, Yongxin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
di: Siegel, Noah Y., et al.
Pubblicazione: (2024) -
SPARSEFIT: Few-shot Prompting with Sparse Fine-tuning for Jointly Generating Predictions and Natural Language Explanations
di: Solano, Jesus, et al.
Pubblicazione: (2023) -
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
di: Arcuschin, Iván, et al.
Pubblicazione: (2026) -
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
di: Siegel, Noah Y., et al.
Pubblicazione: (2025) -
Identification of Entailment and Contradiction Relations between Natural Language Sentences: A Neurosymbolic Approach
di: Feng, Xuyao, et al.
Pubblicazione: (2024)