Truth-value judgment in language models: 'truth directions' are context sensitive
Fuente:
arXiv
Saved in:
| Main Authors: | Schouten, Stefan F., Bloem, Peter, Markov, Ilia, Vossen, Piek |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Constant in HATE: Analyzing Toxicity in Reddit across Topics and Languages
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
Grounding Toxicity in Real-World Events across Languages
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
Unknown Script: Impact of Script on Cross-Lingual Transfer
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024)
LLM Probing with Contrastive Eigenproblems: Improving Understanding and Applicability of CCS
by: Schouten, Stefan F., et al.
Published: (2025)
by: Schouten, Stefan F., et al.
Published: (2025)
Assessing and Refining ChatGPT's Performance in Identifying Targeting and Inappropriate Language: A Comparative Study
by: Baran, Barbarestani, et al.
Published: (2025)
by: Baran, Barbarestani, et al.
Published: (2025)
Understanding and Analyzing Inappropriately Targeting Language in Online Discourse: A Comparative Annotation Study
by: Barbarestani, Baran, et al.
Published: (2025)
by: Barbarestani, Baran, et al.
Published: (2025)
Knowledge acquisition for dialogue agents using reinforcement learning on graph representations
by: Santamaria, Selene Baez, et al.
Published: (2024)
by: Santamaria, Selene Baez, et al.
Published: (2024)
Extracting triples from dialogues for conversational social agents
by: Vossen, Piek, et al.
Published: (2024)
by: Vossen, Piek, et al.
Published: (2024)
Leveraging Open-Source Large Language Models for Native Language Identification
by: Ng, Yee Man, et al.
Published: (2024)
by: Ng, Yee Man, et al.
Published: (2024)
Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection
by: Brook, Joshua Wolfe, et al.
Published: (2025)
by: Brook, Joshua Wolfe, et al.
Published: (2025)
Do Differences in Values Influence Disagreements in Online Discussions?
by: van der Meer, Michiel, et al.
Published: (2023)
by: van der Meer, Michiel, et al.
Published: (2023)
An Empirical Analysis of Diversity in Argument Summarization
by: van der Meer, Michiel, et al.
Published: (2024)
by: van der Meer, Michiel, et al.
Published: (2024)
Berezinskii--Kosterlitz--Thouless transition in a context-sensitive random language model
by: Toji, Yuma, et al.
Published: (2024)
by: Toji, Yuma, et al.
Published: (2024)
Phase transition on a context-sensitive random language model with short range interactions
by: Toji, Yuma, et al.
Published: (2026)
by: Toji, Yuma, et al.
Published: (2026)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
by: Ilia, Evgenia, et al.
Published: (2024)
by: Ilia, Evgenia, et al.
Published: (2024)
Exploring the generalization of LLM truth directions on conversational formats
by: Ichmoukhamedov, Timour, et al.
Published: (2025)
by: Ichmoukhamedov, Timour, et al.
Published: (2025)
A Hybrid Intelligence Method for Argument Mining
by: van der Meer, Michiel, et al.
Published: (2024)
by: van der Meer, Michiel, et al.
Published: (2024)
Lessons in co-creation: the inconvenient truths of inclusive sign language technology development
by: De Meulder, Maartje, et al.
Published: (2024)
by: De Meulder, Maartje, et al.
Published: (2024)
Symbol tuning improves in-context learning in language models
by: Wei, Jerry, et al.
Published: (2023)
by: Wei, Jerry, et al.
Published: (2023)
The simulation of judgment in LLMs
by: Loru, Edoardo, et al.
Published: (2025)
by: Loru, Edoardo, et al.
Published: (2025)
Can large language models explore in-context?
by: Krishnamurthy, Akshay, et al.
Published: (2024)
by: Krishnamurthy, Akshay, et al.
Published: (2024)
Explanation sensitivity to the randomness of large language models: the case of journalistic text classification
by: Bogaert, Jeremie, et al.
Published: (2024)
by: Bogaert, Jeremie, et al.
Published: (2024)
Language models align with human judgments on key grammatical constructions
by: Hu, Jennifer, et al.
Published: (2024)
by: Hu, Jennifer, et al.
Published: (2024)
One ruler to measure them all: Benchmarking multilingual long-context language models
by: Kim, Yekyung, et al.
Published: (2025)
by: Kim, Yekyung, et al.
Published: (2025)
The advantages of context specific language models: the case of the Erasmian Language Model
by: Gonçalves, João, et al.
Published: (2024)
by: Gonçalves, João, et al.
Published: (2024)
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
Strong and weak alignment of large language models with human values
by: Khamassi, Mehdi, et al.
Published: (2024)
by: Khamassi, Mehdi, et al.
Published: (2024)
Cognitive models can reveal interpretable value trade-offs in language models
by: Murthy, Sonia K., et al.
Published: (2025)
by: Murthy, Sonia K., et al.
Published: (2025)
One Thousand and One Pairs: A "novel" challenge for long-context language models
by: Karpinska, Marzena, et al.
Published: (2024)
by: Karpinska, Marzena, et al.
Published: (2024)
Large language models reorganize representational geometry during in-context learning
by: Xiong, Hua-Dong, et al.
Published: (2026)
by: Xiong, Hua-Dong, et al.
Published: (2026)
InkubaLM: A small language model for low-resource African languages
by: Tonja, Atnafu Lambebo, et al.
Published: (2024)
by: Tonja, Atnafu Lambebo, et al.
Published: (2024)
Clustering Internet Memes Through Template Matching and Multi-Dimensional Similarity
by: Bloem, Tygo, et al.
Published: (2025)
by: Bloem, Tygo, et al.
Published: (2025)
A document processing pipeline for the construction of a dataset for topic modeling based on the judgments of the Italian Supreme Court
by: Marulli, Matteo, et al.
Published: (2025)
by: Marulli, Matteo, et al.
Published: (2025)
On the generalization of language models from in-context learning and finetuning: a controlled study
by: Lampinen, Andrew K., et al.
Published: (2025)
by: Lampinen, Andrew K., et al.
Published: (2025)
Truth Neurons
by: Li, Haohang, et al.
Published: (2025)
by: Li, Haohang, et al.
Published: (2025)
Kunz languages for numerical semigroups are context sensitive
by: Delgado, Manuel, et al.
Published: (2023)
by: Delgado, Manuel, et al.
Published: (2023)
The truth is no diaper: Human and AI-generated associations to emotional words
by: Vintar, Špela, et al.
Published: (2025)
by: Vintar, Špela, et al.
Published: (2025)
Inducing lexicons of in-group language with socio-temporal context
by: de Kock, Christine
Published: (2024)
by: de Kock, Christine
Published: (2024)
Enriching language models with graph-based context information to better understand textual data
by: Roethel, Albert, et al.
Published: (2023)
by: Roethel, Albert, et al.
Published: (2023)
Analyzing values about gendered language reform in LLMs' revisions
by: Watson, Jules, et al.
Published: (2025)
by: Watson, Jules, et al.
Published: (2025)
Similar Items
-
The Constant in HATE: Analyzing Toxicity in Reddit across Topics and Languages
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024) -
Grounding Toxicity in Real-World Events across Languages
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024) -
Unknown Script: Impact of Script on Cross-Lingual Transfer
by: Tufa, Wondimagegnhue Tsegaye, et al.
Published: (2024) -
LLM Probing with Contrastive Eigenproblems: Improving Understanding and Applicability of CCS
by: Schouten, Stefan F., et al.
Published: (2025) -
Assessing and Refining ChatGPT's Performance in Identifying Targeting and Inappropriate Language: A Comparative Study
by: Baran, Barbarestani, et al.
Published: (2025)