Measuring How LLMs Internalize Human Psychological Concepts: A preliminary analysis
Fuente:
arXiv
Salvato in:
| Autori principali: | Hamada, Hiro Taiyo, Fujisawa, Ippei, Kawakita, Genji, Yamada, Yuki |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Psychological Concept Neurons: Can Neural Control Bias Probing and Shift Generation in LLMs?
di: Harada, Yuto, et al.
Pubblicazione: (2026)
di: Harada, Yuto, et al.
Pubblicazione: (2026)
How Well Do LLMs Predict Human Behavior? A Measure of their Pretrained Knowledge
di: Gao, Wayne, et al.
Pubblicazione: (2026)
di: Gao, Wayne, et al.
Pubblicazione: (2026)
Scalable Valuation of Human Feedback through Provably Robust Model Alignment
di: Fujisawa, Masahiro, et al.
Pubblicazione: (2025)
di: Fujisawa, Masahiro, et al.
Pubblicazione: (2025)
Signs Beat Floats: Low-Rank Double-Binary Adaptation for On-Device Fine-Tuning
di: Fujisawa, Yoshihiko, et al.
Pubblicazione: (2026)
di: Fujisawa, Yoshihiko, et al.
Pubblicazione: (2026)
Exploring the Frontiers of LLMs in Psychological Applications: A Comprehensive Review
di: Ke, Luoma, et al.
Pubblicazione: (2024)
di: Ke, Luoma, et al.
Pubblicazione: (2024)
ProcBench: Benchmark for Multi-Step Reasoning and Following Procedure
di: Fujisawa, Ippei, et al.
Pubblicazione: (2024)
di: Fujisawa, Ippei, et al.
Pubblicazione: (2024)
On LLMs' Internal Representation of Code Correctness
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
Artificial Neural Nets and the Representation of Human Concepts
di: Freiesleben, Timo
Pubblicazione: (2023)
di: Freiesleben, Timo
Pubblicazione: (2023)
Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection
di: Yamada, Daisuke, et al.
Pubblicazione: (2025)
di: Yamada, Daisuke, et al.
Pubblicazione: (2025)
How Likely Do LLMs with CoT Mimic Human Reasoning?
di: Bao, Guangsheng, et al.
Pubblicazione: (2024)
di: Bao, Guangsheng, et al.
Pubblicazione: (2024)
Explainable Human Activity Recognition: A Unified Review of Concepts and Mechanisms
di: Kundu, Mainak, et al.
Pubblicazione: (2026)
di: Kundu, Mainak, et al.
Pubblicazione: (2026)
Measuring Leakage in Concept-Based Methods: An Information Theoretic Approach
di: Makonnen, Mikael, et al.
Pubblicazione: (2025)
di: Makonnen, Mikael, et al.
Pubblicazione: (2025)
More Than Bits: Multi-Envelope Double Binary Factorization for Extreme Quantization
di: Ichikawa, Yuma, et al.
Pubblicazione: (2025)
di: Ichikawa, Yuma, et al.
Pubblicazione: (2025)
Reinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMs
di: Zhang, Honglin, et al.
Pubblicazione: (2025)
di: Zhang, Honglin, et al.
Pubblicazione: (2025)
MindCraft: How Concept Trees Take Shape In Deep Models
di: Tian, Bowei, et al.
Pubblicazione: (2025)
di: Tian, Bowei, et al.
Pubblicazione: (2025)
Enhancing Output Diversity Improves Conjugate Gradient-based Adversarial Attacks
di: Yamamura, Keiichiro, et al.
Pubblicazione: (2024)
di: Yamamura, Keiichiro, et al.
Pubblicazione: (2024)
Norm Anchors Make Model Edits Last
di: Liu, Mingda, et al.
Pubblicazione: (2026)
di: Liu, Mingda, et al.
Pubblicazione: (2026)
How Good Are LLMs at Processing Tool Outputs?
di: Kate, Kiran, et al.
Pubblicazione: (2025)
di: Kate, Kiran, et al.
Pubblicazione: (2025)
Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
An Analysis of Concept Bottleneck Models: Measuring, Understanding, and Mitigating the Impact of Noisy Annotations
di: Park, Seonghwan, et al.
Pubblicazione: (2025)
di: Park, Seonghwan, et al.
Pubblicazione: (2025)
Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
di: Zhao, Haiyan, et al.
Pubblicazione: (2024)
di: Zhao, Haiyan, et al.
Pubblicazione: (2024)
Learning Displacement-Robust Representations for Landslide Early Warning under Rainfall Forecast Uncertainty
di: Ozeki, Ren, et al.
Pubblicazione: (2026)
di: Ozeki, Ren, et al.
Pubblicazione: (2026)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
di: Kirchhof, Michael, et al.
Pubblicazione: (2025)
di: Kirchhof, Michael, et al.
Pubblicazione: (2025)
NEAT: Concept driven Neuron Attribution in LLMs
di: Kavuri, Vivek Hruday, et al.
Pubblicazione: (2025)
di: Kavuri, Vivek Hruday, et al.
Pubblicazione: (2025)
Coordinate Matrix Machine: A Human-level Concept Learning to Classify Very Similar Documents
di: Sadri, Amin, et al.
Pubblicazione: (2025)
di: Sadri, Amin, et al.
Pubblicazione: (2025)
How to Measure the Intelligence of Large Language Models?
di: Körber, Nils, et al.
Pubblicazione: (2024)
di: Körber, Nils, et al.
Pubblicazione: (2024)
From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
di: Achtibat, Reduan, et al.
Pubblicazione: (2022)
di: Achtibat, Reduan, et al.
Pubblicazione: (2022)
Graph Neural Networks for Gut Microbiome Metaomic data: A preliminary work
di: Irwin, Christopher, et al.
Pubblicazione: (2024)
di: Irwin, Christopher, et al.
Pubblicazione: (2024)
Zero-Shot LLMs in Human-in-the-Loop RL: Replacing Human Feedback for Reward Shaping
di: Nazir, Mohammad Saif, et al.
Pubblicazione: (2025)
di: Nazir, Mohammad Saif, et al.
Pubblicazione: (2025)
Exploring Concept Depth: How Large Language Models Acquire Knowledge and Concept at Different Layers?
di: Jin, Mingyu, et al.
Pubblicazione: (2024)
di: Jin, Mingyu, et al.
Pubblicazione: (2024)
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
di: Braun, Tobias, et al.
Pubblicazione: (2025)
di: Braun, Tobias, et al.
Pubblicazione: (2025)
How Random is Random? Evaluating the Randomness and Humaness of LLMs' Coin Flips
di: Van Koevering, Katherine, et al.
Pubblicazione: (2024)
di: Van Koevering, Katherine, et al.
Pubblicazione: (2024)
On the Internal Representations of Graph Metanetworks
di: Yeom, Taesun, et al.
Pubblicazione: (2025)
di: Yeom, Taesun, et al.
Pubblicazione: (2025)
A Shared Valence Axis Across Modern LLMs and Human EEG: The Saturation Regularity
di: Radwan, Yousef A., et al.
Pubblicazione: (2026)
di: Radwan, Yousef A., et al.
Pubblicazione: (2026)
Subgoal-based Reward Shaping to Improve Efficiency in Reinforcement Learning
di: Okudo, Takato, et al.
Pubblicazione: (2021)
di: Okudo, Takato, et al.
Pubblicazione: (2021)
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
di: Soni, Nikita, et al.
Pubblicazione: (2025)
di: Soni, Nikita, et al.
Pubblicazione: (2025)
Reinforcement-learning robotic sailboats: simulator and preliminary results
di: Vasconcellos, Eduardo Charles, et al.
Pubblicazione: (2024)
di: Vasconcellos, Eduardo Charles, et al.
Pubblicazione: (2024)
Addressing Concept Mislabeling in Concept Bottleneck Models Through Preference Optimization
di: Penaloza, Emiliano, et al.
Pubblicazione: (2025)
di: Penaloza, Emiliano, et al.
Pubblicazione: (2025)
ConceptTracer: Interactive Analysis of Concept Saliency and Selectivity in Neural Representations
di: Knauer, Ricardo, et al.
Pubblicazione: (2026)
di: Knauer, Ricardo, et al.
Pubblicazione: (2026)
How Susceptible are LLMs to Influence in Prompts?
di: Anagnostidis, Sotiris, et al.
Pubblicazione: (2024)
di: Anagnostidis, Sotiris, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Psychological Concept Neurons: Can Neural Control Bias Probing and Shift Generation in LLMs?
di: Harada, Yuto, et al.
Pubblicazione: (2026) -
How Well Do LLMs Predict Human Behavior? A Measure of their Pretrained Knowledge
di: Gao, Wayne, et al.
Pubblicazione: (2026) -
Scalable Valuation of Human Feedback through Provably Robust Model Alignment
di: Fujisawa, Masahiro, et al.
Pubblicazione: (2025) -
Signs Beat Floats: Low-Rank Double-Binary Adaptation for On-Device Fine-Tuning
di: Fujisawa, Yoshihiko, et al.
Pubblicazione: (2026) -
Exploring the Frontiers of LLMs in Psychological Applications: A Comprehensive Review
di: Ke, Luoma, et al.
Pubblicazione: (2024)