Understanding How Value Neurons Shape the Generation of Specified Values in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Su, Yi, Zhang, Jiayi, Yang, Shu, Wang, Xinhai, Hu, Lijie, Wang, Di |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Understanding and Mitigating Cross-lingual Privacy Leakage via Language-specific and Universal Privacy Neurons
di: Dong, Wenshuo, et al.
Pubblicazione: (2025)
di: Dong, Wenshuo, et al.
Pubblicazione: (2025)
Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness
di: Yang, Zhipeng, et al.
Pubblicazione: (2026)
di: Yang, Zhipeng, et al.
Pubblicazione: (2026)
Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
di: Zhang, Jiayi, et al.
Pubblicazione: (2025)
di: Zhang, Jiayi, et al.
Pubblicazione: (2025)
Understanding the Repeat Curse in Large Language Models from a Feature Perspective
di: Yao, Junchi, et al.
Pubblicazione: (2025)
di: Yao, Junchi, et al.
Pubblicazione: (2025)
MoRAL: MoE Augmented LoRA for LLMs' Lifelong Learning
di: Yang, Shu, et al.
Pubblicazione: (2024)
di: Yang, Shu, et al.
Pubblicazione: (2024)
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities
di: Zhang, Xiangxu, et al.
Pubblicazione: (2026)
di: Zhang, Xiangxu, et al.
Pubblicazione: (2026)
Exploring the Personality Traits of LLMs through Latent Features Steering
di: Yang, Shu, et al.
Pubblicazione: (2024)
di: Yang, Shu, et al.
Pubblicazione: (2024)
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
di: Zhang, Hongbo, et al.
Pubblicazione: (2025)
di: Zhang, Hongbo, et al.
Pubblicazione: (2025)
Dialectical Alignment: Resolving the Tension of 3H and Security Threats of LLMs
di: Yang, Shu, et al.
Pubblicazione: (2024)
di: Yang, Shu, et al.
Pubblicazione: (2024)
NeuronMLP: Efficient LLM Inference via Singular Value Decomposition Compression and Tiling on AWS Trainium
di: Song, Dinghong, et al.
Pubblicazione: (2025)
di: Song, Dinghong, et al.
Pubblicazione: (2025)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
di: Zhou, Wenrui, et al.
Pubblicazione: (2025)
di: Zhou, Wenrui, et al.
Pubblicazione: (2025)
MONAL: Model Autophagy Analysis for Modeling Human-AI Interactions
di: Yang, Shu, et al.
Pubblicazione: (2024)
di: Yang, Shu, et al.
Pubblicazione: (2024)
Following the Whispers of Values: Unraveling Neural Mechanisms Behind Value-Oriented Behaviors in LLMs
di: Hu, Ling, et al.
Pubblicazione: (2025)
di: Hu, Ling, et al.
Pubblicazione: (2025)
Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics
di: Ananthanarayanan, Samhruth, et al.
Pubblicazione: (2026)
di: Ananthanarayanan, Samhruth, et al.
Pubblicazione: (2026)
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models
di: Ren, Yuanyi, et al.
Pubblicazione: (2024)
di: Ren, Yuanyi, et al.
Pubblicazione: (2024)
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
di: Jiang, Han, et al.
Pubblicazione: (2024)
di: Jiang, Han, et al.
Pubblicazione: (2024)
AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations
di: Yun, Bhada, et al.
Pubblicazione: (2026)
di: Yun, Bhada, et al.
Pubblicazione: (2026)
CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters
di: Sun, Ao, et al.
Pubblicazione: (2026)
di: Sun, Ao, et al.
Pubblicazione: (2026)
Flames: Benchmarking Value Alignment of LLMs in Chinese
di: Huang, Kexin, et al.
Pubblicazione: (2023)
di: Huang, Kexin, et al.
Pubblicazione: (2023)
Cash or Comfort? How LLMs Value Your Inconvenience
di: Cedro, Mateusz, et al.
Pubblicazione: (2025)
di: Cedro, Mateusz, et al.
Pubblicazione: (2025)
Lost in Literalism: How Supervised Training Shapes Translationese in LLMs
di: Li, Yafu, et al.
Pubblicazione: (2025)
di: Li, Yafu, et al.
Pubblicazione: (2025)
Understanding Reasoning in Chain-of-Thought from the Hopfieldian View
di: Hu, Lijie, et al.
Pubblicazione: (2024)
di: Hu, Lijie, et al.
Pubblicazione: (2024)
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
di: You, Liangliang, et al.
Pubblicazione: (2025)
di: You, Liangliang, et al.
Pubblicazione: (2025)
Enhancing Agentic RL with Progressive Reward Shaping and Value-based Sampling Policy Optimization
di: Su, Jianghao, et al.
Pubblicazione: (2025)
di: Su, Jianghao, et al.
Pubblicazione: (2025)
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
di: Wang, Jun, et al.
Pubblicazione: (2025)
di: Wang, Jun, et al.
Pubblicazione: (2025)
SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs
di: Jie, Shibo, et al.
Pubblicazione: (2025)
di: Jie, Shibo, et al.
Pubblicazione: (2025)
ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models
di: Zhang, Zhaowei, et al.
Pubblicazione: (2023)
di: Zhang, Zhaowei, et al.
Pubblicazione: (2023)
Edu-Values: Towards Evaluating the Chinese Education Values of Large Language Models
di: Zhang, Peiyi, et al.
Pubblicazione: (2024)
di: Zhang, Peiyi, et al.
Pubblicazione: (2024)
Under the Shadow of Babel: How Language Shapes Reasoning in LLMs
di: Wang, Chenxi, et al.
Pubblicazione: (2025)
di: Wang, Chenxi, et al.
Pubblicazione: (2025)
DictLLM: Harnessing Key-Value Data Structures with Large Language Models for Enhanced Medical Diagnostics
di: Guo, YiQiu, et al.
Pubblicazione: (2024)
di: Guo, YiQiu, et al.
Pubblicazione: (2024)
Implicit Values Embedded in How Humans and LLMs Complete Subjective Everyday Tasks
di: Arunasalam, Arjun, et al.
Pubblicazione: (2025)
di: Arunasalam, Arjun, et al.
Pubblicazione: (2025)
ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs
di: Shen, Hua, et al.
Pubblicazione: (2024)
di: Shen, Hua, et al.
Pubblicazione: (2024)
ValueSim: Generating Backstories to Model Individual Value Systems
di: Du, Bangde, et al.
Pubblicazione: (2025)
di: Du, Bangde, et al.
Pubblicazione: (2025)
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
di: Shen, Hua, et al.
Pubblicazione: (2025)
di: Shen, Hua, et al.
Pubblicazione: (2025)
Specifying Genericity through Inclusiveness and Abstractness Continuous Scales
di: Collacciani, Claudia, et al.
Pubblicazione: (2024)
di: Collacciani, Claudia, et al.
Pubblicazione: (2024)
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling
di: Zhang, Zhen, et al.
Pubblicazione: (2026)
di: Zhang, Zhen, et al.
Pubblicazione: (2026)
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
Language Shapes Mental Health Evaluations in Large Language Models
di: Xu, Jiayi, et al.
Pubblicazione: (2026)
di: Xu, Jiayi, et al.
Pubblicazione: (2026)
How Value Induction Reshapes LLM Behaviour
di: Arora, Arnav, et al.
Pubblicazione: (2026)
di: Arora, Arnav, et al.
Pubblicazione: (2026)
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Understanding and Mitigating Cross-lingual Privacy Leakage via Language-specific and Universal Privacy Neurons
di: Dong, Wenshuo, et al.
Pubblicazione: (2025) -
Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness
di: Yang, Zhipeng, et al.
Pubblicazione: (2026) -
Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
di: Zhang, Jiayi, et al.
Pubblicazione: (2025) -
Understanding the Repeat Curse in Large Language Models from a Feature Perspective
di: Yao, Junchi, et al.
Pubblicazione: (2025) -
MoRAL: MoE Augmented LoRA for LLMs' Lifelong Learning
di: Yang, Shu, et al.
Pubblicazione: (2024)