Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Yuxin, Wan, Chaoqun, Zhang, Yonggang, Wang, Wenxiao, Lin, Binbin, He, Xiaofei, Shen, Xu, Ye, Jieping |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SciPIP: An LLM-based Scientific Paper Idea Proposer
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
by: Yan, Shaotian, et al.
Published: (2025)
by: Yan, Shaotian, et al.
Published: (2025)
Concise and Organized Perception Facilitates Reasoning in Large Language Models
by: Liu, Junjie, et al.
Published: (2023)
by: Liu, Junjie, et al.
Published: (2023)
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach
by: Fan, Sinan, et al.
Published: (2025)
by: Fan, Sinan, et al.
Published: (2025)
Interpreting and Improving Large Language Models in Arithmetic Calculation
by: Zhang, Wei, et al.
Published: (2024)
by: Zhang, Wei, et al.
Published: (2024)
Controlling Thinking Speed in Reasoning Models
by: Lin, Zhengkai, et al.
Published: (2025)
by: Lin, Zhengkai, et al.
Published: (2025)
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning
by: Huang, Chenxi, et al.
Published: (2025)
by: Huang, Chenxi, et al.
Published: (2025)
JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
by: Feng, Junlan, et al.
Published: (2025)
by: Feng, Junlan, et al.
Published: (2025)
Model Compression and Efficient Inference for Large Language Models: A Survey
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental Learning
by: Chen, Minghao, et al.
Published: (2024)
by: Chen, Minghao, et al.
Published: (2024)
From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Rethinking All Evidence: Enhancing Trustworthy Retrieval-Augmented Generation via Conflict-Driven Summarization
by: Chen, Juan, et al.
Published: (2025)
by: Chen, Juan, et al.
Published: (2025)
CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention
by: Sun, Yuxi, et al.
Published: (2025)
by: Sun, Yuxi, et al.
Published: (2025)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
by: Christopoulou, Fenia, et al.
Published: (2024)
by: Christopoulou, Fenia, et al.
Published: (2024)
Delving into the Reversal Curse: How Far Can Large Language Models Generalize?
by: Lin, Zhengkai, et al.
Published: (2024)
by: Lin, Zhengkai, et al.
Published: (2024)
Sparse Activation Editing for Reliable Instruction Following in Narratives
by: Zhao, Runcong, et al.
Published: (2025)
by: Zhao, Runcong, et al.
Published: (2025)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
by: Banayeeanzade, Amin, et al.
Published: (2025)
by: Banayeeanzade, Amin, et al.
Published: (2025)
SAC-KG: Exploiting Large Language Models as Skilled Automatic Constructors for Domain Knowledge Graphs
by: Chen, Hanzhu, et al.
Published: (2024)
by: Chen, Hanzhu, et al.
Published: (2024)
Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection
by: Wang, Bing, et al.
Published: (2026)
by: Wang, Bing, et al.
Published: (2026)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
Activations as Features: Probing LLMs for Generalizable Essay Scoring Representations
by: Chi, Jinwei, et al.
Published: (2025)
by: Chi, Jinwei, et al.
Published: (2025)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
by: Peng, Dan, et al.
Published: (2025)
by: Peng, Dan, et al.
Published: (2025)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
by: Shi, Zhiyuan, et al.
Published: (2026)
by: Shi, Zhiyuan, et al.
Published: (2026)
Harmonic LLMs are Trustworthy
by: Kersting, Nicholas S., et al.
Published: (2024)
by: Kersting, Nicholas S., et al.
Published: (2024)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
by: Hong, Junyuan, et al.
Published: (2024)
by: Hong, Junyuan, et al.
Published: (2024)
Inference-time Alignment via Sparse Junction Steering
by: Hu, Runyi, et al.
Published: (2026)
by: Hu, Runyi, et al.
Published: (2026)
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
by: Zhong, Chengzhi, et al.
Published: (2025)
by: Zhong, Chengzhi, et al.
Published: (2025)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
by: Wu, Xinwei, et al.
Published: (2025)
by: Wu, Xinwei, et al.
Published: (2025)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
by: Fang, Yi, et al.
Published: (2026)
by: Fang, Yi, et al.
Published: (2026)
Know Your Needs Better: Towards Structured Understanding of Marketer Demands with Analogical Reasoning Augmented LLMs
by: Wang, Junjie, et al.
Published: (2024)
by: Wang, Junjie, et al.
Published: (2024)
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
by: Sun, Zhongxiang, et al.
Published: (2026)
by: Sun, Zhongxiang, et al.
Published: (2026)
Enhancing Idiomatic Representation in Multiple Languages via an Adaptive Contrastive Triplet Loss
by: He, Wei, et al.
Published: (2024)
by: He, Wei, et al.
Published: (2024)
Enhancing the Traditional Chinese Medicine Capabilities of Large Language Model through Reinforcement Learning from AI Feedback
by: Yu, Song, et al.
Published: (2024)
by: Yu, Song, et al.
Published: (2024)
Explore the Reasoning Capability of LLMs in the Chess Testbed
by: Wang, Shu, et al.
Published: (2024)
by: Wang, Shu, et al.
Published: (2024)
Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs
by: Zhang, Zhuoxuan, et al.
Published: (2025)
by: Zhang, Zhuoxuan, et al.
Published: (2025)
Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents
by: Alvarez, Aitor Arronte, et al.
Published: (2026)
by: Alvarez, Aitor Arronte, et al.
Published: (2026)
Less is More: Sparse Watermarking in LLMs with Enhanced Text Quality
by: Hoang, Duy C., et al.
Published: (2024)
by: Hoang, Duy C., et al.
Published: (2024)
Achieving Sparse Activation in Small Language Models
by: Song, Jifeng, et al.
Published: (2024)
by: Song, Jifeng, et al.
Published: (2024)
Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension Perception
by: Huang, Yuncheng, et al.
Published: (2023)
by: Huang, Yuncheng, et al.
Published: (2023)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
Similar Items
-
SciPIP: An LLM-based Scientific Paper Idea Proposer
by: Wang, Wenxiao, et al.
Published: (2024) -
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
by: Yan, Shaotian, et al.
Published: (2025) -
Concise and Organized Perception Facilitates Reasoning in Large Language Models
by: Liu, Junjie, et al.
Published: (2023) -
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach
by: Fan, Sinan, et al.
Published: (2025) -
Interpreting and Improving Large Language Models in Arithmetic Calculation
by: Zhang, Wei, et al.
Published: (2024)