Saved in:
| Main Authors: | Wu, Xinwei, Li, Haojie, Liu, Hongyu, Ji, Xinyu, Li, Ruohan, Chen, Yule, Zhang, Yigeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.23121 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Context-Memory Conflicts in LLMs through Dynamic Cognitive Reconciliation Decoding
by: Zhou, Yigeng, et al.
Published: (2026)
by: Zhou, Yigeng, et al.
Published: (2026)
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs
by: Li, Wu, et al.
Published: (2026)
by: Li, Wu, et al.
Published: (2026)
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
by: Guo, Ruiling, et al.
Published: (2025)
by: Guo, Ruiling, et al.
Published: (2025)
LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
by: Ren, Qibing, et al.
Published: (2024)
by: Ren, Qibing, et al.
Published: (2024)
Textual Self-attention Network: Test-Time Preference Optimization through Textual Gradient-based Attention
by: Mo, Shibing, et al.
Published: (2025)
by: Mo, Shibing, et al.
Published: (2025)
Mind the Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering
by: Liu, Yaokun, et al.
Published: (2026)
by: Liu, Yaokun, et al.
Published: (2026)
Positive and Risky Message Assessment for Music Products
by: Zhang, Yigeng, et al.
Published: (2023)
by: Zhang, Yigeng, et al.
Published: (2023)
LLM Factoscope: Uncovering LLMs' Factual Discernment through Inner States Analysis
by: He, Jinwen, et al.
Published: (2023)
by: He, Jinwen, et al.
Published: (2023)
Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs
by: Zhang, Zhuoxuan, et al.
Published: (2025)
by: Zhang, Zhuoxuan, et al.
Published: (2025)
JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
by: Feng, Junlan, et al.
Published: (2025)
by: Feng, Junlan, et al.
Published: (2025)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
by: Hong, Junyuan, et al.
Published: (2024)
by: Hong, Junyuan, et al.
Published: (2024)
TEG-DB: A Comprehensive Dataset and Benchmark of Textual-Edge Graphs
by: Li, Zhuofeng, et al.
Published: (2024)
by: Li, Zhuofeng, et al.
Published: (2024)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
by: Banayeeanzade, Amin, et al.
Published: (2025)
by: Banayeeanzade, Amin, et al.
Published: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Unveiling the Competitive Dynamics: A Comparative Evaluation of American and Chinese LLMs
by: Jiang, Zhenhui, et al.
Published: (2024)
by: Jiang, Zhenhui, et al.
Published: (2024)
Act-Adaptive Margin: Dynamically Calibrating Reward Models for Subjective Ambiguity
by: Fang, Feiteng, et al.
Published: (2025)
by: Fang, Feiteng, et al.
Published: (2025)
CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
by: Wu, Chengwei, et al.
Published: (2025)
by: Wu, Chengwei, et al.
Published: (2025)
DISC: Plug-and-Play Decoding Intervention with Similarity of Characters for Chinese Spelling Check
by: Qiao, Ziheng, et al.
Published: (2024)
by: Qiao, Ziheng, et al.
Published: (2024)
Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
by: Chen, Hongyu, et al.
Published: (2025)
by: Chen, Hongyu, et al.
Published: (2025)
Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
by: Xuan, Richmond Sin Jing, et al.
Published: (2025)
by: Xuan, Richmond Sin Jing, et al.
Published: (2025)
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering
by: Keluskar, Aryan, et al.
Published: (2024)
by: Keluskar, Aryan, et al.
Published: (2024)
Multi-objective Large Language Model Alignment with Hierarchical Experts
by: Li, Zhuo, et al.
Published: (2025)
by: Li, Zhuo, et al.
Published: (2025)
Harmonic LLMs are Trustworthy
by: Kersting, Nicholas S., et al.
Published: (2024)
by: Kersting, Nicholas S., et al.
Published: (2024)
Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
by: Noughabi, Havva Alizadeh, et al.
Published: (2025)
nvBench 2.0: Resolving Ambiguity in Text-to-Visualization through Stepwise Reasoning
by: Luo, Tianqi, et al.
Published: (2025)
by: Luo, Tianqi, et al.
Published: (2025)
Understanding Textual Capability Degradation in Speech LLMs via Parameter Importance Analysis
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
by: Xiao, Yuxin, et al.
Published: (2024)
by: Xiao, Yuxin, et al.
Published: (2024)
FragileFlow: Spectral Control of Correct-but-Fragile Predictions for Foundation Model Robustness
by: Li, Zhuoyun, et al.
Published: (2026)
by: Li, Zhuoyun, et al.
Published: (2026)
Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs
by: Wu, Xinwei, et al.
Published: (2026)
by: Wu, Xinwei, et al.
Published: (2026)
Logic-Regularized Verifier Elicits Reasoning from LLMs
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
by: Chen, Yankai, et al.
Published: (2025)
by: Chen, Yankai, et al.
Published: (2025)
Persona-Aware Alignment Framework for Personalized Dialogue Generation
by: Li, Guanrong, et al.
Published: (2025)
by: Li, Guanrong, et al.
Published: (2025)
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
by: Mo, Lingbo, et al.
Published: (2023)
by: Mo, Lingbo, et al.
Published: (2023)
Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMs
by: Shore, Amber, et al.
Published: (2025)
by: Shore, Amber, et al.
Published: (2025)
SAFEFLOW: A Principled Protocol for Trustworthy and Transactional Autonomous Agent Systems
by: Li, Peiran, et al.
Published: (2025)
by: Li, Peiran, et al.
Published: (2025)
Aligning LLMs through Multi-perspective User Preference Ranking-based Feedback for Programming Question Answering
by: Yang, Hongyu, et al.
Published: (2024)
by: Yang, Hongyu, et al.
Published: (2024)
OrderBkd: Textual backdoor attack through repositioning
by: Alekseevskaia, Irina, et al.
Published: (2024)
by: Alekseevskaia, Irina, et al.
Published: (2024)
Scaling Textual Gradients via Sampling-Based Momentum
by: Ding, Zixin, et al.
Published: (2025)
by: Ding, Zixin, et al.
Published: (2025)
Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch
by: Wen, Xueru, et al.
Published: (2025)
by: Wen, Xueru, et al.
Published: (2025)
Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
by: Tan, Nelvin, et al.
Published: (2025)
by: Tan, Nelvin, et al.
Published: (2025)
Similar Items
-
Mitigating Context-Memory Conflicts in LLMs through Dynamic Cognitive Reconciliation Decoding
by: Zhou, Yigeng, et al.
Published: (2026) -
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs
by: Li, Wu, et al.
Published: (2026) -
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
by: Guo, Ruiling, et al.
Published: (2025) -
LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
by: Ren, Qibing, et al.
Published: (2024) -
Textual Self-attention Network: Test-Time Preference Optimization through Textual Gradient-based Attention
by: Mo, Shibing, et al.
Published: (2025)