Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Xinwei, Li, Haojie, Liu, Hongyu, Ji, Xinyu, Li, Ruohan, Chen, Yule, Zhang, Yigeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mitigating Context-Memory Conflicts in LLMs through Dynamic Cognitive Reconciliation Decoding
di: Zhou, Yigeng, et al.
Pubblicazione: (2026)
di: Zhou, Yigeng, et al.
Pubblicazione: (2026)
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs
di: Li, Wu, et al.
Pubblicazione: (2026)
di: Li, Wu, et al.
Pubblicazione: (2026)
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
di: Guo, Ruiling, et al.
Pubblicazione: (2025)
di: Guo, Ruiling, et al.
Pubblicazione: (2025)
LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
di: Ren, Qibing, et al.
Pubblicazione: (2024)
di: Ren, Qibing, et al.
Pubblicazione: (2024)
Mind the Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering
di: Liu, Yaokun, et al.
Pubblicazione: (2026)
di: Liu, Yaokun, et al.
Pubblicazione: (2026)
Textual Self-attention Network: Test-Time Preference Optimization through Textual Gradient-based Attention
di: Mo, Shibing, et al.
Pubblicazione: (2025)
di: Mo, Shibing, et al.
Pubblicazione: (2025)
LLM Factoscope: Uncovering LLMs' Factual Discernment through Inner States Analysis
di: He, Jinwen, et al.
Pubblicazione: (2023)
di: He, Jinwen, et al.
Pubblicazione: (2023)
Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs
di: Zhang, Zhuoxuan, et al.
Pubblicazione: (2025)
di: Zhang, Zhuoxuan, et al.
Pubblicazione: (2025)
JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
di: Feng, Junlan, et al.
Pubblicazione: (2025)
di: Feng, Junlan, et al.
Pubblicazione: (2025)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
di: Hong, Junyuan, et al.
Pubblicazione: (2024)
di: Hong, Junyuan, et al.
Pubblicazione: (2024)
Positive and Risky Message Assessment for Music Products
di: Zhang, Yigeng, et al.
Pubblicazione: (2023)
di: Zhang, Yigeng, et al.
Pubblicazione: (2023)
Unveiling the Competitive Dynamics: A Comparative Evaluation of American and Chinese LLMs
di: Jiang, Zhenhui, et al.
Pubblicazione: (2024)
di: Jiang, Zhenhui, et al.
Pubblicazione: (2024)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2025)
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2025)
TEG-DB: A Comprehensive Dataset and Benchmark of Textual-Edge Graphs
di: Li, Zhuofeng, et al.
Pubblicazione: (2024)
di: Li, Zhuofeng, et al.
Pubblicazione: (2024)
Act-Adaptive Margin: Dynamically Calibrating Reward Models for Subjective Ambiguity
di: Fang, Feiteng, et al.
Pubblicazione: (2025)
di: Fang, Feiteng, et al.
Pubblicazione: (2025)
CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
di: Wu, Chengwei, et al.
Pubblicazione: (2025)
di: Wu, Chengwei, et al.
Pubblicazione: (2025)
Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
di: Xuan, Richmond Sin Jing, et al.
Pubblicazione: (2025)
di: Xuan, Richmond Sin Jing, et al.
Pubblicazione: (2025)
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering
di: Keluskar, Aryan, et al.
Pubblicazione: (2024)
di: Keluskar, Aryan, et al.
Pubblicazione: (2024)
Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
di: Noughabi, Havva Alizadeh, et al.
Pubblicazione: (2025)
di: Noughabi, Havva Alizadeh, et al.
Pubblicazione: (2025)
DISC: Plug-and-Play Decoding Intervention with Similarity of Characters for Chinese Spelling Check
di: Qiao, Ziheng, et al.
Pubblicazione: (2024)
di: Qiao, Ziheng, et al.
Pubblicazione: (2024)
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
di: Xiao, Yuxin, et al.
Pubblicazione: (2024)
di: Xiao, Yuxin, et al.
Pubblicazione: (2024)
nvBench 2.0: Resolving Ambiguity in Text-to-Visualization through Stepwise Reasoning
di: Luo, Tianqi, et al.
Pubblicazione: (2025)
di: Luo, Tianqi, et al.
Pubblicazione: (2025)
Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
di: Chen, Hongyu, et al.
Pubblicazione: (2025)
di: Chen, Hongyu, et al.
Pubblicazione: (2025)
Understanding Textual Capability Degradation in Speech LLMs via Parameter Importance Analysis
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
di: Mo, Lingbo, et al.
Pubblicazione: (2023)
di: Mo, Lingbo, et al.
Pubblicazione: (2023)
Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMs
di: Shore, Amber, et al.
Pubblicazione: (2025)
di: Shore, Amber, et al.
Pubblicazione: (2025)
Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
di: Chen, Yankai, et al.
Pubblicazione: (2025)
di: Chen, Yankai, et al.
Pubblicazione: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
SAFEFLOW: A Principled Protocol for Trustworthy and Transactional Autonomous Agent Systems
di: Li, Peiran, et al.
Pubblicazione: (2025)
di: Li, Peiran, et al.
Pubblicazione: (2025)
Harmonic LLMs are Trustworthy
di: Kersting, Nicholas S., et al.
Pubblicazione: (2024)
di: Kersting, Nicholas S., et al.
Pubblicazione: (2024)
Persona-Aware Alignment Framework for Personalized Dialogue Generation
di: Li, Guanrong, et al.
Pubblicazione: (2025)
di: Li, Guanrong, et al.
Pubblicazione: (2025)
Logic-Regularized Verifier Elicits Reasoning from LLMs
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
OrderBkd: Textual backdoor attack through repositioning
di: Alekseevskaia, Irina, et al.
Pubblicazione: (2024)
di: Alekseevskaia, Irina, et al.
Pubblicazione: (2024)
Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
di: Tan, Nelvin, et al.
Pubblicazione: (2025)
di: Tan, Nelvin, et al.
Pubblicazione: (2025)
Multi-objective Large Language Model Alignment with Hierarchical Experts
di: Li, Zhuo, et al.
Pubblicazione: (2025)
di: Li, Zhuo, et al.
Pubblicazione: (2025)
CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention
di: Sun, Yuxi, et al.
Pubblicazione: (2025)
di: Sun, Yuxi, et al.
Pubblicazione: (2025)
Beyond Textual Context: Structural Graph Encoding with Adaptive Space Alignment to alleviate the hallucination of LLMs
di: Zhang, Yifang, et al.
Pubblicazione: (2025)
di: Zhang, Yifang, et al.
Pubblicazione: (2025)
Aligning LLMs through Multi-perspective User Preference Ranking-based Feedback for Programming Question Answering
di: Yang, Hongyu, et al.
Pubblicazione: (2024)
di: Yang, Hongyu, et al.
Pubblicazione: (2024)
Scaling Textual Gradients via Sampling-Based Momentum
di: Ding, Zixin, et al.
Pubblicazione: (2025)
di: Ding, Zixin, et al.
Pubblicazione: (2025)
Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs
di: Wu, Xinwei, et al.
Pubblicazione: (2026)
di: Wu, Xinwei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Mitigating Context-Memory Conflicts in LLMs through Dynamic Cognitive Reconciliation Decoding
di: Zhou, Yigeng, et al.
Pubblicazione: (2026) -
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs
di: Li, Wu, et al.
Pubblicazione: (2026) -
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
di: Guo, Ruiling, et al.
Pubblicazione: (2025) -
LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
di: Ren, Qibing, et al.
Pubblicazione: (2024) -
Mind the Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering
di: Liu, Yaokun, et al.
Pubblicazione: (2026)