How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
Fuente:
arXiv
Guardado en:
| Autores principales: | Mo, Lingbo, Wang, Boshi, Chen, Muhao, Sun, Huan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
por: Jones, Jaylen, et al.
Publicado: (2024)
por: Jones, Jaylen, et al.
Publicado: (2024)
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
por: Zhang, Kai, et al.
Publicado: (2023)
por: Zhang, Kai, et al.
Publicado: (2023)
A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
por: Mo, Lingbo, et al.
Publicado: (2024)
por: Mo, Lingbo, et al.
Publicado: (2024)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
por: Tong, Terry, et al.
Publicado: (2025)
por: Tong, Terry, et al.
Publicado: (2025)
How Fragile is Relation Extraction under Entity Replacements?
por: Wang, Yiwei, et al.
Publicado: (2023)
por: Wang, Yiwei, et al.
Publicado: (2023)
Defend LLMs Through Self-Consciousness
por: Huang, Boshi, et al.
Publicado: (2025)
por: Huang, Boshi, et al.
Publicado: (2025)
ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?
por: Hui, Zheng, et al.
Publicado: (2024)
por: Hui, Zheng, et al.
Publicado: (2024)
CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention
por: Sun, Yuxi, et al.
Publicado: (2025)
por: Sun, Yuxi, et al.
Publicado: (2025)
HierTOD: A Task-Oriented Dialogue System Driven by Hierarchical Goals
por: Mo, Lingbo, et al.
Publicado: (2024)
por: Mo, Lingbo, et al.
Publicado: (2024)
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
por: Schoene, Annika M, et al.
Publicado: (2025)
por: Schoene, Annika M, et al.
Publicado: (2025)
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
por: Li, Bangzheng, et al.
Publicado: (2023)
por: Li, Bangzheng, et al.
Publicado: (2023)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
por: Xu, Nan, et al.
Publicado: (2024)
por: Xu, Nan, et al.
Publicado: (2024)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
por: Banayeeanzade, Amin, et al.
Publicado: (2025)
por: Banayeeanzade, Amin, et al.
Publicado: (2025)
ToolBridge: An Open-Source Dataset to Equip LLMs with External Tool Capabilities
por: Jin, Zhenchao, et al.
Publicado: (2024)
por: Jin, Zhenchao, et al.
Publicado: (2024)
JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
por: Feng, Junlan, et al.
Publicado: (2025)
por: Feng, Junlan, et al.
Publicado: (2025)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
por: Wu, Xinwei, et al.
Publicado: (2025)
por: Wu, Xinwei, et al.
Publicado: (2025)
REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading
por: Zhao, Chengshuai, et al.
Publicado: (2026)
por: Zhao, Chengshuai, et al.
Publicado: (2026)
Tucano 2 Cool: Better Open Source LLMs for Portuguese
por: Corrêa, Nicholas Kluge, et al.
Publicado: (2026)
por: Corrêa, Nicholas Kluge, et al.
Publicado: (2026)
A Framework to Assess Multilingual Vulnerabilities of LLMs
por: Tang, Likai, et al.
Publicado: (2025)
por: Tang, Likai, et al.
Publicado: (2025)
Harmonic LLMs are Trustworthy
por: Kersting, Nicholas S., et al.
Publicado: (2024)
por: Kersting, Nicholas S., et al.
Publicado: (2024)
Is Open-Source There Yet? A Comparative Study on Commercial and Open-Source LLMs in Their Ability to Label Chest X-Ray Reports
por: Dorfner, Felix J., et al.
Publicado: (2024)
por: Dorfner, Felix J., et al.
Publicado: (2024)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
por: Xu, Jiashu, et al.
Publicado: (2023)
por: Xu, Jiashu, et al.
Publicado: (2023)
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
por: Handa, Divij, et al.
Publicado: (2024)
por: Handa, Divij, et al.
Publicado: (2024)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
por: Wang, Boshi, et al.
Publicado: (2024)
por: Wang, Boshi, et al.
Publicado: (2024)
Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English
por: Anand, Avinash, et al.
Publicado: (2024)
por: Anand, Avinash, et al.
Publicado: (2024)
Code Execution as Grounded Supervision for LLM Reasoning
por: Jung, Dongwon, et al.
Publicado: (2025)
por: Jung, Dongwon, et al.
Publicado: (2025)
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
por: Xiao, Yuxin, et al.
Publicado: (2024)
por: Xiao, Yuxin, et al.
Publicado: (2024)
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering
por: Keluskar, Aryan, et al.
Publicado: (2024)
por: Keluskar, Aryan, et al.
Publicado: (2024)
SudoLM: Learning Access Control of Parametric Knowledge with Authorization Alignment
por: Liu, Qin, et al.
Publicado: (2024)
por: Liu, Qin, et al.
Publicado: (2024)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
por: Hong, Junyuan, et al.
Publicado: (2024)
por: Hong, Junyuan, et al.
Publicado: (2024)
Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
por: Golde, Jonas, et al.
Publicado: (2023)
por: Golde, Jonas, et al.
Publicado: (2023)
Extreme Speech Classification in the Era of LLMs: Exploring Open-Source and Proprietary Models
por: Mahajan, Sarthak, et al.
Publicado: (2025)
por: Mahajan, Sarthak, et al.
Publicado: (2025)
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
por: Wang, Huandong, et al.
Publicado: (2025)
por: Wang, Huandong, et al.
Publicado: (2025)
DeepEdit: Knowledge Editing as Decoding with Constraints
por: Wang, Yiwei, et al.
Publicado: (2024)
por: Wang, Yiwei, et al.
Publicado: (2024)
ExaRanker-Open: Synthetic Explanation for IR using Open-Source LLMs
por: Ferraretto, Fernando, et al.
Publicado: (2024)
por: Ferraretto, Fernando, et al.
Publicado: (2024)
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
por: Zhou, Weibo, et al.
Publicado: (2025)
por: Zhou, Weibo, et al.
Publicado: (2025)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
por: Wen, Xiaofei, et al.
Publicado: (2025)
por: Wen, Xiaofei, et al.
Publicado: (2025)
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
por: Liu, Xiaogeng, et al.
Publicado: (2023)
por: Liu, Xiaogeng, et al.
Publicado: (2023)
Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents
por: Alvarez, Aitor Arronte, et al.
Publicado: (2026)
por: Alvarez, Aitor Arronte, et al.
Publicado: (2026)
Red Teaming Language Models for Processing Contradictory Dialogues
por: Wen, Xiaofei, et al.
Publicado: (2024)
por: Wen, Xiaofei, et al.
Publicado: (2024)
Ejemplares similares
-
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
por: Jones, Jaylen, et al.
Publicado: (2024) -
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
por: Zhang, Kai, et al.
Publicado: (2023) -
A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
por: Mo, Lingbo, et al.
Publicado: (2024) -
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
por: Tong, Terry, et al.
Publicado: (2025) -
How Fragile is Relation Extraction under Entity Replacements?
por: Wang, Yiwei, et al.
Publicado: (2023)