Good Parenting is all you need -- Multi-agentic LLM Hallucination Mitigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kwartler, Ted, Berman, Matthew, Aqrawi, Alan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)
von: Kwartler, Ted, et al.
Veröffentlicht: (2024)
von: Kwartler, Ted, et al.
Veröffentlicht: (2024)
Well, that escalated quickly: The Single-Turn Crescendo Attack (STCA)
von: Aqrawi, Alan, et al.
Veröffentlicht: (2024)
von: Aqrawi, Alan, et al.
Veröffentlicht: (2024)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
von: Chakraborty, Trishna, et al.
Veröffentlicht: (2024)
von: Chakraborty, Trishna, et al.
Veröffentlicht: (2024)
LLM Dataset Inference: Did you train on my dataset?
von: Maini, Pratyush, et al.
Veröffentlicht: (2024)
von: Maini, Pratyush, et al.
Veröffentlicht: (2024)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
von: Singh, Himanshu, et al.
Veröffentlicht: (2026)
von: Singh, Himanshu, et al.
Veröffentlicht: (2026)
Is poisoning a real threat to LLM alignment? Maybe more so than you think
von: Pathmanathan, Pankayaraj, et al.
Veröffentlicht: (2024)
von: Pathmanathan, Pankayaraj, et al.
Veröffentlicht: (2024)
Ingest-And-Ground: Dispelling Hallucinations from Continually-Pretrained LLMs with RAG
von: Fang, Chenhao, et al.
Veröffentlicht: (2024)
von: Fang, Chenhao, et al.
Veröffentlicht: (2024)
Multi-use LLM Watermarking and the False Detection Problem
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
WorldCup Sampling for Multi-bit LLM Watermarking
von: Wang, Yidan, et al.
Veröffentlicht: (2026)
von: Wang, Yidan, et al.
Veröffentlicht: (2026)
How Good is Post-Hoc Watermarking With Language Model Rephrasing?
von: Fernandez, Pierre, et al.
Veröffentlicht: (2025)
von: Fernandez, Pierre, et al.
Veröffentlicht: (2025)
The Double-edged Sword of LLM-based Data Reconstruction: Understanding and Mitigating Contextual Vulnerability in Word-level Differential Privacy Text Sanitization
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
von: Assogba, Yannick, et al.
Veröffentlicht: (2026)
von: Assogba, Yannick, et al.
Veröffentlicht: (2026)
HE is all you need: Compressing FHE Ciphertexts using Additive HE
von: Mahdavi, Rasoul Akhavan, et al.
Veröffentlicht: (2023)
von: Mahdavi, Rasoul Akhavan, et al.
Veröffentlicht: (2023)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
Mitigating Jailbreaks with Intent-Aware LLMs
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
DAVE: A Policy-Enforcing LLM Spokesperson for Secure Multi-Document Data Sharing
von: Brinkhege, René, et al.
Veröffentlicht: (2026)
von: Brinkhege, René, et al.
Veröffentlicht: (2026)
CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems
von: Wen, Yan, et al.
Veröffentlicht: (2025)
von: Wen, Yan, et al.
Veröffentlicht: (2025)
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
von: Wu, Yiran, et al.
Veröffentlicht: (2025)
von: Wu, Yiran, et al.
Veröffentlicht: (2025)
SoK: Privacy Risks and Mitigations in Retrieval-Augmented Generation Systems
von: Bodea, Andreea-Elena, et al.
Veröffentlicht: (2026)
von: Bodea, Andreea-Elena, et al.
Veröffentlicht: (2026)
Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
von: Luo, Zhifan, et al.
Veröffentlicht: (2025)
von: Luo, Zhifan, et al.
Veröffentlicht: (2025)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
MBTSAD: Mitigating Backdoors in Language Models Based on Token Splitting and Attention Distillation
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
An Investigation on Group Query Hallucination Attacks
von: Miao, Kehao, et al.
Veröffentlicht: (2025)
von: Miao, Kehao, et al.
Veröffentlicht: (2025)
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
von: He, Xuanli, et al.
Veröffentlicht: (2024)
von: He, Xuanli, et al.
Veröffentlicht: (2024)
WaterPool: A Watermark Mitigating Trade-offs among Imperceptibility, Efficacy and Robustness
von: Huang, Baizhou, et al.
Veröffentlicht: (2024)
von: Huang, Baizhou, et al.
Veröffentlicht: (2024)
Mitigating Cyber Risk in the Age of Open-Weight LLMs: Policy Gaps and Technical Realities
von: de Gregorio, Alfonso
Veröffentlicht: (2025)
von: de Gregorio, Alfonso
Veröffentlicht: (2025)
LLM Reinforcement in Context
von: Rivasseau, Thomas
Veröffentlicht: (2025)
von: Rivasseau, Thomas
Veröffentlicht: (2025)
LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
von: Liang, Buyun, et al.
Veröffentlicht: (2025)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
von: Liang, Buyun, et al.
Veröffentlicht: (2026)
PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
von: Hughes, Anthony, et al.
Veröffentlicht: (2025)
von: Hughes, Anthony, et al.
Veröffentlicht: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
Watermarking LLM Agent Trajectories
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
von: Meng, Wenlong, et al.
Veröffentlicht: (2026)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
von: Noah, Amit Finkman, et al.
Veröffentlicht: (2024)
von: Noah, Amit Finkman, et al.
Veröffentlicht: (2024)
Proactive defense against LLM Jailbreak
von: Zhao, Weiliang, et al.
Veröffentlicht: (2025)
von: Zhao, Weiliang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)
von: Kwartler, Ted, et al.
Veröffentlicht: (2024) -
Well, that escalated quickly: The Single-Turn Crescendo Attack (STCA)
von: Aqrawi, Alan, et al.
Veröffentlicht: (2024) -
Cross-Modal Safety Alignment: Is textual unlearning all you need?
von: Chakraborty, Trishna, et al.
Veröffentlicht: (2024) -
LLM Dataset Inference: Did you train on my dataset?
von: Maini, Pratyush, et al.
Veröffentlicht: (2024) -
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)