VEXA: Evidence-Grounded and Persona-Adaptive Explanations for Scam Risk Sensemaking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | An, Heajun, Ng, Connor, Dulal, Sandesh Sharma, Kim, Junghwan, Cho, Jin-Hee |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
Slow is Fast! Dissecting Ethereum's Slow Liquidity Drain Scams
von: Tran, Minh Trung, et al.
Veröffentlicht: (2025)
von: Tran, Minh Trung, et al.
Veröffentlicht: (2025)
Persona-Model Collapse in Emergent Misalignment
von: Costa, Davi Bastos, et al.
Veröffentlicht: (2026)
von: Costa, Davi Bastos, et al.
Veröffentlicht: (2026)
Confidence Elicitation: A New Attack Vector for Large Language Models
von: Formento, Brian, et al.
Veröffentlicht: (2025)
von: Formento, Brian, et al.
Veröffentlicht: (2025)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
Copyright-Protected Language Generation via Adaptive Model Fusion
von: Abad, Javier, et al.
Veröffentlicht: (2024)
von: Abad, Javier, et al.
Veröffentlicht: (2024)
Malicious and Unintentional Disclosure Risks in Large Language Models for Code Generation
von: Rabin, Rafiqul, et al.
Veröffentlicht: (2025)
von: Rabin, Rafiqul, et al.
Veröffentlicht: (2025)
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
A Framework for Cost-Effective and Self-Adaptive LLM Shaking and Recovery Mechanism
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
Adaptive Pre-training Data Detection for Large Language Models via Surprising Tokens
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model Queries
von: Huang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
von: Huang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
von: Ma, Avery, et al.
Veröffentlicht: (2025)
von: Ma, Avery, et al.
Veröffentlicht: (2025)
DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
von: Wang, Lionel Z., et al.
Veröffentlicht: (2026)
von: Wang, Lionel Z., et al.
Veröffentlicht: (2026)
SPML: A DSL for Defending Language Models Against Prompt Attacks
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2024)
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2024)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
von: Yoon, Do-hyeon, et al.
Veröffentlicht: (2025)
von: Yoon, Do-hyeon, et al.
Veröffentlicht: (2025)
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2025)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
von: Kim, Taeyoun, et al.
Veröffentlicht: (2024)
von: Kim, Taeyoun, et al.
Veröffentlicht: (2024)
Data-centric NLP Backdoor Defense from the Lens of Memorization
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
Graphene: Infrastructure Security Posture Analysis with AI-generated Attack Graphs
von: Jin, Xin, et al.
Veröffentlicht: (2023)
von: Jin, Xin, et al.
Veröffentlicht: (2023)
Exploiting Code Symmetries for Learning Program Semantics
von: Pei, Kexin, et al.
Veröffentlicht: (2023)
von: Pei, Kexin, et al.
Veröffentlicht: (2023)
Explaining the Model, Protecting Your Data: Revealing and Mitigating the Data Privacy Risks of Post-Hoc Model Explanations via Membership Inference
von: Huang, Catherine, et al.
Veröffentlicht: (2024)
von: Huang, Catherine, et al.
Veröffentlicht: (2024)
The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
von: Yan, Jun, et al.
Veröffentlicht: (2023)
von: Yan, Jun, et al.
Veröffentlicht: (2023)
Graph Neural Network Explanations are Fragile
von: Li, Jiate, et al.
Veröffentlicht: (2024)
von: Li, Jiate, et al.
Veröffentlicht: (2024)
LiSA: Lifelong Safety Adaptation via Conservative Policy Induction
von: Kim, Minbeom, et al.
Veröffentlicht: (2026)
von: Kim, Minbeom, et al.
Veröffentlicht: (2026)
A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
von: Xin, Rui, et al.
Veröffentlicht: (2025)
von: Xin, Rui, et al.
Veröffentlicht: (2025)
Representation Bending for Large Language Model Safety
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
Distinguishing Scams and Fraud with Ensemble Learning
von: Chadalavada, Isha, et al.
Veröffentlicht: (2024)
von: Chadalavada, Isha, et al.
Veröffentlicht: (2024)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
von: Lee, Seanie, et al.
Veröffentlicht: (2024)
von: Lee, Seanie, et al.
Veröffentlicht: (2024)
Gandalf the Red: Adaptive Security for LLMs
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
von: Pfister, Niklas, et al.
Veröffentlicht: (2025)
International Students and Scams: At Risk Abroad
von: Zhang, Katherine, et al.
Veröffentlicht: (2025)
von: Zhang, Katherine, et al.
Veröffentlicht: (2025)
Graph Reconstruction from Differentially Private GNN Explanations
von: Sahoo, Rishi Raj, et al.
Veröffentlicht: (2026)
von: Sahoo, Rishi Raj, et al.
Veröffentlicht: (2026)
Dual Explanations via Subgraph Matching for Malware Detection
von: Shokouhinejad, Hossein, et al.
Veröffentlicht: (2025)
von: Shokouhinejad, Hossein, et al.
Veröffentlicht: (2025)
Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation
von: Morasso, Cristian, et al.
Veröffentlicht: (2026)
von: Morasso, Cristian, et al.
Veröffentlicht: (2026)
Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents
von: Wang, Chunxiao
Veröffentlicht: (2026)
von: Wang, Chunxiao
Veröffentlicht: (2026)
Adaptive Instruction Composition for Automated LLM Red-Teaming
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
von: Zymet, Jesse, et al.
Veröffentlicht: (2026)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
von: Paulus, Anselm, et al.
Veröffentlicht: (2024)
von: Paulus, Anselm, et al.
Veröffentlicht: (2024)
Watermarking Counterfactual Explanations
von: Guo, Hangzhi, et al.
Veröffentlicht: (2024)
von: Guo, Hangzhi, et al.
Veröffentlicht: (2024)
LLM Cyber Evaluations Don't Capture Real-World Risk
von: Lukošiūtė, Kamilė, et al.
Veröffentlicht: (2025)
von: Lukošiūtė, Kamilė, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
von: Kim, Heegyu, et al.
Veröffentlicht: (2024) -
Slow is Fast! Dissecting Ethereum's Slow Liquidity Drain Scams
von: Tran, Minh Trung, et al.
Veröffentlicht: (2025) -
Persona-Model Collapse in Emergent Misalignment
von: Costa, Davi Bastos, et al.
Veröffentlicht: (2026) -
Confidence Elicitation: A New Attack Vector for Large Language Models
von: Formento, Brian, et al.
Veröffentlicht: (2025) -
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)