Autonomous Chain-of-Thought Distillation for Graph-Based Fraud Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yuan, Hu, Jun, Hooi, Bryan, He, Bingsheng, Chen, Cheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks
di: Xu, Naen, et al.
Pubblicazione: (2026)
di: Xu, Naen, et al.
Pubblicazione: (2026)
KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection
di: Li, Yuexin, et al.
Pubblicazione: (2024)
di: Li, Yuexin, et al.
Pubblicazione: (2024)
Privacy in Large Language Models: Attacks, Defenses and Future Directions
di: Li, Haoran, et al.
Pubblicazione: (2023)
di: Li, Haoran, et al.
Pubblicazione: (2023)
CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems
di: Wen, Yan, et al.
Pubblicazione: (2025)
di: Wen, Yan, et al.
Pubblicazione: (2025)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
di: You, Wenhao, et al.
Pubblicazione: (2025)
di: You, Wenhao, et al.
Pubblicazione: (2025)
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
Automated Phishing Detection Using URLs and Webpages
di: Wang, Huilin, et al.
Pubblicazione: (2024)
di: Wang, Huilin, et al.
Pubblicazione: (2024)
From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
di: Liu, Junhao, et al.
Pubblicazione: (2025)
di: Liu, Junhao, et al.
Pubblicazione: (2025)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
di: Wu, Tianyi, et al.
Pubblicazione: (2025)
di: Wu, Tianyi, et al.
Pubblicazione: (2025)
Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory
di: Li, Haoran, et al.
Pubblicazione: (2024)
di: Li, Haoran, et al.
Pubblicazione: (2024)
Argus: Reorchestrating Static Analysis via a Multi-Agent Ensemble for Full-Chain Security Vulnerability Detection
di: Liang, Zi, et al.
Pubblicazione: (2026)
di: Liang, Zi, et al.
Pubblicazione: (2026)
AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
di: Li, Yuexin, et al.
Pubblicazione: (2026)
di: Li, Yuexin, et al.
Pubblicazione: (2026)
"Yes, My LoRD." Guiding Language Model Extraction with Locality Reinforced Distillation
di: Liang, Zi, et al.
Pubblicazione: (2024)
di: Liang, Zi, et al.
Pubblicazione: (2024)
MBTSAD: Mitigating Backdoors in Language Models Based on Token Splitting and Attention Distillation
di: Ding, Yidong, et al.
Pubblicazione: (2025)
di: Ding, Yidong, et al.
Pubblicazione: (2025)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
di: Zhu, He, et al.
Pubblicazione: (2026)
di: Zhu, He, et al.
Pubblicazione: (2026)
GraphSteal: Structural Knowledge Stealing from Graph RAG via Traversal Reconstruction
di: Gu, Jinze, et al.
Pubblicazione: (2026)
di: Gu, Jinze, et al.
Pubblicazione: (2026)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
di: Li, Xiang, et al.
Pubblicazione: (2025)
di: Li, Xiang, et al.
Pubblicazione: (2025)
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
di: Rad, Melissa Kazemi, et al.
Pubblicazione: (2025)
di: Rad, Melissa Kazemi, et al.
Pubblicazione: (2025)
"I Strongly Suspect This Website Is a Scam": Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents
di: Roy, Soham, et al.
Pubblicazione: (2026)
di: Roy, Soham, et al.
Pubblicazione: (2026)
Self and Cross-Model Distillation for LLMs: Effective Methods for Refusal Pattern Alignment
di: Li, Jie, et al.
Pubblicazione: (2024)
di: Li, Jie, et al.
Pubblicazione: (2024)
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
di: Pisano, Matthew, et al.
Pubblicazione: (2023)
di: Pisano, Matthew, et al.
Pubblicazione: (2023)
Rethinking Backdoor Detection Evaluation for Language Models
di: Yan, Jun, et al.
Pubblicazione: (2024)
di: Yan, Jun, et al.
Pubblicazione: (2024)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
di: Liu, Kangwei, et al.
Pubblicazione: (2025)
di: Liu, Kangwei, et al.
Pubblicazione: (2025)
LLM-based Privacy Data Augmentation Guided by Knowledge Distillation with a Distribution Tutor for Medical Text Classification
di: Song, Yiping, et al.
Pubblicazione: (2024)
di: Song, Yiping, et al.
Pubblicazione: (2024)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
di: Batra, Shourya, et al.
Pubblicazione: (2025)
di: Batra, Shourya, et al.
Pubblicazione: (2025)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
CAPID: Context-Aware PII Detection for Question-Answering Systems
di: Ponomarenko, Mariia, et al.
Pubblicazione: (2026)
di: Ponomarenko, Mariia, et al.
Pubblicazione: (2026)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
di: Zhao, Tianzhe, et al.
Pubblicazione: (2025)
di: Zhao, Tianzhe, et al.
Pubblicazione: (2025)
DiDOTS: Knowledge Distillation from Large-Language-Models for Dementia Obfuscation in Transcribed Speech
di: Woszczyk, Dominika, et al.
Pubblicazione: (2024)
di: Woszczyk, Dominika, et al.
Pubblicazione: (2024)
Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method
di: Zhang, Weichao, et al.
Pubblicazione: (2024)
di: Zhang, Weichao, et al.
Pubblicazione: (2024)
Fingerprinting LLMs via Prompt Injection
di: Hu, Yuepeng, et al.
Pubblicazione: (2025)
di: Hu, Yuepeng, et al.
Pubblicazione: (2025)
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
Chain-of-Lure: A Universal Jailbreak Attack Framework using Unconstrained Synthetic Narratives
di: Chang, Wenhan, et al.
Pubblicazione: (2025)
di: Chang, Wenhan, et al.
Pubblicazione: (2025)
Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning
di: Chen, Chaoran, et al.
Pubblicazione: (2026)
di: Chen, Chaoran, et al.
Pubblicazione: (2026)
Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System
di: He, Haorui, et al.
Pubblicazione: (2025)
di: He, Haorui, et al.
Pubblicazione: (2025)
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
di: Cyberey, Hannah, et al.
Pubblicazione: (2025)
di: Cyberey, Hannah, et al.
Pubblicazione: (2025)
GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
di: Meng, Wenlong, et al.
Pubblicazione: (2025)
di: Meng, Wenlong, et al.
Pubblicazione: (2025)
Tag&Tab: Pretraining Data Detection in Large Language Models Using Keyword-Based Membership Inference Attack
di: Antebi, Sagiv, et al.
Pubblicazione: (2025)
di: Antebi, Sagiv, et al.
Pubblicazione: (2025)
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
di: He, Xuanli, et al.
Pubblicazione: (2024)
di: He, Xuanli, et al.
Pubblicazione: (2024)
Documenti analoghi
-
FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks
di: Xu, Naen, et al.
Pubblicazione: (2026) -
KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection
di: Li, Yuexin, et al.
Pubblicazione: (2024) -
Privacy in Large Language Models: Attacks, Defenses and Future Directions
di: Li, Haoran, et al.
Pubblicazione: (2023) -
CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems
di: Wen, Yan, et al.
Pubblicazione: (2025) -
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
di: You, Wenhao, et al.
Pubblicazione: (2025)