Enabling Contextual Soft Moderation on Social Media through Contrastive Textual Deviation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Paudel, Pujan, Saeed, Mohammad Hammas, Auger, Rebecca, Wells, Chris, Stringhini, Gianluca |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries
von: Paudel, Pujan, et al.
Veröffentlicht: (2025)
von: Paudel, Pujan, et al.
Veröffentlicht: (2025)
TUBERAIDER: Attributing Coordinated Hate Attacks on YouTube Videos to their Source Communities
von: Saeed, Mohammad Hammas, et al.
Veröffentlicht: (2023)
von: Saeed, Mohammad Hammas, et al.
Veröffentlicht: (2023)
Unraveling the Web of Disinformation: Exploring the Larger Context of State-Sponsored Influence Campaigns on Twitter
von: Saeed, Mohammad Hammas, et al.
Veröffentlicht: (2024)
von: Saeed, Mohammad Hammas, et al.
Veröffentlicht: (2024)
PIXELMOD: Improving Soft Moderation of Visual Misleading Information on Twitter
von: Paudel, Pujan, et al.
Veröffentlicht: (2024)
von: Paudel, Pujan, et al.
Veröffentlicht: (2024)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
Multi-Granularity Tibetan Textual Adversarial Attack Method Based on Masked Language Model
von: Cao, Xi, et al.
Veröffentlicht: (2024)
von: Cao, Xi, et al.
Veröffentlicht: (2024)
Adversarial Text Generation with Dynamic Contextual Perturbation
von: Waghela, Hetvi, et al.
Veröffentlicht: (2025)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2025)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
von: Zhu, He, et al.
Veröffentlicht: (2026)
von: Zhu, He, et al.
Veröffentlicht: (2026)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
von: German, Eyal, et al.
Veröffentlicht: (2025)
von: German, Eyal, et al.
Veröffentlicht: (2025)
Pay Attention to the Robustness of Chinese Minority Language Models! Syllable-level Textual Adversarial Attack on Tibetan Script
von: Cao, Xi, et al.
Veröffentlicht: (2024)
von: Cao, Xi, et al.
Veröffentlicht: (2024)
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
von: Béjar, Mario Rodríguez, et al.
Veröffentlicht: (2026)
von: Béjar, Mario Rodríguez, et al.
Veröffentlicht: (2026)
Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory
von: Li, Haoran, et al.
Veröffentlicht: (2024)
von: Li, Haoran, et al.
Veröffentlicht: (2024)
GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory
von: Fan, Wei, et al.
Veröffentlicht: (2024)
von: Fan, Wei, et al.
Veröffentlicht: (2024)
Reversible Jump Attack to Textual Classifiers with Modification Reduction
von: Ni, Mingze, et al.
Veröffentlicht: (2024)
von: Ni, Mingze, et al.
Veröffentlicht: (2024)
Claim-Guided Textual Backdoor Attack for Practical Applications
von: Song, Minkyoo, et al.
Veröffentlicht: (2024)
von: Song, Minkyoo, et al.
Veröffentlicht: (2024)
Multi-use LLM Watermarking and the False Detection Problem
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
von: An, Li, et al.
Veröffentlicht: (2025)
von: An, Li, et al.
Veröffentlicht: (2025)
The Double-edged Sword of LLM-based Data Reconstruction: Understanding and Mitigating Contextual Vulnerability in Word-level Differential Privacy Text Sanitization
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
Contextualized Privacy Defense for LLM Agents
von: Wen, Yule, et al.
Veröffentlicht: (2026)
von: Wen, Yule, et al.
Veröffentlicht: (2026)
FLAME: Flexible LLM-Assisted Moderation Engine
von: Bakulin, Ivan, et al.
Veröffentlicht: (2025)
von: Bakulin, Ivan, et al.
Veröffentlicht: (2025)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
von: Zhang, Fangyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Fangyuan, et al.
Veröffentlicht: (2024)
Publicly-Detectable Watermarking for Language Models
von: Fairoze, Jaiden, et al.
Veröffentlicht: (2023)
von: Fairoze, Jaiden, et al.
Veröffentlicht: (2023)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
von: Geng, Jiahui, et al.
Veröffentlicht: (2025)
von: Geng, Jiahui, et al.
Veröffentlicht: (2025)
Contextual Agent Security: A Policy for Every Purpose
von: Tsai, Lillian, et al.
Veröffentlicht: (2025)
von: Tsai, Lillian, et al.
Veröffentlicht: (2025)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
von: Das, Saswat, et al.
Veröffentlicht: (2025)
von: Das, Saswat, et al.
Veröffentlicht: (2025)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
Adversarial Robustness through Dynamic Ensemble Learning
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
von: Mirbagheri, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Mirbagheri, Mohammad Reza, et al.
Veröffentlicht: (2025)
A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
von: Xin, Rui, et al.
Veröffentlicht: (2025)
von: Xin, Rui, et al.
Veröffentlicht: (2025)
A Curious Case of Searching for the Correlation between Training Data and Adversarial Robustness of Transformer Textual Models
von: Dang, Cuong, et al.
Veröffentlicht: (2024)
von: Dang, Cuong, et al.
Veröffentlicht: (2024)
Say Something Else: Rethinking Contextual Privacy as Information Sufficiency
von: Xiao, Yunze, et al.
Veröffentlicht: (2026)
von: Xiao, Yunze, et al.
Veröffentlicht: (2026)
Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID
von: Han, Xia, et al.
Veröffentlicht: (2025)
von: Han, Xia, et al.
Veröffentlicht: (2025)
From Theory to Practice: Evaluating Data Poisoning Attacks and Defenses in In-Context Learning on Social Media Health Discourse
von: Jhuma, Rabeya Amin, et al.
Veröffentlicht: (2025)
von: Jhuma, Rabeya Amin, et al.
Veröffentlicht: (2025)
LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
von: Ullah, Saad, et al.
Veröffentlicht: (2023)
von: Ullah, Saad, et al.
Veröffentlicht: (2023)
Enhance Robustness of Language Models Against Variation Attack through Graph Integration
von: Xiong, Zi, et al.
Veröffentlicht: (2024)
von: Xiong, Zi, et al.
Veröffentlicht: (2024)
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LOKI: Proactively Discovering Online Scam Websites by Mining Toxic Search Queries
von: Paudel, Pujan, et al.
Veröffentlicht: (2025) -
TUBERAIDER: Attributing Coordinated Hate Attacks on YouTube Videos to their Source Communities
von: Saeed, Mohammad Hammas, et al.
Veröffentlicht: (2023) -
Unraveling the Web of Disinformation: Exploring the Larger Context of State-Sponsored Influence Campaigns on Twitter
von: Saeed, Mohammad Hammas, et al.
Veröffentlicht: (2024) -
PIXELMOD: Improving Soft Moderation of Visual Misleading Information on Twitter
von: Paudel, Pujan, et al.
Veröffentlicht: (2024) -
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)