Saved in:
| Main Authors: | Qi, Peigui, Tang, Kunsheng, Yu, Yanpu, Wu, Jialin, Song, Yide, Zhou, Wenbo, Huang, Zhicong, Hong, Cheng, Zhang, Weiming, Yu, Nenghai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.06502 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
by: Qi, Peigui, et al.
Published: (2025)
by: Qi, Peigui, et al.
Published: (2025)
Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
by: Wu, Jialin, et al.
Published: (2026)
by: Wu, Jialin, et al.
Published: (2026)
GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language Models
by: Tang, Kunsheng, et al.
Published: (2024)
by: Tang, Kunsheng, et al.
Published: (2024)
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Membership Inference Attacks against Large Audio Language Models
by: Dong, Jia-Kai, et al.
Published: (2026)
by: Dong, Jia-Kai, et al.
Published: (2026)
SP-Guard: Selective Prompt-adaptive Guidance for Safe Text-to-Image Generation
by: Yu, Sumin, et al.
Published: (2025)
by: Yu, Sumin, et al.
Published: (2025)
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025)
by: Wang, Kongxin, et al.
Published: (2025)
Knowledge Distillation in RNN-Attention Models for Early Prediction of Student Performance
by: Leelaluk, Sukrit, et al.
Published: (2024)
by: Leelaluk, Sukrit, et al.
Published: (2024)
Capture-Calibrate-Coach: A Graph-Based Framework for Knowledge Monitoring Estimation and Adaptive Feedback
by: Li, Gen, et al.
Published: (2026)
by: Li, Gen, et al.
Published: (2026)
TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion Sensitivity
by: Chen, Yuzhuo, et al.
Published: (2025)
by: Chen, Yuzhuo, et al.
Published: (2025)
Adaptive Defense Orchestration for RAG: A Sentinel-Strategist Architecture against Multi-Vector Attacks
by: Pallerla, Pranav, et al.
Published: (2026)
by: Pallerla, Pranav, et al.
Published: (2026)
Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
LECTOR: Summarizing E-book Reading Content for Personalized Student Support
by: Zapata, Erwin Daniel López, et al.
Published: (2025)
by: Zapata, Erwin Daniel López, et al.
Published: (2025)
VoiceSHIELD-Small: Real-Time Malicious Speech Detection and Transcription
by: Ranjan, Sumit, et al.
Published: (2026)
by: Ranjan, Sumit, et al.
Published: (2026)
RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code
by: Chen, Jiachi, et al.
Published: (2024)
by: Chen, Jiachi, et al.
Published: (2024)
Single-Agent vs. Multi-Agent LLM Strategies for Automated Student Reflection Assessment
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
by: Wang, Yanshu, et al.
Published: (2026)
by: Wang, Yanshu, et al.
Published: (2026)
In Defense of the Turing Test and its Legacy
by: Gonçalves, Bernardo
Published: (2025)
by: Gonçalves, Bernardo
Published: (2025)
Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy Games
by: Qi, Runnan, et al.
Published: (2025)
by: Qi, Runnan, et al.
Published: (2025)
DeRAG: Black-box Adversarial Attacks on Multiple Retrieval-Augmented Generation Applications via Prompt Injection
by: Wang, Jerry, et al.
Published: (2025)
by: Wang, Jerry, et al.
Published: (2025)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
by: Chona, Alankrit, et al.
Published: (2026)
by: Chona, Alankrit, et al.
Published: (2026)
Architecture-Agnostic Feature Synergy for Universal Defense Against Heterogeneous Generative Threats
by: Zhang, Bingxue, et al.
Published: (2026)
by: Zhang, Bingxue, et al.
Published: (2026)
GOT-JEPA: Generic Object Tracking with Model Adaptation and Occlusion Handling using Joint-Embedding Predictive Architecture
by: Chen, Shih-Fang, et al.
Published: (2026)
by: Chen, Shih-Fang, et al.
Published: (2026)
GOT-Edit: Geometry-Aware Generic Object Tracking via Online Model Editing
by: Chen, Shih-Fang, et al.
Published: (2026)
by: Chen, Shih-Fang, et al.
Published: (2026)
Whose wife is it anyway? Assessing bias against same-gender relationships in machine translation
by: Stewart, Ian, et al.
Published: (2024)
by: Stewart, Ian, et al.
Published: (2024)
Simulated Human Learning in a Dynamic, Partially-Observed, Time-Series Environment
by: Jiang, Jeffrey, et al.
Published: (2025)
by: Jiang, Jeffrey, et al.
Published: (2025)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
by: Chakraborty, Amit, et al.
Published: (2025)
by: Chakraborty, Amit, et al.
Published: (2025)
Robust Uncertainty Quantification for Factual Generation of Large Language Models
by: Zhang, Yuhao, et al.
Published: (2026)
by: Zhang, Yuhao, et al.
Published: (2026)
On the Evolution of A.I. and Machine Learning: Towards a Meta-level Measuring and Understanding Impact, Influence, and Leadership at Premier A.I. Conferences
by: Audibert, Rafael B., et al.
Published: (2022)
by: Audibert, Rafael B., et al.
Published: (2022)
Enhancing Computer Programming Education with LLMs: A Study on Effective Prompt Engineering for Python Code Generation
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study
by: Xu, Luyao, et al.
Published: (2026)
by: Xu, Luyao, et al.
Published: (2026)
From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring
by: Nguyen, Minh Hoang, et al.
Published: (2026)
by: Nguyen, Minh Hoang, et al.
Published: (2026)
Why we need an AI-resilient society
by: Bartz-Beielstein, Thomas
Published: (2019)
by: Bartz-Beielstein, Thomas
Published: (2019)
Standardization of Post-Publication Code Verification by Journals is Possible with the Support of the Community
by: Lopez-Moreno, Susana, et al.
Published: (2026)
by: Lopez-Moreno, Susana, et al.
Published: (2026)
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
by: Yu, Hao, et al.
Published: (2026)
by: Yu, Hao, et al.
Published: (2026)
Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
by: Wang, Haochuan Kevin, et al.
Published: (2026)
by: Wang, Haochuan Kevin, et al.
Published: (2026)
Detecting Inpainted Video with Frequency Domain Insights
by: Tang, Quanhui, et al.
Published: (2024)
by: Tang, Quanhui, et al.
Published: (2024)
Closing the SNAP Gap: Identifying Under-Enrollment in High-Poverty ZIP Codes
by: Ray, Auyona
Published: (2025)
by: Ray, Auyona
Published: (2025)
Making High-Level AI Design Decisions Explicit Using a Binary Stream System-Designation Approach
by: Mossbridge, Julia
Published: (2024)
by: Mossbridge, Julia
Published: (2024)
Similar Items
-
SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
by: Qi, Peigui, et al.
Published: (2025) -
Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
by: Wu, Jialin, et al.
Published: (2026) -
GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language Models
by: Tang, Kunsheng, et al.
Published: (2024) -
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts
by: Young, Richard J., et al.
Published: (2026) -
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)