Saved in:
| Main Authors: | Li, Zongjie, Qiu, Wenying, Ma, Pingchuan, Li, Yichen, Li, You, He, Sijia, Jiang, Baozheng, Wang, Shuai, Gu, Weixi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.01723 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models
by: Wang, Xunguang, et al.
Published: (2023)
by: Wang, Xunguang, et al.
Published: (2023)
Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges
by: Ji, Zimo, et al.
Published: (2025)
by: Ji, Zimo, et al.
Published: (2025)
Testing and Understanding Erroneous Planning in LLM Agents through Synthesized User Inputs
by: Ji, Zhenlan, et al.
Published: (2024)
by: Ji, Zhenlan, et al.
Published: (2024)
WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making
by: Li, Zongjie, et al.
Published: (2026)
by: Li, Zongjie, et al.
Published: (2026)
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
by: Wang, Xunguang, et al.
Published: (2025)
by: Wang, Xunguang, et al.
Published: (2025)
SkillReducer: Optimizing LLM Agent Skills for Token Efficiency
by: Gao, Yudong, et al.
Published: (2026)
by: Gao, Yudong, et al.
Published: (2026)
Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models
by: Wang, Xunguang, et al.
Published: (2026)
by: Wang, Xunguang, et al.
Published: (2026)
Empirical Study of Code Large Language Models for Binary Security Patch Detection
by: Li, Qingyuan, et al.
Published: (2025)
by: Li, Qingyuan, et al.
Published: (2025)
How Multi-Modal LLMs Reshape Visual Deep Learning Testing? A Comprehensive Study Through the Lens of Image Mutation
by: Wang, Liwen, et al.
Published: (2024)
by: Wang, Liwen, et al.
Published: (2024)
Taming Various Privilege Escalation in LLM-Based Agent Systems: A Mandatory Access Control Framework
by: Ji, Zimo, et al.
Published: (2026)
by: Ji, Zimo, et al.
Published: (2026)
OpenVNA: A Framework for Analyzing the Behavior of Multimodal Language Understanding System under Noisy Scenarios
by: Yuan, Ziqi, et al.
Published: (2024)
by: Yuan, Ziqi, et al.
Published: (2024)
Split and Merge: Aligning Position Biases in LLM-based Evaluators
by: Li, Zongjie, et al.
Published: (2023)
by: Li, Zongjie, et al.
Published: (2023)
Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models
by: Li, Zongjie, et al.
Published: (2025)
by: Li, Zongjie, et al.
Published: (2025)
EAMET: Robust Massive Model Editing via Embedding Alignment Optimization
by: Dai, Yanbo, et al.
Published: (2025)
by: Dai, Yanbo, et al.
Published: (2025)
API-guided Dataset Synthesis to Finetune Large Code Models
by: Li, Zongjie, et al.
Published: (2024)
by: Li, Zongjie, et al.
Published: (2024)
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
by: Ji, Zimo, et al.
Published: (2025)
by: Ji, Zimo, et al.
Published: (2025)
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
by: Wang, Xunguang, et al.
Published: (2024)
by: Wang, Xunguang, et al.
Published: (2024)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
by: Wang, Xunguang, et al.
Published: (2025)
by: Wang, Xunguang, et al.
Published: (2025)
Do Large Language Models have Problem-Solving Capability under Incomplete Information Scenarios?
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Toward Intelligent Electronic-Photonic Design Automation for Large-Scale Photonic Integrated Circuits: from Device Inverse Design to Physical Layout Generation
by: Zhou, Hongjian, et al.
Published: (2025)
by: Zhou, Hongjian, et al.
Published: (2025)
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
by: Dang, Yunkai, et al.
Published: (2024)
by: Dang, Yunkai, et al.
Published: (2024)
Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
by: Ji, Zimo, et al.
Published: (2026)
by: Ji, Zimo, et al.
Published: (2026)
Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMs
by: Li, Zongjie, et al.
Published: (2025)
by: Li, Zongjie, et al.
Published: (2025)
SEAL: Subspace-Anchored Watermarks for LLM Ownership
by: Dai, Yanbo, et al.
Published: (2025)
by: Dai, Yanbo, et al.
Published: (2025)
Beyond Accuracy: An Empirical Study on Unit Testing in Open-source Deep Learning Projects
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Disabling Self-Correction in Retrieval-Augmented Generation via Stealthy Retriever Poisoning
by: Dai, Yanbo, et al.
Published: (2025)
by: Dai, Yanbo, et al.
Published: (2025)
Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical Study
by: Hou, Guanyu, et al.
Published: (2025)
by: Hou, Guanyu, et al.
Published: (2025)
Design and Empirical Study of a Large Language Model-Based Multi-Agent Investment System for Chinese Public REITs
by: Li, Zheng
Published: (2026)
by: Li, Zheng
Published: (2026)
Network simulation tools for unmanned aerial vehicle communications: A survey
by: Weiwei Jiang, et al.
Published: (2024)
by: Weiwei Jiang, et al.
Published: (2024)
Toward Adaptive Reasoning in Large Language Models with Thought Rollback
by: Chen, Sijia, et al.
Published: (2024)
by: Chen, Sijia, et al.
Published: (2024)
From Evaluation to Enhancement: Large Language Models for Zero-Knowledge Proof Code Generation
by: Xue, Zhantong, et al.
Published: (2025)
by: Xue, Zhantong, et al.
Published: (2025)
TianHui: A Domain-Specific Large Language Model for Diverse Traditional Chinese Medicine Scenarios
by: Yin, Ji, et al.
Published: (2025)
by: Yin, Ji, et al.
Published: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
CodeJudge: Evaluating Code Generation with Large Language Models
by: Tong, Weixi, et al.
Published: (2024)
by: Tong, Weixi, et al.
Published: (2024)
CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios
by: Ouyang, Zetian, et al.
Published: (2024)
by: Ouyang, Zetian, et al.
Published: (2024)
Eliminating Information Leakage in Hard Concept Bottleneck Models with Supervised, Hierarchical Concept Learning
by: Sun, Ao, et al.
Published: (2024)
by: Sun, Ao, et al.
Published: (2024)
Efficient Differentiable Causal Discovery via Reliable Super-Structure Learning
by: Ma, Pingchuan, et al.
Published: (2026)
by: Ma, Pingchuan, et al.
Published: (2026)
Feature engineering vs. deep learning for paper section identification: Toward applications in Chinese medical literature
by: Zhou, Sijia, et al.
Published: (2024)
by: Zhou, Sijia, et al.
Published: (2024)
Exploring the Effectiveness of LLMs in Automated Logging Generation: An Empirical Study
by: Li, Yichen, et al.
Published: (2023)
by: Li, Yichen, et al.
Published: (2023)
ADEPT-Z: Zero-Shot Automated Circuit Topology Search for Pareto-Optimal Photonic Tensor Cores
by: Jiang, Ziyang, et al.
Published: (2024)
by: Jiang, Ziyang, et al.
Published: (2024)
Similar Items
-
InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models
by: Wang, Xunguang, et al.
Published: (2023) -
Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges
by: Ji, Zimo, et al.
Published: (2025) -
Testing and Understanding Erroneous Planning in LLM Agents through Synthesized User Inputs
by: Ji, Zhenlan, et al.
Published: (2024) -
WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making
by: Li, Zongjie, et al.
Published: (2026) -
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
by: Wang, Xunguang, et al.
Published: (2025)