Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Zhaomin, Du, Mingzhe, Ng, See-Kiong, He, Bingsheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prompt Optimization with Human Feedback
by: Lin, Xiaoqiang, et al.
Published: (2024)
by: Lin, Xiaoqiang, et al.
Published: (2024)
EmoTrack: Robust Depression Tracking from Counseling Transcripts across Session Regimes
by: Wu, Zhaomin, et al.
Published: (2026)
by: Wu, Zhaomin, et al.
Published: (2026)
VertiBench: Advancing Feature Distribution Diversity in Vertical Federated Learning Benchmarks
by: Wu, Zhaomin, et al.
Published: (2023)
by: Wu, Zhaomin, et al.
Published: (2023)
Learning Relational Tabular Data without Shared Features
by: Wu, Zhaomin, et al.
Published: (2025)
by: Wu, Zhaomin, et al.
Published: (2025)
LLM DNA: Tracing Model Evolution via Functional Representations
by: Wu, Zhaomin, et al.
Published: (2025)
by: Wu, Zhaomin, et al.
Published: (2025)
Federated Data-Efficient Instruction Tuning for Large Language Models
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Federated Transformer: Multi-Party Vertical Federated Learning on Practical Fuzzily Linked Data
by: Wu, Zhaomin, et al.
Published: (2024)
by: Wu, Zhaomin, et al.
Published: (2024)
Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars
by: Wu, Zhaoxuan, et al.
Published: (2024)
by: Wu, Zhaoxuan, et al.
Published: (2024)
Rating Quality of Diverse Time Series Data by Meta-learning from LLM Judgment
by: Wu, Shunyu, et al.
Published: (2025)
by: Wu, Shunyu, et al.
Published: (2025)
Low Resource Reconstruction Attacks Through Benign Prompts
by: Yarkoni, Sol, et al.
Published: (2025)
by: Yarkoni, Sol, et al.
Published: (2025)
FD-LLM: Large Language Model for Fault Diagnosis of Machines
by: Qaid, Hamzah A. A. M., et al.
Published: (2024)
by: Qaid, Hamzah A. A. M., et al.
Published: (2024)
Dipper: Diversity in Prompts for Producing Large Language Model Ensembles in Reasoning tasks
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
by: Zhao, James Xu, et al.
Published: (2025)
by: Zhao, James Xu, et al.
Published: (2025)
DGP: A Dual-Granularity Prompting Framework for Fraud Detection with Graph-Enhanced LLMs
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
SafeRedir: Prompt Embedding Redirection for Robust Unlearning in Image Generation Models
by: Liu, Renyang, et al.
Published: (2026)
by: Liu, Renyang, et al.
Published: (2026)
Multi-Modal One-Shot Federated Ensemble Learning for Medical Data with Vision Large Language Model
by: Wang, Naibo, et al.
Published: (2025)
by: Wang, Naibo, et al.
Published: (2025)
Vertical Federated Learning in Practice: The Good, the Bad, and the Ugly
by: Wu, Zhaomin, et al.
Published: (2025)
by: Wu, Zhaomin, et al.
Published: (2025)
Model-based Large Language Model Customization as Service
by: Wu, Zhaomin, et al.
Published: (2024)
by: Wu, Zhaomin, et al.
Published: (2024)
BACE-RUL: A Bi-directional Adversarial Network with Covariate Encoding for Machine Remaining Useful Life Prediction
by: Zhang, Zekai, et al.
Published: (2025)
by: Zhang, Zekai, et al.
Published: (2025)
Lightweight Time Series Data Valuation on Time Series Foundation Models via In-Context Finetuning
by: Wu, Shunyu, et al.
Published: (2025)
by: Wu, Shunyu, et al.
Published: (2025)
Ferret: Federated Full-Parameter Tuning at Scale for Large Language Models
by: Shu, Yao, et al.
Published: (2024)
by: Shu, Yao, et al.
Published: (2024)
PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection
by: Camarato, Steffen J., et al.
Published: (2026)
by: Camarato, Steffen J., et al.
Published: (2026)
How Does Response Length Affect Long-Form Factuality
by: Zhao, James Xu, et al.
Published: (2025)
by: Zhao, James Xu, et al.
Published: (2025)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
On Newton's Method to Unlearn Neural Networks
by: Bui, Nhung, et al.
Published: (2024)
by: Bui, Nhung, et al.
Published: (2024)
Auto-Prompt Ensemble for LLM Judge
by: Li, Jiajie, et al.
Published: (2025)
by: Li, Jiajie, et al.
Published: (2025)
GLOW: Graph-Language Co-Reasoning for Agentic Workflow Performance Prediction
by: Guan, Wei, et al.
Published: (2025)
by: Guan, Wei, et al.
Published: (2025)
Integrating Time Series into LLMs via Multi-layer Steerable Embedding Fusion for Enhanced Forecasting
by: Chen, Zhuomin, et al.
Published: (2025)
by: Chen, Zhuomin, et al.
Published: (2025)
Beyond Prompt: Fine-grained Simulation of Cognitively Impaired Standardized Patients via Stochastic Steering
by: Zhang, Weikang, et al.
Published: (2026)
by: Zhang, Weikang, et al.
Published: (2026)
WaveMoE: A Wavelet-Enhanced Mixture-of-Experts Foundation Model for Time Series Forecasting
by: Wu, Shunyu, et al.
Published: (2026)
by: Wu, Shunyu, et al.
Published: (2026)
Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
by: Guo, Jizhou, et al.
Published: (2025)
by: Guo, Jizhou, et al.
Published: (2025)
Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference
by: Adamska, Marta, et al.
Published: (2025)
by: Adamska, Marta, et al.
Published: (2025)
Source Attribution for Large Language Model-Generated Data
by: Wang, Jingtan, et al.
Published: (2023)
by: Wang, Jingtan, et al.
Published: (2023)
Operationalizing Data Minimization for Privacy-Preserving LLM Prompting
by: Zhou, Jijie, et al.
Published: (2025)
by: Zhou, Jijie, et al.
Published: (2025)
Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs
by: Ashok, Arjun, et al.
Published: (2025)
by: Ashok, Arjun, et al.
Published: (2025)
From Prompts to Power: Measuring the Energy Footprint of LLM Inference
by: Caravaca, Francisco, et al.
Published: (2025)
by: Caravaca, Francisco, et al.
Published: (2025)
Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems
by: Ross, Brendan Leigh, et al.
Published: (2025)
by: Ross, Brendan Leigh, et al.
Published: (2025)
Edge Prompt Tuning for Graph Neural Networks
by: Fu, Xingbo, et al.
Published: (2025)
by: Fu, Xingbo, et al.
Published: (2025)
How Not to Detect Prompt Injections with an LLM
by: Choudhary, Sarthak, et al.
Published: (2025)
by: Choudhary, Sarthak, et al.
Published: (2025)
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
by: Kumarappan, Adarsh, et al.
Published: (2025)
by: Kumarappan, Adarsh, et al.
Published: (2025)
Similar Items
-
Prompt Optimization with Human Feedback
by: Lin, Xiaoqiang, et al.
Published: (2024) -
EmoTrack: Robust Depression Tracking from Counseling Transcripts across Session Regimes
by: Wu, Zhaomin, et al.
Published: (2026) -
VertiBench: Advancing Feature Distribution Diversity in Vertical Federated Learning Benchmarks
by: Wu, Zhaomin, et al.
Published: (2023) -
Learning Relational Tabular Data without Shared Features
by: Wu, Zhaomin, et al.
Published: (2025) -
LLM DNA: Tracing Model Evolution via Functional Representations
by: Wu, Zhaomin, et al.
Published: (2025)