Gespeichert in:
| Hauptverfasser: | Nirman, Diana Bar-Or, Weizman, Ariel, Azaria, Amos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2412.11625 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
von: Ezra, Elon, et al.
Veröffentlicht: (2025)
von: Ezra, Elon, et al.
Veröffentlicht: (2025)
Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support Setting
von: Kayser, Maxime, et al.
Veröffentlicht: (2024)
von: Kayser, Maxime, et al.
Veröffentlicht: (2024)
One Token to Fool LLM-as-a-Judge
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
TALL -- A Trainable Architecture for Enhancing LLM Performance in Low-Resource Languages
von: Ofer, Moshe, et al.
Veröffentlicht: (2025)
von: Ofer, Moshe, et al.
Veröffentlicht: (2025)
RAFT: Realistic Attacks to Fool Text Detectors
von: Wang, James, et al.
Veröffentlicht: (2024)
von: Wang, James, et al.
Veröffentlicht: (2024)
Fooling the Textual Fooler via Randomizing Latent Representations
von: Hoang, Duy C., et al.
Veröffentlicht: (2023)
von: Hoang, Duy C., et al.
Veröffentlicht: (2023)
Language Model Re-rankers are Fooled by Lexical Similarities
von: Hagström, Lovisa, et al.
Veröffentlicht: (2025)
von: Hagström, Lovisa, et al.
Veröffentlicht: (2025)
Don't be a Fool: Pooling Strategies in Offensive Language Detection from User-Intended Adversarial Attacks
von: Yu, Seunguk, et al.
Veröffentlicht: (2024)
von: Yu, Seunguk, et al.
Veröffentlicht: (2024)
Too Big to Fool: Resisting Deception in Language Models
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
von: Mukhopadhyay, Souradeep, et al.
Veröffentlicht: (2025)
von: Mukhopadhyay, Souradeep, et al.
Veröffentlicht: (2025)
Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
von: Lermen, Simon, et al.
Veröffentlicht: (2025)
von: Lermen, Simon, et al.
Veröffentlicht: (2025)
Ruffle&Riley: Insights from Designing and Evaluating a Large Language Model-Based Conversational Tutoring System
von: Schmucker, Robin, et al.
Veröffentlicht: (2024)
von: Schmucker, Robin, et al.
Veröffentlicht: (2024)
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
von: Yamamura, Atsushi, et al.
Veröffentlicht: (2024)
von: Yamamura, Atsushi, et al.
Veröffentlicht: (2024)
The State Of TTS: A Case Study with Human Fooling Rates
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
von: Qian, Yusu, et al.
Veröffentlicht: (2024)
von: Qian, Yusu, et al.
Veröffentlicht: (2024)
Dagger Behind Smile: Fool LLMs with a Happy Ending Story
von: Song, Xurui, et al.
Veröffentlicht: (2025)
von: Song, Xurui, et al.
Veröffentlicht: (2025)
Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack
von: Ding, Peng, et al.
Veröffentlicht: (2025)
von: Ding, Peng, et al.
Veröffentlicht: (2025)
Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
von: Jiang, Bowen, et al.
Veröffentlicht: (2025)
von: Jiang, Bowen, et al.
Veröffentlicht: (2025)
StyleFool: Fooling Video Classification Systems via Style Transfer
von: Cao, Yuxin, et al.
Veröffentlicht: (2022)
von: Cao, Yuxin, et al.
Veröffentlicht: (2022)
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
von: Yao, Yang, et al.
Veröffentlicht: (2025)
von: Yao, Yang, et al.
Veröffentlicht: (2025)
FeatureFool: Zero-Query Fooling of Video Models via Feature Map
von: Tang, Duoxun, et al.
Veröffentlicht: (2025)
von: Tang, Duoxun, et al.
Veröffentlicht: (2025)
Don't Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning
von: Ravie, Navin Sriram, et al.
Veröffentlicht: (2026)
von: Ravie, Navin Sriram, et al.
Veröffentlicht: (2026)
Fool Me Once: A Case of Recurrent Delirium in the Setting of Buprenorphine Use
von: Yasmeen Abdo, et al.
Veröffentlicht: (2025)
von: Yasmeen Abdo, et al.
Veröffentlicht: (2025)
A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
von: Ding, Peng, et al.
Veröffentlicht: (2023)
von: Ding, Peng, et al.
Veröffentlicht: (2023)
Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors
von: Pedrotti, Andrea, et al.
Veröffentlicht: (2025)
von: Pedrotti, Andrea, et al.
Veröffentlicht: (2025)
Language Models Optimized to Fool Detectors Still Have a Distinct Style (And How to Change It)
von: Soto, Rafael Rivera, et al.
Veröffentlicht: (2025)
von: Soto, Rafael Rivera, et al.
Veröffentlicht: (2025)
Unmasking Digital Falsehoods: A Comparative Analysis of LLM-Based Misinformation Detection Strategies
von: Huang, Tianyi, et al.
Veröffentlicht: (2025)
von: Huang, Tianyi, et al.
Veröffentlicht: (2025)
Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
von: Saiem, Bijoy Ahmed, et al.
Veröffentlicht: (2024)
The Psychology of Falsehood: A Human-Centric Survey of Misinformation Detection
von: Nandi, Arghodeep, et al.
Veröffentlicht: (2025)
von: Nandi, Arghodeep, et al.
Veröffentlicht: (2025)
Poor Fools
von: Daily, Jay E.
Veröffentlicht: (1972)
von: Daily, Jay E.
Veröffentlicht: (1972)
"This Wasn't Made for Me": Recentering User Experience and Emotional Impact in the Evaluation of ASR Bias
von: Liang, Siyu, et al.
Veröffentlicht: (2026)
von: Liang, Siyu, et al.
Veröffentlicht: (2026)
Hear Me, See Me, Understand Me: Audio-Visual Autism Behavior Recognition
von: Deng, Shijian, et al.
Veröffentlicht: (2024)
von: Deng, Shijian, et al.
Veröffentlicht: (2024)
The Energy of Falsehood: Detecting Hallucinations via Diffusion Model Likelihoods
von: Gautam, Arpit Singh, et al.
Veröffentlicht: (2026)
von: Gautam, Arpit Singh, et al.
Veröffentlicht: (2026)
Use Me Wisely: AI-Driven Assessment for LLM Prompting Skills Development
von: Ognibene, Dimitri, et al.
Veröffentlicht: (2025)
von: Ognibene, Dimitri, et al.
Veröffentlicht: (2025)
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals
von: Wu, Yue, et al.
Veröffentlicht: (2023)
von: Wu, Yue, et al.
Veröffentlicht: (2023)
Me-Agent: A Personalized Mobile Agent with Two-Level User Habit Learning for Enhanced Interaction
von: Wang, Shuoxin, et al.
Veröffentlicht: (2026)
von: Wang, Shuoxin, et al.
Veröffentlicht: (2026)
Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution
von: Cao, Zouying, et al.
Veröffentlicht: (2025)
von: Cao, Zouying, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
von: Ezra, Elon, et al.
Veröffentlicht: (2025) -
Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support Setting
von: Kayser, Maxime, et al.
Veröffentlicht: (2024) -
One Token to Fool LLM-as-a-Judge
von: Zhao, Yulai, et al.
Veröffentlicht: (2025) -
TALL -- A Trainable Architecture for Enhancing LLM Performance in Low-Resource Languages
von: Ofer, Moshe, et al.
Veröffentlicht: (2025) -
RAFT: Realistic Attacks to Fool Text Detectors
von: Wang, James, et al.
Veröffentlicht: (2024)