Saved in:
| Main Authors: | Kaneko, Masahiro, Niwa, Ayana, Baldwin, Timothy |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.01291 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
by: Niwa, Ayana, et al.
Published: (2025)
by: Niwa, Ayana, et al.
Published: (2025)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
by: Kaneko, Masahiro
Published: (2026)
by: Kaneko, Masahiro
Published: (2026)
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
by: Koike, Ryuto, et al.
Published: (2025)
by: Koike, Ryuto, et al.
Published: (2025)
JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs
by: Shiotani, Taihei, et al.
Published: (2026)
by: Shiotani, Taihei, et al.
Published: (2026)
A Little Leak Will Sink a Great Ship: Survey of Transparency for Large Language Models from Start to Finish
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
Multi-view autoencoders for Fake News Detection
by: Pereira, Ingryd V. S. T., et al.
Published: (2025)
by: Pereira, Ingryd V. S. T., et al.
Published: (2025)
A Multi-Label Dataset of French Fake News: Human and Machine Insights
by: Icard, Benjamin, et al.
Published: (2024)
by: Icard, Benjamin, et al.
Published: (2024)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Generalization Gaps in Political Fake News Detection: An Empirical Study on the LIAR Dataset
by: Hasan, S Mahmudul, et al.
Published: (2025)
by: Hasan, S Mahmudul, et al.
Published: (2025)
Detection of Human and Machine-Authored Fake News in Urdu
by: Ali, Muhammad Zain, et al.
Published: (2024)
by: Ali, Muhammad Zain, et al.
Published: (2024)
Beyond the Resumé: A Rubric-Aware Automatic Interview System for Information Elicitation
by: Stuart, Harry, et al.
Published: (2026)
by: Stuart, Harry, et al.
Published: (2026)
The Gaps between Pre-train and Downstream Settings in Bias Evaluation and Debiasing
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
Eagle: Ethical Dataset Given from Real Interactions
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
Benchmarking Gender and Political Bias in Large Language Models
by: Yang, Jinrui, et al.
Published: (2025)
by: Yang, Jinrui, et al.
Published: (2025)
A Self-Learning Multimodal Approach for Fake News Detection
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
by: Niwa, Ayana, et al.
Published: (2024)
by: Niwa, Ayana, et al.
Published: (2024)
A Regularized LSTM Method for Detecting Fake News Articles
by: Camelia, Tanjina Sultana, et al.
Published: (2024)
by: Camelia, Tanjina Sultana, et al.
Published: (2024)
Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
by: Wei, Zhipeng, et al.
Published: (2024)
by: Wei, Zhipeng, et al.
Published: (2024)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
by: Fang, Zhicheng, et al.
Published: (2026)
by: Fang, Zhicheng, et al.
Published: (2026)
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
by: Cheng, Yixin, et al.
Published: (2024)
by: Cheng, Yixin, et al.
Published: (2024)
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
Bangla Fake News Detection Based On Multichannel Combined CNN-LSTM
by: George, Md. Zahin Hossain, et al.
Published: (2025)
by: George, Md. Zahin Hossain, et al.
Published: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
by: Hu, Hanjiang, et al.
Published: (2025)
by: Hu, Hanjiang, et al.
Published: (2025)
A New Method for Cross-Lingual-based Semantic Role Labeling
by: Ebrahimi, Mohammad, et al.
Published: (2024)
by: Ebrahimi, Mohammad, et al.
Published: (2024)
FaKnow: A Unified Library for Fake News Detection
by: Zhu, Yiyuan, et al.
Published: (2024)
by: Zhu, Yiyuan, et al.
Published: (2024)
Don't Ignore the Tail: Decoupling top-K Probabilities for Efficient Language Model Distillation
by: Dasgupta, Sayantan, et al.
Published: (2026)
by: Dasgupta, Sayantan, et al.
Published: (2026)
Advanced Text Analytics -- Graph Neural Network for Fake News Detection in Social Media
by: Patel, Anantram, et al.
Published: (2025)
by: Patel, Anantram, et al.
Published: (2025)
Employing Sentence Space Embedding for Classification of Data Stream from Fake News Domain
by: Zyblewski, Paweł, et al.
Published: (2024)
by: Zyblewski, Paweł, et al.
Published: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
Neighborhood-Order Learning Graph Attention Network for Fake News Detection
by: Lakzaei, Batool, et al.
Published: (2025)
by: Lakzaei, Batool, et al.
Published: (2025)
HSFN: Hierarchical Selection for Fake News Detection building Heterogeneous Ensemble
by: Coutinho, Sara B., et al.
Published: (2025)
by: Coutinho, Sara B., et al.
Published: (2025)
Cross-Modal Augmentation for Few-Shot Multimodal Fake News Detection
by: Jiang, Ye, et al.
Published: (2024)
by: Jiang, Ye, et al.
Published: (2024)
Multi-EuP: The Multilingual European Parliament Dataset for Analysis of Bias in Information Retrieval
by: Yang, Jinrui, et al.
Published: (2023)
by: Yang, Jinrui, et al.
Published: (2023)
Enhancing Bangla Fake News Detection Using Bidirectional Gated Recurrent Units and Deep Learning Techniques
by: Roy, Utsha, et al.
Published: (2024)
by: Roy, Utsha, et al.
Published: (2024)
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
An Empirical Study of Multi-Generation Sampling for Jailbreak Detection in Large Language Models
by: Luo, Hanrui, et al.
Published: (2026)
by: Luo, Hanrui, et al.
Published: (2026)
Similar Items
-
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
by: Kaneko, Masahiro, et al.
Published: (2025) -
Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
by: Kaneko, Masahiro, et al.
Published: (2025) -
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
by: Niwa, Ayana, et al.
Published: (2025) -
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
by: Kaneko, Masahiro
Published: (2026) -
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
by: Koike, Ryuto, et al.
Published: (2025)