ToBlend: Token-Level Blending With an Ensemble of LLMs to Attack AI-Generated Text Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Fan, Kwak, Haewoon, An, Jisun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
di: Huang, Fan, et al.
Pubblicazione: (2026)
di: Huang, Fan, et al.
Pubblicazione: (2026)
Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability
di: Huang, Fan, et al.
Pubblicazione: (2026)
di: Huang, Fan, et al.
Pubblicazione: (2026)
ChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?
di: Huang, Fan, et al.
Pubblicazione: (2024)
di: Huang, Fan, et al.
Pubblicazione: (2024)
Can Lessons From Human Teams Be Applied to Multi-Agent Systems? The Role of Structure, Diversity, and Interaction Dynamics
di: Muralidharan, Rasika, et al.
Pubblicazione: (2025)
di: Muralidharan, Rasika, et al.
Pubblicazione: (2025)
Can we trust the evaluation on ChatGPT?
di: Aiyappa, Rachith, et al.
Pubblicazione: (2023)
di: Aiyappa, Rachith, et al.
Pubblicazione: (2023)
Benchmarking zero-shot stance detection with FlanT5-XXL: Insights from training data, prompting, and decoding strategies into its near-SoTA performance
di: Aiyappa, Rachith, et al.
Pubblicazione: (2024)
di: Aiyappa, Rachith, et al.
Pubblicazione: (2024)
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
di: Kachwala, Zoher, et al.
Pubblicazione: (2026)
di: Kachwala, Zoher, et al.
Pubblicazione: (2026)
CogBias: Measuring and Mitigating Cognitive Bias in Large Language Models
di: Huang, Fan, et al.
Pubblicazione: (2026)
di: Huang, Fan, et al.
Pubblicazione: (2026)
A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies
di: Qi, Weihong, et al.
Pubblicazione: (2025)
di: Qi, Weihong, et al.
Pubblicazione: (2025)
XChoice: Explainable Evaluation of AI-Human Alignment in LLM-based Constrained Choice Decision Making
di: Qi, Weihong, et al.
Pubblicazione: (2026)
di: Qi, Weihong, et al.
Pubblicazione: (2026)
SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs
di: Hu, Yebowen, et al.
Pubblicazione: (2024)
di: Hu, Yebowen, et al.
Pubblicazione: (2024)
Somatic in the East, Psychological in the West?: Investigating Clinically-Grounded Cross-Cultural Depression Symptom Expression in LLMs
di: Sakai, Shintaro, et al.
Pubblicazione: (2025)
di: Sakai, Shintaro, et al.
Pubblicazione: (2025)
Rematch: Robust and Efficient Matching of Local Knowledge Graphs to Improve Structural and Semantic Similarity
di: Kachwala, Zoher, et al.
Pubblicazione: (2024)
di: Kachwala, Zoher, et al.
Pubblicazione: (2024)
When to Ensemble: Identifying Token-Level Points for Stable and Fast LLM Ensembling
di: Yun, Heecheol, et al.
Pubblicazione: (2025)
di: Yun, Heecheol, et al.
Pubblicazione: (2025)
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
di: Koneru, Sai, et al.
Pubblicazione: (2024)
di: Koneru, Sai, et al.
Pubblicazione: (2024)
Ever-Evolving Memory by Blending and Refining the Past
di: Kim, Seo Hyun, et al.
Pubblicazione: (2024)
di: Kim, Seo Hyun, et al.
Pubblicazione: (2024)
Domain Gating Ensemble Networks for AI-Generated Text Detection
di: Tripathi, Arihant, et al.
Pubblicazione: (2025)
di: Tripathi, Arihant, et al.
Pubblicazione: (2025)
Adversarial Attacks on AI-Generated Text Detection Models: A Token Probability-Based Approach Using Embeddings
di: Kadhim, Ahmed K., et al.
Pubblicazione: (2025)
di: Kadhim, Ahmed K., et al.
Pubblicazione: (2025)
Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses
di: Hussain, Khizar, et al.
Pubblicazione: (2026)
di: Hussain, Khizar, et al.
Pubblicazione: (2026)
Fine-Grained Detection of AI-Generated Text Using Sentence-Level Segmentation
di: Teja, Lekkala Sai, et al.
Pubblicazione: (2025)
di: Teja, Lekkala Sai, et al.
Pubblicazione: (2025)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
di: Huang, Zeyu, et al.
Pubblicazione: (2025)
di: Huang, Zeyu, et al.
Pubblicazione: (2025)
AI Generated Text Detection
di: Alikhanov, Adilkhan, et al.
Pubblicazione: (2026)
di: Alikhanov, Adilkhan, et al.
Pubblicazione: (2026)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
di: Teja, Lekkala Sai, et al.
Pubblicazione: (2025)
di: Teja, Lekkala Sai, et al.
Pubblicazione: (2025)
The Perfect Blend: Redefining RLHF with Mixture of Judges
di: Xu, Tengyu, et al.
Pubblicazione: (2024)
di: Xu, Tengyu, et al.
Pubblicazione: (2024)
LuxVeri at GenAI Detection Task 1: Inverse Perplexity Weighted Ensemble for Robust Detection of AI-Generated Text across English and Multilingual Contexts
di: Mobin, Md Kamrujjaman, et al.
Pubblicazione: (2025)
di: Mobin, Md Kamrujjaman, et al.
Pubblicazione: (2025)
Sarang at DEFACTIFY 4.0: Detecting AI-Generated Text Using Noised Data and an Ensemble of DeBERTa Models
di: Trivedi, Avinash, et al.
Pubblicazione: (2025)
di: Trivedi, Avinash, et al.
Pubblicazione: (2025)
Vietnamese AI Generated Text Detection
di: Tran, Quang-Dan, et al.
Pubblicazione: (2024)
di: Tran, Quang-Dan, et al.
Pubblicazione: (2024)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
di: Lu, Xiaoding, et al.
Pubblicazione: (2024)
di: Lu, Xiaoding, et al.
Pubblicazione: (2024)
LLMs Can Infer Political Alignment from Online Conversations
di: Lee, Byunghwee, et al.
Pubblicazione: (2026)
di: Lee, Byunghwee, et al.
Pubblicazione: (2026)
Multi-Facet Blending for Faceted Query-by-Example Retrieval
di: Do, Heejin, et al.
Pubblicazione: (2024)
di: Do, Heejin, et al.
Pubblicazione: (2024)
LuxVeri at GenAI Detection Task 3: Cross-Domain Detection of AI-Generated Text Using Inverse Perplexity-Weighted Ensemble of Fine-Tuned Transformer Models
di: Mobin, Md Kamrujjaman, et al.
Pubblicazione: (2025)
di: Mobin, Md Kamrujjaman, et al.
Pubblicazione: (2025)
Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
di: Li, Yanhong, et al.
Pubblicazione: (2025)
di: Li, Yanhong, et al.
Pubblicazione: (2025)
Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
di: Xiao, Feng, et al.
Pubblicazione: (2025)
di: Xiao, Feng, et al.
Pubblicazione: (2025)
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
di: Han, Chao, et al.
Pubblicazione: (2025)
di: Han, Chao, et al.
Pubblicazione: (2025)
Text Generation Beyond Discrete Token Sampling
di: Zhuang, Yufan, et al.
Pubblicazione: (2025)
di: Zhuang, Yufan, et al.
Pubblicazione: (2025)
Multiscale Positive-Unlabeled Detection of AI-Generated Texts
di: Tian, Yuchuan, et al.
Pubblicazione: (2023)
di: Tian, Yuchuan, et al.
Pubblicazione: (2023)
Fine-tuned Large Language Models (LLMs): Improved Prompt Injection Attacks Detection
di: Rahman, Md Abdur, et al.
Pubblicazione: (2024)
di: Rahman, Md Abdur, et al.
Pubblicazione: (2024)
JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency
di: Cai, Aichen, et al.
Pubblicazione: (2026)
di: Cai, Aichen, et al.
Pubblicazione: (2026)
Detecting Machine-Generated Texts: Not Just "AI vs Humans" and Explainability is Complicated
di: Ji, Jiazhou, et al.
Pubblicazione: (2024)
di: Ji, Jiazhou, et al.
Pubblicazione: (2024)
Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
di: Yu, Yao-Ching, et al.
Pubblicazione: (2024)
di: Yu, Yao-Ching, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
di: Huang, Fan, et al.
Pubblicazione: (2026) -
Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability
di: Huang, Fan, et al.
Pubblicazione: (2026) -
ChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?
di: Huang, Fan, et al.
Pubblicazione: (2024) -
Can Lessons From Human Teams Be Applied to Multi-Agent Systems? The Role of Structure, Diversity, and Interaction Dynamics
di: Muralidharan, Rasika, et al.
Pubblicazione: (2025) -
Can we trust the evaluation on ChatGPT?
di: Aiyappa, Rachith, et al.
Pubblicazione: (2023)