Can We Trust LLM Detectors?
Fuente:
arXiv
Saved in:
| Main Authors: | Sandhan, Jivnesh, Jaiswal, Harshit, Cheng, Fei, Murawaki, Yugo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Persona Jailbreaking in Large Language Models
by: Sandhan, Jivnesh, et al.
Published: (2026)
by: Sandhan, Jivnesh, et al.
Published: (2026)
CAPE: Context-Aware Personality Evaluation Framework for Large Language Models
by: Sandhan, Jivnesh, et al.
Published: (2025)
by: Sandhan, Jivnesh, et al.
Published: (2025)
Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis
by: Matta, Shiho, et al.
Published: (2024)
by: Matta, Shiho, et al.
Published: (2024)
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
by: Zhong, Chengzhi, et al.
Published: (2025)
by: Zhong, Chengzhi, et al.
Published: (2025)
LLM-REVal: Can We Trust LLM Reviewers Yet?
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Beyond English-Centric LLMs: What Language Do Multilingual Language Models Think in?
by: Zhong, Chengzhi, et al.
Published: (2024)
by: Zhong, Chengzhi, et al.
Published: (2024)
Principal Component Analysis as a Sanity Check for Bayesian Phylolinguistic Reconstruction
by: Murawaki, Yugo
Published: (2024)
by: Murawaki, Yugo
Published: (2024)
Can We Trust a Black-box LLM? LLM Untrustworthy Boundary Detection via Bias-Diffusion and Multi-Agent Reinforcement Learning
by: Zhou, Xiaotian, et al.
Published: (2026)
by: Zhou, Xiaotian, et al.
Published: (2026)
Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models
by: Yan, Ruiyi, et al.
Published: (2025)
by: Yan, Ruiyi, et al.
Published: (2025)
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages
by: Ray, Pretam, et al.
Published: (2024)
by: Ray, Pretam, et al.
Published: (2024)
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
by: Harshit
Published: (2025)
by: Harshit
Published: (2025)
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
by: Prandi, Matteo, et al.
Published: (2025)
by: Prandi, Matteo, et al.
Published: (2025)
Efficient Provably Secure Linguistic Steganography via Range Coding
by: Yan, Ruiyi, et al.
Published: (2026)
by: Yan, Ruiyi, et al.
Published: (2026)
Still Not There: Can LLMs Outperform Smaller Task-Specific Seq2Seq Models on the Poetry-to-Prose Conversion Task?
by: Das, Kunal Kingkar, et al.
Published: (2025)
by: Das, Kunal Kingkar, et al.
Published: (2025)
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
by: Utami, Nabelanita, et al.
Published: (2026)
by: Utami, Nabelanita, et al.
Published: (2026)
Are We on the Right Way to Assessing LLM-as-a-Judge?
by: Feng, Yuanning, et al.
Published: (2025)
by: Feng, Yuanning, et al.
Published: (2025)
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
by: Chaudhary, Manav, et al.
Published: (2024)
by: Chaudhary, Manav, et al.
Published: (2024)
Anchored Sliding Window: Toward Robust and Imperceptible Linguistic Steganography
by: Yan, Ruiyi, et al.
Published: (2026)
by: Yan, Ruiyi, et al.
Published: (2026)
Can LLMs Generate Visualizations with Dataless Prompts?
by: Coelho, Darius, et al.
Published: (2024)
by: Coelho, Darius, et al.
Published: (2024)
Can We Locate and Prevent Stereotypes in LLMs?
by: D'Souza, Alex
Published: (2026)
by: D'Souza, Alex
Published: (2026)
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
LLMs Can Plan Only If We Tell Them
by: Sel, Bilgehan, et al.
Published: (2025)
by: Sel, Bilgehan, et al.
Published: (2025)
Citations and Trust in LLM Generated Responses
by: Ding, Yifan, et al.
Published: (2025)
by: Ding, Yifan, et al.
Published: (2025)
Can We Verify Step by Step for Incorrect Answer Detection?
by: Xu, Xin, et al.
Published: (2024)
by: Xu, Xin, et al.
Published: (2024)
Can We Edit LLMs for Long-Tail Biomedical Knowledge?
by: Yi, Xinhao, et al.
Published: (2025)
by: Yi, Xinhao, et al.
Published: (2025)
We Can't Understand AI Using our Existing Vocabulary
by: Hewitt, John, et al.
Published: (2025)
by: Hewitt, John, et al.
Published: (2025)
Dialogue You Can Trust: Human and AI Perspectives on Generated Conversations
by: Ebubechukwu, Ike, et al.
Published: (2024)
by: Ebubechukwu, Ike, et al.
Published: (2024)
LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?
by: Cheng, Xueqi, et al.
Published: (2026)
by: Cheng, Xueqi, et al.
Published: (2026)
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
by: Sinha, Aarush, et al.
Published: (2026)
by: Sinha, Aarush, et al.
Published: (2026)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
by: Hong, Junyuan, et al.
Published: (2024)
by: Hong, Junyuan, et al.
Published: (2024)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
by: Wang, Yidong, et al.
Published: (2025)
by: Wang, Yidong, et al.
Published: (2025)
Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking
by: Sarkar, Sujoy, et al.
Published: (2025)
by: Sarkar, Sujoy, et al.
Published: (2025)
ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
by: Waghjale, Siddhant, et al.
Published: (2024)
by: Waghjale, Siddhant, et al.
Published: (2024)
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
by: Aswal, Darpan, et al.
Published: (2025)
by: Aswal, Darpan, et al.
Published: (2025)
Enhancing Robustness of LLM-Synthetic Text Detectors for Academic Writing: A Comprehensive Analysis
by: Dou, Zhicheng, et al.
Published: (2024)
by: Dou, Zhicheng, et al.
Published: (2024)
How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?
by: Yamaguchi, Atsuki, et al.
Published: (2024)
by: Yamaguchi, Atsuki, et al.
Published: (2024)
MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text
by: Li, Chenjun, et al.
Published: (2026)
by: Li, Chenjun, et al.
Published: (2026)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
by: Sun, Xin, et al.
Published: (2026)
by: Sun, Xin, et al.
Published: (2026)
Similar Items
-
Persona Jailbreaking in Large Language Models
by: Sandhan, Jivnesh, et al.
Published: (2026) -
CAPE: Context-Aware Personality Evaluation Framework for Large Language Models
by: Sandhan, Jivnesh, et al.
Published: (2025) -
Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis
by: Matta, Shiho, et al.
Published: (2024) -
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
by: Zhong, Chengzhi, et al.
Published: (2025) -
LLM-REVal: Can We Trust LLM Reviewers Yet?
by: Li, Rui, et al.
Published: (2025)