RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Xin, Wu, Junchao, Yang, Shu, Zhan, Runzhe, Wu, Zeyu, Luo, Ziyang, Wang, Di, Yang, Min, Chao, Lidia S., Wong, Derek F. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
von: Wu, Junchao, et al.
Veröffentlicht: (2023)
von: Wu, Junchao, et al.
Veröffentlicht: (2023)
Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
Rethinking Prompt-based Debiasing in Large Language Models
von: Yang, Xinyi, et al.
Veröffentlicht: (2025)
von: Yang, Xinyi, et al.
Veröffentlicht: (2025)
Neuron-Aware Data Selection In Instruction Tuning For Large Language Models
von: Chen, Xin, et al.
Veröffentlicht: (2026)
von: Chen, Xin, et al.
Veröffentlicht: (2026)
Prefix Text as a Yarn: Eliciting Non-English Alignment in Foundation Language Model
von: Zhan, Runzhe, et al.
Veröffentlicht: (2024)
von: Zhan, Runzhe, et al.
Veröffentlicht: (2024)
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025)
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025)
Let's Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model
von: Xu, Haoyun, et al.
Veröffentlicht: (2024)
von: Xu, Haoyun, et al.
Veröffentlicht: (2024)
Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety
von: Huang, Yuyi, et al.
Veröffentlicht: (2025)
von: Huang, Yuyi, et al.
Veröffentlicht: (2025)
Intrinsic Model Weaknesses: How Priming Attacks Unveil Vulnerabilities in Large Language Models
von: Huang, Yuyi, et al.
Veröffentlicht: (2025)
von: Huang, Yuyi, et al.
Veröffentlicht: (2025)
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
von: Ma, Jingkun, et al.
Veröffentlicht: (2024)
von: Ma, Jingkun, et al.
Veröffentlicht: (2024)
Is Long-to-Short a Free Lunch? Investigating Inconsistency and Reasoning Efficiency in LRMs
von: Yang, Shu, et al.
Veröffentlicht: (2025)
von: Yang, Shu, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
von: Zhang, Jiayi, et al.
Veröffentlicht: (2025)
von: Zhang, Jiayi, et al.
Veröffentlicht: (2025)
Understanding Aha Moments: from External Observations to Internal Mechanisms
von: Yang, Shu, et al.
Veröffentlicht: (2025)
von: Yang, Shu, et al.
Veröffentlicht: (2025)
Fraud-R1 : A Multi-Round Benchmark for Assessing the Robustness of LLM Against Augmented Fraud and Phishing Inducements
von: Yang, Shu, et al.
Veröffentlicht: (2025)
von: Yang, Shu, et al.
Veröffentlicht: (2025)
Investigating CoT Monitorability in Large Reasoning Models
von: Yang, Shu, et al.
Veröffentlicht: (2025)
von: Yang, Shu, et al.
Veröffentlicht: (2025)
Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation
von: Sun, Yanming, et al.
Veröffentlicht: (2025)
von: Sun, Yanming, et al.
Veröffentlicht: (2025)
DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection
von: Wu, Junchao, et al.
Veröffentlicht: (2026)
von: Wu, Junchao, et al.
Veröffentlicht: (2026)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
von: Wang, Shanshan, et al.
Veröffentlicht: (2025)
von: Wang, Shanshan, et al.
Veröffentlicht: (2025)
Writing Patterns Reveal a Hidden Division of Labor in Scientific Teams
von: Yang, Lulin, et al.
Veröffentlicht: (2025)
von: Yang, Lulin, et al.
Veröffentlicht: (2025)
Repr Types: One Abstraction to Rule Them All
von: Palmkvist, Viktor, et al.
Veröffentlicht: (2024)
von: Palmkvist, Viktor, et al.
Veröffentlicht: (2024)
Concept Factorization via Self-Representation and Adaptive Graph Structure Learning
von: Yang, Zhengqin, et al.
Veröffentlicht: (2025)
von: Yang, Zhengqin, et al.
Veröffentlicht: (2025)
Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluation
von: Li, Tong, et al.
Veröffentlicht: (2025)
von: Li, Tong, et al.
Veröffentlicht: (2025)
HiddenGuard: Fine-Grained Safe Generation with Specialized Representation Router
von: Mei, Lingrui, et al.
Veröffentlicht: (2024)
von: Mei, Lingrui, et al.
Veröffentlicht: (2024)
Making RL with Preference-based Feedback Efficient via Randomization
von: Wu, Runzhe, et al.
Veröffentlicht: (2023)
von: Wu, Runzhe, et al.
Veröffentlicht: (2023)
The Pattern of Exotic Hidden-Heavy Hadrons Revealed
von: Bruschini, Roberto
Veröffentlicht: (2025)
von: Bruschini, Roberto
Veröffentlicht: (2025)
Self-Guard: Empower the LLM to Safeguard Itself
von: Wang, Zezhong, et al.
Veröffentlicht: (2023)
von: Wang, Zezhong, et al.
Veröffentlicht: (2023)
Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models
von: Gong, Xilin, et al.
Veröffentlicht: (2026)
von: Gong, Xilin, et al.
Veröffentlicht: (2026)
Finite Satisfiability of the Two-Variable Guarded Fragment with Transitive Guards and Related Variants
von: Kieronski, Emanuel, et al.
Veröffentlicht: (2016)
von: Kieronski, Emanuel, et al.
Veröffentlicht: (2016)
On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning
von: Ye, Xiaotian, et al.
Veröffentlicht: (2026)
von: Ye, Xiaotian, et al.
Veröffentlicht: (2026)
Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination
von: He, Xiaoqi, et al.
Veröffentlicht: (2026)
von: He, Xiaoqi, et al.
Veröffentlicht: (2026)
Steer-to-Detect: Probing Hidden Representations for Detection of LLM-Generated Texts
von: Liang, Luxu, et al.
Veröffentlicht: (2026)
von: Liang, Luxu, et al.
Veröffentlicht: (2026)
Harnessing LLM to Attack LLM-Guarded Text-to-Image Models
von: Deng, Yimo, et al.
Veröffentlicht: (2023)
von: Deng, Yimo, et al.
Veröffentlicht: (2023)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
Language-Image Alignment with Fixed Text Encoders
von: Yang, Jingfeng, et al.
Veröffentlicht: (2025)
von: Yang, Jingfeng, et al.
Veröffentlicht: (2025)
Hidden Division of Labor in Scientific Teams Revealed Through 1.6 Million LaTeX Files
von: Pei, Jiaxin, et al.
Veröffentlicht: (2025)
von: Pei, Jiaxin, et al.
Veröffentlicht: (2025)
Can ChatGPT Really Understand Modern Chinese Poetry?
von: Wang, Shanshan, et al.
Veröffentlicht: (2026)
von: Wang, Shanshan, et al.
Veröffentlicht: (2026)
A Two-Stage Prediction-Aware Contrastive Learning Framework for Multi-Intent NLU
von: Chen, Guanhua, et al.
Veröffentlicht: (2024)
von: Chen, Guanhua, et al.
Veröffentlicht: (2024)
What is the Best Way for ChatGPT to Translate Poetry?
von: Wang, Shanshan, et al.
Veröffentlicht: (2024)
von: Wang, Shanshan, et al.
Veröffentlicht: (2024)
Unveiling LLMs' Metaphorical Understanding: Exploring Conceptual Irrelevance, Context Leveraging and Syntactic Influence
von: Ye, Fengying, et al.
Veröffentlicht: (2025)
von: Ye, Fengying, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
von: Wu, Junchao, et al.
Veröffentlicht: (2023) -
Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
von: Wu, Junchao, et al.
Veröffentlicht: (2024) -
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
von: Wu, Junchao, et al.
Veröffentlicht: (2024) -
Rethinking Prompt-based Debiasing in Large Language Models
von: Yang, Xinyi, et al.
Veröffentlicht: (2025) -
Neuron-Aware Data Selection In Instruction Tuning For Large Language Models
von: Chen, Xin, et al.
Veröffentlicht: (2026)