Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Beining, Zubiaga, Arkaitz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Technical Report on the Pangram AI-Generated Text Classifier
von: Emi, Bradley, et al.
Veröffentlicht: (2024)
von: Emi, Bradley, et al.
Veröffentlicht: (2024)
CATER: Leveraging LLM to Pioneer a Multidimensional, Reference-Independent Paradigm in Translation Quality Evaluation
von: IIDA, Kurando, et al.
Veröffentlicht: (2024)
von: IIDA, Kurando, et al.
Veröffentlicht: (2024)
Dynamic Demonstration Retrieval and Cognitive Understanding for Emotional Support Conversation
von: Xu, Zhe, et al.
Veröffentlicht: (2024)
von: Xu, Zhe, et al.
Veröffentlicht: (2024)
Pun Unintended: LLMs and the Illusion of Humor Understanding
von: Zangari, Alessandro, et al.
Veröffentlicht: (2025)
von: Zangari, Alessandro, et al.
Veröffentlicht: (2025)
HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent
von: Xu, Weijie, et al.
Veröffentlicht: (2024)
von: Xu, Weijie, et al.
Veröffentlicht: (2024)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
Raw Text is All you Need: Knowledge-intensive Multi-turn Instruction Tuning for Large Language Model
von: Hou, Xia, et al.
Veröffentlicht: (2024)
von: Hou, Xia, et al.
Veröffentlicht: (2024)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning
von: Ming, Xiaoyang, et al.
Veröffentlicht: (2026)
von: Ming, Xiaoyang, et al.
Veröffentlicht: (2026)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
von: Muñoz-Ortiz, Alberto, et al.
Veröffentlicht: (2023)
von: Muñoz-Ortiz, Alberto, et al.
Veröffentlicht: (2023)
A Study on Bias Detection and Classification in Natural Language Processing
von: Evans, Ana Sofia, et al.
Veröffentlicht: (2024)
von: Evans, Ana Sofia, et al.
Veröffentlicht: (2024)
Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
von: Zong, Chang, et al.
Veröffentlicht: (2024)
von: Zong, Chang, et al.
Veröffentlicht: (2024)
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
Topeax -- An Improved Clustering Topic Model with Density Peak Detection and Lexical-Semantic Term Importance
von: Kardos, Márton
Veröffentlicht: (2026)
von: Kardos, Márton
Veröffentlicht: (2026)
Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
von: Görge, Rebekka, et al.
Veröffentlicht: (2025)
von: Görge, Rebekka, et al.
Veröffentlicht: (2025)
EmPO: Emotion Grounding for Empathetic Response Generation through Preference Optimization
von: Sotolar, Ondrej, et al.
Veröffentlicht: (2024)
von: Sotolar, Ondrej, et al.
Veröffentlicht: (2024)
Efficient Adaptive Rejection Sampling for Accelerating Speculative Decoding in Large Language Models
von: Sun, Chendong, et al.
Veröffentlicht: (2025)
von: Sun, Chendong, et al.
Veröffentlicht: (2025)
S2vNTM: Semi-supervised vMF Neural Topic Modeling
von: Xu, Weijie, et al.
Veröffentlicht: (2023)
von: Xu, Weijie, et al.
Veröffentlicht: (2023)
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
von: Fang, Xi, et al.
Veröffentlicht: (2025)
von: Fang, Xi, et al.
Veröffentlicht: (2025)
MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models
von: Mehta, Rahul, et al.
Veröffentlicht: (2024)
von: Mehta, Rahul, et al.
Veröffentlicht: (2024)
Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation
von: Rouzegar, Hamidreza, et al.
Veröffentlicht: (2024)
von: Rouzegar, Hamidreza, et al.
Veröffentlicht: (2024)
Do LLMs Truly Understand When a Precedent Is Overruled?
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Investigating on RLHF methodology
von: Kutalev, Alexey, et al.
Veröffentlicht: (2024)
von: Kutalev, Alexey, et al.
Veröffentlicht: (2024)
Performance Evaluation of Sentiment Analysis on Text and Emoji Data Using End-to-End, Transfer Learning, Distributed and Explainable AI Models
von: Velampalli, Sirisha, et al.
Veröffentlicht: (2025)
von: Velampalli, Sirisha, et al.
Veröffentlicht: (2025)
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
von: Sela, Omer
Veröffentlicht: (2026)
von: Sela, Omer
Veröffentlicht: (2026)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context
von: An, Tao
Veröffentlicht: (2025)
von: An, Tao
Veröffentlicht: (2025)
One Agent to Serve All: a Lite-Adaptive Stylized AI Assistant for Millions of Multi-Style Official Accounts
von: Fan, Xingyu, et al.
Veröffentlicht: (2025)
von: Fan, Xingyu, et al.
Veröffentlicht: (2025)
A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents
von: Soman, Sumit, et al.
Veröffentlicht: (2025)
von: Soman, Sumit, et al.
Veröffentlicht: (2025)
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
von: Alam, Firoj, et al.
Veröffentlicht: (2025)
von: Alam, Firoj, et al.
Veröffentlicht: (2025)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
von: Schmidt, Craig W., et al.
Veröffentlicht: (2025)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2025)
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
Reducing Hallucinations in Summarization via Reinforcement Learning with Entity Hallucination Index
von: Katwe, Praveenkumar, et al.
Veröffentlicht: (2025)
von: Katwe, Praveenkumar, et al.
Veröffentlicht: (2025)
Bielik 11B v2 Technical Report
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025)
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
Multi-chain Graph Refinement and Selection for Reliable Reasoning in Large Language Models
von: Yang, Yujiao, et al.
Veröffentlicht: (2025)
von: Yang, Yujiao, et al.
Veröffentlicht: (2025)
The Impact of Role Design in In-Context Learning for Large Language Models
von: Rouzegar, Hamidreza, et al.
Veröffentlicht: (2025)
von: Rouzegar, Hamidreza, et al.
Veröffentlicht: (2025)
Semantic Synergy: Unlocking Policy Insights and Learning Pathways Through Advanced Skill Mapping
von: Koundouri, Phoebe, et al.
Veröffentlicht: (2025)
von: Koundouri, Phoebe, et al.
Veröffentlicht: (2025)
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2024)
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2024)
TREX: Tokenizer Regression for Optimal Data Mixture
von: Won, Inho, et al.
Veröffentlicht: (2026)
von: Won, Inho, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Technical Report on the Pangram AI-Generated Text Classifier
von: Emi, Bradley, et al.
Veröffentlicht: (2024) -
CATER: Leveraging LLM to Pioneer a Multidimensional, Reference-Independent Paradigm in Translation Quality Evaluation
von: IIDA, Kurando, et al.
Veröffentlicht: (2024) -
Dynamic Demonstration Retrieval and Cognitive Understanding for Emotional Support Conversation
von: Xu, Zhe, et al.
Veröffentlicht: (2024) -
Pun Unintended: LLMs and the Illusion of Humor Understanding
von: Zangari, Alessandro, et al.
Veröffentlicht: (2025) -
HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent
von: Xu, Weijie, et al.
Veröffentlicht: (2024)