Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data
Fuente:
arXiv
Guardado en:
| Autor principal: | Lübbers, Christopher Lee |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Human Understanding of Paraphrase Types in Large Language Models
por: Meier, Dominik, et al.
Publicado: (2024)
por: Meier, Dominik, et al.
Publicado: (2024)
Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents
por: Kim, San, et al.
Publicado: (2024)
por: Kim, San, et al.
Publicado: (2024)
Neural Machine Translation for Malayalam Paraphrase Generation
por: Varghese, Christeena, et al.
Publicado: (2024)
por: Varghese, Christeena, et al.
Publicado: (2024)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
por: Michail, Andrianos, et al.
Publicado: (2024)
por: Michail, Andrianos, et al.
Publicado: (2024)
SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
por: Yun, Janghyeon, et al.
Publicado: (2025)
por: Yun, Janghyeon, et al.
Publicado: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
por: Ashuach, Tomer, et al.
Publicado: (2025)
por: Ashuach, Tomer, et al.
Publicado: (2025)
DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory
por: Zhou, Wenxuan, et al.
Publicado: (2025)
por: Zhou, Wenxuan, et al.
Publicado: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
Big City Bias: Evaluating the Impact of Metropolitan Size on Computational Job Market Abilities of Language Models
por: Campanella, Charlie, et al.
Publicado: (2024)
por: Campanella, Charlie, et al.
Publicado: (2024)
After Retrieval, Before Generation: Enhancing the Trustworthiness of Large Language Models in Retrieval-Augmented Generation
por: Dai, Xinbang, et al.
Publicado: (2025)
por: Dai, Xinbang, et al.
Publicado: (2025)
Unveiling Attractor Cycles in Large Language Models: A Dynamical Systems View of Successive Paraphrasing
por: Wang, Zhilin, et al.
Publicado: (2025)
por: Wang, Zhilin, et al.
Publicado: (2025)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
por: Cui, Hyang
Publicado: (2025)
por: Cui, Hyang
Publicado: (2025)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
por: Dejl, Adam, et al.
Publicado: (2025)
por: Dejl, Adam, et al.
Publicado: (2025)
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
por: Hinterleitner, Lukas, et al.
Publicado: (2026)
por: Hinterleitner, Lukas, et al.
Publicado: (2026)
TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
por: Liu, Aiwei, et al.
Publicado: (2024)
por: Liu, Aiwei, et al.
Publicado: (2024)
Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models
por: Chen, Yanbing, et al.
Publicado: (2024)
por: Chen, Yanbing, et al.
Publicado: (2024)
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
por: Hu, Yuxuan, et al.
Publicado: (2025)
por: Hu, Yuxuan, et al.
Publicado: (2025)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
por: Ashuach, Tomer, et al.
Publicado: (2024)
por: Ashuach, Tomer, et al.
Publicado: (2024)
Lacuna Language Learning: Leveraging RNNs for Ranked Text Completion in Digitized Coptic Manuscripts
por: Levine, Lauren, et al.
Publicado: (2024)
por: Levine, Lauren, et al.
Publicado: (2024)
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
por: Dugan, Liam, et al.
Publicado: (2024)
por: Dugan, Liam, et al.
Publicado: (2024)
Synthia: Scalable Grounded Persona Generation from Social Media Data
por: Rahimzadeh, Vahid, et al.
Publicado: (2025)
por: Rahimzadeh, Vahid, et al.
Publicado: (2025)
OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation
por: Yu, Jinzheng, et al.
Publicado: (2025)
por: Yu, Jinzheng, et al.
Publicado: (2025)
SynSym: A Synthetic Data Generation Framework for Psychiatric Symptom Identification
por: Kang, Migyeong, et al.
Publicado: (2026)
por: Kang, Migyeong, et al.
Publicado: (2026)
Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension and Text Generation
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
por: Costa-jussà, Marta R., et al.
Publicado: (2024)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
por: Wang, Xintao, et al.
Publicado: (2026)
por: Wang, Xintao, et al.
Publicado: (2026)
Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
por: Xu, Beining, et al.
Publicado: (2025)
por: Xu, Beining, et al.
Publicado: (2025)
Unifying the Scope of Bridging Anaphora Types in English: Bridging Annotations in ARRAU and GUM
por: Levine, Lauren, et al.
Publicado: (2024)
por: Levine, Lauren, et al.
Publicado: (2024)
Enhancing Emotion Prediction in News Headlines: Insights from ChatGPT and Seq2Seq Models for Free-Text Generation
por: Gao, Ge, et al.
Publicado: (2024)
por: Gao, Ge, et al.
Publicado: (2024)
Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology
por: Kocbek, Primož, et al.
Publicado: (2025)
por: Kocbek, Primož, et al.
Publicado: (2025)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
por: Dinh, Tu Anh, et al.
Publicado: (2024)
por: Dinh, Tu Anh, et al.
Publicado: (2024)
Please Make it Sound like Human: Encoder-Decoder vs. Decoder-Only Transformers for AI-to-Human Text Style Transfer
por: Paneru, Utsav
Publicado: (2026)
por: Paneru, Utsav
Publicado: (2026)
Investigating the Impact of Text Summarization on Topic Modeling
por: Khandelwal, Trishia
Publicado: (2024)
por: Khandelwal, Trishia
Publicado: (2024)
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering
por: Muller, Sacha, et al.
Publicado: (2024)
por: Muller, Sacha, et al.
Publicado: (2024)
From Classification to Ranking: Enhancing LLM Reasoning Capabilities for MBTI Personality Detection
por: Cao, Yuan, et al.
Publicado: (2026)
por: Cao, Yuan, et al.
Publicado: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
por: CH-Wang, Sky, et al.
Publicado: (2025)
por: CH-Wang, Sky, et al.
Publicado: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
por: Liu, Zhongxin, et al.
Publicado: (2025)
por: Liu, Zhongxin, et al.
Publicado: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
por: Tong, Jingqi, et al.
Publicado: (2025)
por: Tong, Jingqi, et al.
Publicado: (2025)
LLMs and the Human Condition
por: Wallis, Peter
Publicado: (2024)
por: Wallis, Peter
Publicado: (2024)
Ejemplares similares
-
Towards Human Understanding of Paraphrase Types in Large Language Models
por: Meier, Dominik, et al.
Publicado: (2024) -
Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents
por: Kim, San, et al.
Publicado: (2024) -
Neural Machine Translation for Malayalam Paraphrase Generation
por: Varghese, Christeena, et al.
Publicado: (2024) -
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
por: Michail, Andrianos, et al.
Publicado: (2024) -
SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
por: Yun, Janghyeon, et al.
Publicado: (2025)