PRODIGy: a PROfile-based DIalogue Generation dataset
Fuente:
arXiv
Guardado en:
| Autores principales: | Occhipinti, Daniela, Tekiroglu, Serra Sinem, Guerini, Marco |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When Harry Meets Superman: The Role of The Interlocutor in Persona-Based Dialogue Generation
por: Occhipinti, Daniela, et al.
Publicado: (2025)
por: Occhipinti, Daniela, et al.
Publicado: (2025)
Fine-tuning with HED-IT: The impact of human post-editing for dialogical language models
por: Occhipinti, Daniela, et al.
Publicado: (2024)
por: Occhipinti, Daniela, et al.
Publicado: (2024)
CrisiText: A dataset of warning messages for LLM training in emergency communication
por: Gonella, Giacomo, et al.
Publicado: (2025)
por: Gonella, Giacomo, et al.
Publicado: (2025)
Face the Facts! Evaluating RAG-based Pipelines for Professional Fact-Checking
por: Russo, Daniel, et al.
Publicado: (2024)
por: Russo, Daniel, et al.
Publicado: (2024)
Basque and Spanish Counter Narrative Generation: Data Creation and Evaluation
por: Bengoetxea, Jaione, et al.
Publicado: (2024)
por: Bengoetxea, Jaione, et al.
Publicado: (2024)
Assisted Counterspeech Writing at the Crossroads of Hate Speech and Misinformation
por: Martone, Genoveffa, et al.
Publicado: (2026)
por: Martone, Genoveffa, et al.
Publicado: (2026)
Don't Stop the Multi-Party! On Generating Synthetic Written Multi-Party Conversations with Constraints
por: Penzo, Nicolò, et al.
Publicado: (2025)
por: Penzo, Nicolò, et al.
Publicado: (2025)
NLP for Counterspeech against Hate: A Survey and How-To Guide
por: Bonaldi, Helena, et al.
Publicado: (2024)
por: Bonaldi, Helena, et al.
Publicado: (2024)
Putting Context in Context: the Impact of Discussion Structure on Text Classification
por: Penzo, Nicolò, et al.
Publicado: (2024)
por: Penzo, Nicolò, et al.
Publicado: (2024)
Do LLMs suffer from Multi-Party Hangover? A Diagnostic Approach to Addressee Recognition and Response Selection in Conversations
por: Penzo, Nicolò, et al.
Publicado: (2024)
por: Penzo, Nicolò, et al.
Publicado: (2024)
LLMberjack: Guided Trimming of Debate Trees for Multi-Party Conversation Creation
por: Bottona, Leonardo, et al.
Publicado: (2026)
por: Bottona, Leonardo, et al.
Publicado: (2026)
A document processing pipeline for the construction of a dataset for topic modeling based on the judgments of the Italian Supreme Court
por: Marulli, Matteo, et al.
Publicado: (2025)
por: Marulli, Matteo, et al.
Publicado: (2025)
Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering
por: Bonaldi, Helena, et al.
Publicado: (2024)
por: Bonaldi, Helena, et al.
Publicado: (2024)
BUSTER: a "BUSiness Transaction Entity Recognition" dataset
por: Zugarini, Andrea, et al.
Publicado: (2024)
por: Zugarini, Andrea, et al.
Publicado: (2024)
Figurative Archive: an open dataset and web-based application for the study of metaphor
por: Bressler, Maddalena, et al.
Publicado: (2025)
por: Bressler, Maddalena, et al.
Publicado: (2025)
Grammar and Gameplay-aligned RL for Game Description Generation with LLMs
por: Tanaka, Tsunehiko, et al.
Publicado: (2025)
por: Tanaka, Tsunehiko, et al.
Publicado: (2025)
I Want to Break Free! Persuasion and Anti-Social Behavior of LLMs in Multi-Agent Settings with Social Hierarchy
por: Campedelli, Gian Maria, et al.
Publicado: (2024)
por: Campedelli, Gian Maria, et al.
Publicado: (2024)
OpenStaxQA: A multilingual dataset based on open-source college textbooks
por: Gupta, Pranav
Publicado: (2025)
por: Gupta, Pranav
Publicado: (2025)
Estimating Knowledge in Large Language Models Without Generating a Single Token
por: Gottesman, Daniela, et al.
Publicado: (2024)
por: Gottesman, Daniela, et al.
Publicado: (2024)
Automated essay scoring in Arabic: a dataset and analysis of a BERT-based system
por: Ghazawi, Rayed, et al.
Publicado: (2024)
por: Ghazawi, Rayed, et al.
Publicado: (2024)
GenFighter: A Generative and Evolutive Textual Attack Removal
por: Islam, Md Athikul, et al.
Publicado: (2024)
por: Islam, Md Athikul, et al.
Publicado: (2024)
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
por: Weck, Benno, et al.
Publicado: (2023)
por: Weck, Benno, et al.
Publicado: (2023)
Toxic language detection: a systematic review of Arabic datasets
por: Bensalem, Imene, et al.
Publicado: (2023)
por: Bensalem, Imene, et al.
Publicado: (2023)
MAEBE: Multi-Agent Emergent Behavior Framework
por: Erisken, Sinem, et al.
Publicado: (2025)
por: Erisken, Sinem, et al.
Publicado: (2025)
Exa-PSD: a new Persian sentiment analysis dataset on Twitter
por: Ghaderi, Seyed Himan, et al.
Publicado: (2026)
por: Ghaderi, Seyed Himan, et al.
Publicado: (2026)
EnzChemRED, a rich enzyme chemistry relation extraction dataset
por: Lai, Po-Ting, et al.
Publicado: (2024)
por: Lai, Po-Ting, et al.
Publicado: (2024)
A Turkish Educational Crossword Puzzle Generator
por: Zeinalipour, Kamyar, et al.
Publicado: (2024)
por: Zeinalipour, Kamyar, et al.
Publicado: (2024)
A social context-aware graph-based multimodal attentive learning framework for disaster content classification during emergencies: a benchmark dataset and method
por: Dar, Shahid Shafi, et al.
Publicado: (2024)
por: Dar, Shahid Shafi, et al.
Publicado: (2024)
ADI-20: Arabic Dialect Identification dataset and models
por: Elleuch, Haroun, et al.
Publicado: (2025)
por: Elleuch, Haroun, et al.
Publicado: (2025)
TANQ: An open domain dataset of table answered questions
por: Akhtar, Mubashara, et al.
Publicado: (2024)
por: Akhtar, Mubashara, et al.
Publicado: (2024)
Automating Turkish Educational Quiz Generation Using Large Language Models
por: Zeinalipour, Kamyar, et al.
Publicado: (2024)
por: Zeinalipour, Kamyar, et al.
Publicado: (2024)
ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
por: Jourdan, Léane, et al.
Publicado: (2025)
por: Jourdan, Léane, et al.
Publicado: (2025)
MaterioMiner -- An ontology-based text mining dataset for extraction of process-structure-property entities
por: Durmaz, Ali Riza, et al.
Publicado: (2024)
por: Durmaz, Ali Riza, et al.
Publicado: (2024)
Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara
por: Diarra, Yacouba, et al.
Publicado: (2025)
por: Diarra, Yacouba, et al.
Publicado: (2025)
SecureBreak -- A dataset towards safe and secure models
por: Arazzi, Marco, et al.
Publicado: (2026)
por: Arazzi, Marco, et al.
Publicado: (2026)
XNLIeu: a dataset for cross-lingual NLI in Basque
por: Heredia, Maite, et al.
Publicado: (2024)
por: Heredia, Maite, et al.
Publicado: (2024)
HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America
por: Ivetta, Guido, et al.
Publicado: (2025)
por: Ivetta, Guido, et al.
Publicado: (2025)
Fostering Natural Conversation in Large Language Models with NICO: a Natural Interactive COnversation dataset
por: Sun, Renliang, et al.
Publicado: (2024)
por: Sun, Renliang, et al.
Publicado: (2024)
DepressionEmo: A novel dataset for multilabel classification of depression emotions
por: Rahman, Abu Bakar Siddiqur, et al.
Publicado: (2024)
por: Rahman, Abu Bakar Siddiqur, et al.
Publicado: (2024)
GLeMM: A large-scale multilingual dataset for morphological research
por: Nabil, Hathout, et al.
Publicado: (2026)
por: Nabil, Hathout, et al.
Publicado: (2026)
Ejemplares similares
-
When Harry Meets Superman: The Role of The Interlocutor in Persona-Based Dialogue Generation
por: Occhipinti, Daniela, et al.
Publicado: (2025) -
Fine-tuning with HED-IT: The impact of human post-editing for dialogical language models
por: Occhipinti, Daniela, et al.
Publicado: (2024) -
CrisiText: A dataset of warning messages for LLM training in emergency communication
por: Gonella, Giacomo, et al.
Publicado: (2025) -
Face the Facts! Evaluating RAG-based Pipelines for Professional Fact-Checking
por: Russo, Daniel, et al.
Publicado: (2024) -
Basque and Spanish Counter Narrative Generation: Data Creation and Evaluation
por: Bengoetxea, Jaione, et al.
Publicado: (2024)