SurveySum: A Dataset for Summarizing Multiple Scientific Articles into a Survey Section
Fuente:
arXiv
Saved in:
| Main Authors: | Fernandes, Leandro Carísio, Guedes, Gustavo Bartz, Laitz, Thiago Soares, Almeida, Thales Sales, Nogueira, Rodrigo, Lotufo, Roberto, Pereira, Jayr |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PublicHearingBR: A Brazilian Portuguese Dataset of Public Hearing Transcripts for Summarization of Long Documents
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
InRanker: Distilled Rankers for Zero-shot Information Retrieval
by: Laitz, Thiago, et al.
Published: (2024)
by: Laitz, Thiago, et al.
Published: (2024)
ExaRanker-Open: Synthetic Explanation for IR using Open-Source LLMs
by: Ferraretto, Fernando, et al.
Published: (2024)
by: Ferraretto, Fernando, et al.
Published: (2024)
Measuring Cross-lingual Transfer in Bytes
by: de Souza, Leandro Rodrigues, et al.
Published: (2024)
by: de Souza, Leandro Rodrigues, et al.
Published: (2024)
Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
by: Abonizio, Hugo, et al.
Published: (2025)
by: Abonizio, Hugo, et al.
Published: (2025)
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Classification and Clustering of Sentence-Level Embeddings of Scientific Articles Generated by Contrastive Learning
by: Guedes, Gustavo Bartz, et al.
Published: (2024)
by: Guedes, Gustavo Bartz, et al.
Published: (2024)
Sabiá-3 Technical Report
by: Abonizio, Hugo, et al.
Published: (2024)
by: Abonizio, Hugo, et al.
Published: (2024)
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Quati: A Brazilian Portuguese Information Retrieval Dataset from Native Speakers
by: Bueno, Mirelle, et al.
Published: (2024)
by: Bueno, Mirelle, et al.
Published: (2024)
JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections
by: Pereira, Jayr, et al.
Published: (2026)
by: Pereira, Jayr, et al.
Published: (2026)
Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Domain-Adaptive Dense Retrieval for Brazilian Legal Search
by: Pereira, Jayr, et al.
Published: (2026)
by: Pereira, Jayr, et al.
Published: (2026)
RAISE: Reasoning Agent for Interactive SQL Exploration
by: Granado, Fernando, et al.
Published: (2025)
by: Granado, Fernando, et al.
Published: (2025)
Check-Eval: A Checklist-based Approach for Evaluating Text Quality
by: Pereira, Jayr, et al.
Published: (2024)
by: Pereira, Jayr, et al.
Published: (2024)
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)
by: Bonás, Giovana Kerche, et al.
Published: (2026)
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
by: Junior, Roseval Malaquias, et al.
Published: (2026)
by: Junior, Roseval Malaquias, et al.
Published: (2026)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
by: Pires, Ramon, et al.
Published: (2026)
by: Pires, Ramon, et al.
Published: (2026)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Sabiá-4 Technical Report
by: Laitz, Thiago, et al.
Published: (2026)
by: Laitz, Thiago, et al.
Published: (2026)
Lissard: Long and Simple Sequential Reasoning Datasets
by: Bueno, Mirelle, et al.
Published: (2024)
by: Bueno, Mirelle, et al.
Published: (2024)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
Sabiá-2: A New Generation of Portuguese Large Language Models
by: Almeida, Thales Sales, et al.
Published: (2024)
by: Almeida, Thales Sales, et al.
Published: (2024)
PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
MLissard: Multilingual Long and Simple Sequential Reasoning Benchmarks
by: Bueno, Mirelle, et al.
Published: (2024)
by: Bueno, Mirelle, et al.
Published: (2024)
ptt5-v2: A Closer Look at Continued Pretraining of T5 Models for the Portuguese Language
by: Piau, Marcos, et al.
Published: (2024)
by: Piau, Marcos, et al.
Published: (2024)
The State and Fate of Summarization Datasets: A Survey
by: Dahan, Noam, et al.
Published: (2024)
by: Dahan, Noam, et al.
Published: (2024)
Predictive Authoring for Brazilian Portuguese Augmentative and Alternative Communication
by: Pereira, Jayr, et al.
Published: (2023)
by: Pereira, Jayr, et al.
Published: (2023)
Survey on Abstractive Text Summarization: Dataset, Models, and Metrics
by: Nnadi, Gospel Ozioma, et al.
Published: (2024)
by: Nnadi, Gospel Ozioma, et al.
Published: (2024)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
NoticIA: A Clickbait Article Summarization Dataset in Spanish
by: García-Ferrero, Iker, et al.
Published: (2024)
by: García-Ferrero, Iker, et al.
Published: (2024)
INACIA: Integrating Large Language Models in Brazilian Audit Courts: Opportunities and Challenges
by: Pereira, Jayr, et al.
Published: (2024)
by: Pereira, Jayr, et al.
Published: (2024)
MovieSum: An Abstractive Summarization Dataset for Movie Screenplays
by: Saxena, Rohit, et al.
Published: (2024)
by: Saxena, Rohit, et al.
Published: (2024)
JurisTCU: A Brazilian Portuguese Information Retrieval Dataset with Query Relevance Judgments
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
by: Fernandes, Leandro Carísio, et al.
Published: (2025)
Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines
by: Abonizio, Hugo, et al.
Published: (2026)
by: Abonizio, Hugo, et al.
Published: (2026)
Retailers and carriers’ viewpoint on Sorocaba’s city logistics: a spatial analysis
by: Thales Stevan Guedes Furquim
Published: (2020)
by: Thales Stevan Guedes Furquim
Published: (2020)
The interplay between domain specialization and model size
by: Junior, Roseval Malaquias, et al.
Published: (2025)
by: Junior, Roseval Malaquias, et al.
Published: (2025)
Similar Items
-
PublicHearingBR: A Brazilian Portuguese Dataset of Public Hearing Transcripts for Summarization of Long Documents
by: Fernandes, Leandro Carísio, et al.
Published: (2024) -
InRanker: Distilled Rankers for Zero-shot Information Retrieval
by: Laitz, Thiago, et al.
Published: (2024) -
ExaRanker-Open: Synthetic Explanation for IR using Open-Source LLMs
by: Ferraretto, Fernando, et al.
Published: (2024) -
Measuring Cross-lingual Transfer in Bytes
by: de Souza, Leandro Rodrigues, et al.
Published: (2024) -
Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
by: Abonizio, Hugo, et al.
Published: (2025)