Introducing Three New Benchmark Datasets for Hierarchical Text Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Toit, Jaco du, Redelinghuys, Herman, Dunaiski, Marcel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Combining Language and Topic Models for Hierarchical Text Classification
by: Toit, Jaco du, et al.
Published: (2025)
by: Toit, Jaco du, et al.
Published: (2025)
Exploring User Retrieval Integration towards Large Language Models for Cross-Domain Sequential Recommendation
by: Shen, Tingjia, et al.
Published: (2024)
by: Shen, Tingjia, et al.
Published: (2024)
Reliable Part-of-Speech Tagging of Historical Corpora through Set-Valued Prediction
by: Heid, Stefan, et al.
Published: (2020)
by: Heid, Stefan, et al.
Published: (2020)
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
by: Gautam, Sushant, et al.
Published: (2024)
by: Gautam, Sushant, et al.
Published: (2024)
Evaluation of Table Representations to Answer Questions from Tables in Documents : A Case Study using 3GPP Specifications
by: Roychowdhury, Sujoy, et al.
Published: (2024)
by: Roychowdhury, Sujoy, et al.
Published: (2024)
Suppressing Domain-Specific Hallucination in Construction LLMs: A Knowledge Graph Foundation for GraphRAG and QLoRA on River and Sediment Control Technical Standards
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
SURE-RAG: Sufficiency and Uncertainty-Aware Evidence Verification for Selective Retrieval-Augmented Generation
by: Qiu, Jingxi, et al.
Published: (2026)
by: Qiu, Jingxi, et al.
Published: (2026)
Comparison of Metadata Representation Models for Knowledge Graph Embeddings
by: Egami, Shusaku, et al.
Published: (2025)
by: Egami, Shusaku, et al.
Published: (2025)
MUDY: Multi-Granular Dynamic Candidate Contextualization for Unsupervised Keyphrase Extraction
by: Kang, Hyeongu, et al.
Published: (2026)
by: Kang, Hyeongu, et al.
Published: (2026)
Using LLM-Based Approaches to Enhance and Automate Topic Labeling
by: Khandelwal, Trishia
Published: (2025)
by: Khandelwal, Trishia
Published: (2025)
Numbers Matter! Bringing Quantity-awareness to Retrieval Systems
by: Almasian, Satya, et al.
Published: (2024)
by: Almasian, Satya, et al.
Published: (2024)
HiPS: Hierarchical PDF Segmentation of Textbooks
by: Wehnert, Sabine, et al.
Published: (2025)
by: Wehnert, Sabine, et al.
Published: (2025)
Language Models and Retrieval Augmented Generation for Automated Structured Data Extraction from Diagnostic Reports
by: Jabal, Mohamed Sobhi, et al.
Published: (2024)
by: Jabal, Mohamed Sobhi, et al.
Published: (2024)
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
by: Teixeira, Tiago, et al.
Published: (2026)
by: Teixeira, Tiago, et al.
Published: (2026)
A Method for Detecting Legal Article Competition for Korean Criminal Law Using a Case-augmented Mention Graph
by: An, Seonho, et al.
Published: (2024)
by: An, Seonho, et al.
Published: (2024)
Using text embedding models as text classifiers with medical data
by: Goel, Rishabh
Published: (2024)
by: Goel, Rishabh
Published: (2024)
Learning variant product relationship and variation attributes from e-commerce website structures
by: Herrero-Vidal, Pedro, et al.
Published: (2024)
by: Herrero-Vidal, Pedro, et al.
Published: (2024)
Knowledge Distillation of Domain-adapted LLMs for Question-Answering in Telecom
by: Sen, Rishika, et al.
Published: (2025)
by: Sen, Rishika, et al.
Published: (2025)
From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering
by: Santos, José Guilherme Marques dos, et al.
Published: (2026)
by: Santos, José Guilherme Marques dos, et al.
Published: (2026)
Annif at the GermEval-2025 LLMs4Subjects Task: Traditional XMTC Augmented by Efficient LLMs
by: Suominen, Osma, et al.
Published: (2025)
by: Suominen, Osma, et al.
Published: (2025)
RecaLLM: Addressing the Lost-in-Thought Phenomenon with Explicit In-Context Retrieval
by: Whitecross, Kyle, et al.
Published: (2026)
by: Whitecross, Kyle, et al.
Published: (2026)
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
by: Couturier, Camille, et al.
Published: (2025)
by: Couturier, Camille, et al.
Published: (2025)
Topic Is Not Agenda: A Citation-Community Audit of Text Embeddings
by: Yoo, Junseon
Published: (2026)
by: Yoo, Junseon
Published: (2026)
skLEP: A Slovak General Language Understanding Benchmark
by: Šuppa, Marek, et al.
Published: (2025)
by: Šuppa, Marek, et al.
Published: (2025)
Disaster Question Answering with LoRA Efficiency and Accurate End Position
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
LLMs in the Loop: Leveraging Large Language Model Annotations for Active Learning in Low-Resource Languages
by: Kholodna, Nataliia, et al.
Published: (2024)
by: Kholodna, Nataliia, et al.
Published: (2024)
KnowThyself: An Agentic Assistant for LLM Interpretability
by: Prasai, Suraj, et al.
Published: (2025)
by: Prasai, Suraj, et al.
Published: (2025)
Uncovering the Limitations of Query Performance Prediction: Failures, Insights, and Implications for Selective Query Processing
by: Chifu, Adrian-Gabriel, et al.
Published: (2025)
by: Chifu, Adrian-Gabriel, et al.
Published: (2025)
ATANT v1.1: Positioning Continuity Evaluation Against Memory, Long-Context, and Agentic-Memory Benchmarks
by: Tanguturi, Samuel Sameer
Published: (2026)
by: Tanguturi, Samuel Sameer
Published: (2026)
A Language Model based Framework for New Concept Placement in Ontologies
by: Dong, Hang, et al.
Published: (2024)
by: Dong, Hang, et al.
Published: (2024)
HELIOS: Harmonizing Early Fusion, Late Fusion, and LLM Reasoning for Multi-Granular Table-Text Retrieval
by: Park, Sungho, et al.
Published: (2026)
by: Park, Sungho, et al.
Published: (2026)
MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature
by: Ratul, Md Toyaha Rahman, et al.
Published: (2026)
by: Ratul, Md Toyaha Rahman, et al.
Published: (2026)
Algorithmic Trust and Compliance: Benchmarking Brand Notability for UK iGaming Entities in Generative Search Engines
by: Oruesagasti, Julen
Published: (2026)
by: Oruesagasti, Julen
Published: (2026)
Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review
by: Rombach, Alexander Michael, et al.
Published: (2024)
by: Rombach, Alexander Michael, et al.
Published: (2024)
HiFi-RAG: Hierarchical Content Filtering and Two-Pass Generation for Open-Domain RAG
by: Nuengsigkapian, Cattalyya
Published: (2025)
by: Nuengsigkapian, Cattalyya
Published: (2025)
Document Understanding for Healthcare Referrals
by: Mistry, Jimit, et al.
Published: (2023)
by: Mistry, Jimit, et al.
Published: (2023)
Detection of ChatGPT Fake Science with the xFakeSci Learning Algorithm
by: Hamed, Ahmed Abdeen, et al.
Published: (2023)
by: Hamed, Ahmed Abdeen, et al.
Published: (2023)
MDKeyChunker: Single-Call LLM Enrichment with Rolling Keys and Key-Based Restructuring for High-Accuracy RAG
by: Mangla, Bhavik
Published: (2026)
by: Mangla, Bhavik
Published: (2026)
An efficient domain-independent approach for supervised keyphrase extraction and ranking
by: Ramaswamy, Sriraghavendra
Published: (2024)
by: Ramaswamy, Sriraghavendra
Published: (2024)
CUE-R: Beyond the Final Answer in Retrieval-Augmented Generation
by: Jain, Siddharth, et al.
Published: (2026)
by: Jain, Siddharth, et al.
Published: (2026)
Similar Items
-
Combining Language and Topic Models for Hierarchical Text Classification
by: Toit, Jaco du, et al.
Published: (2025) -
Exploring User Retrieval Integration towards Large Language Models for Cross-Domain Sequential Recommendation
by: Shen, Tingjia, et al.
Published: (2024) -
Reliable Part-of-Speech Tagging of Historical Corpora through Set-Valued Prediction
by: Heid, Stefan, et al.
Published: (2020) -
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
by: Gautam, Sushant, et al.
Published: (2024) -
Evaluation of Table Representations to Answer Questions from Tables in Documents : A Case Study using 3GPP Specifications
by: Roychowdhury, Sujoy, et al.
Published: (2024)