Less Is More? When Dataset Context Hurts LLM-Generated Dataset Descriptions
Fuente:
arXiv
Salvato in:
| Autori principali: | Gan, Lisa-Yao, Das, Arunav, Walker, Johanna, Diepold, Klaus, Simperl, Elena |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Keywords are not always the key: A metadata field analysis for natural language search on open data portals
di: Gan, Lisa-Yao, et al.
Pubblicazione: (2025)
di: Gan, Lisa-Yao, et al.
Pubblicazione: (2025)
Less is More: Efficient Time Series Dataset Condensation via Two-fold Modal Matching--Extended Version
di: Miao, Hao, et al.
Pubblicazione: (2024)
di: Miao, Hao, et al.
Pubblicazione: (2024)
AutoDDG: Automated Dataset Description Generation using Large Language Models
di: Zhang, Haoxiang, et al.
Pubblicazione: (2025)
di: Zhang, Haoxiang, et al.
Pubblicazione: (2025)
Smaller and More Flexible Cuckoo Filters
di: Schmitz, Johanna Elena, et al.
Pubblicazione: (2025)
di: Schmitz, Johanna Elena, et al.
Pubblicazione: (2025)
LAKEGEN: A LLM-based Tabular Corpus Generator for Evaluating Dataset Discovery in Data Lakes
di: Dai, Zhenwei, et al.
Pubblicazione: (2025)
di: Dai, Zhenwei, et al.
Pubblicazione: (2025)
Prompting Datasets: Data Discovery with Conversational Agents
di: Walker, Johanna, et al.
Pubblicazione: (2023)
di: Walker, Johanna, et al.
Pubblicazione: (2023)
CogPic: A Multimodal Dataset for Early Cognitive Impairment Assessment via Picture Description Tasks
di: Wu, Liuyu, et al.
Pubblicazione: (2026)
di: Wu, Liuyu, et al.
Pubblicazione: (2026)
Is Long Context All You Need? Leveraging LLM's Extended Context for NL2SQL
di: Chung, Yeounoh, et al.
Pubblicazione: (2025)
di: Chung, Yeounoh, et al.
Pubblicazione: (2025)
FlexiDataGen: An Adaptive LLM Framework for Dynamic Semantic Dataset Generation in Sensitive Domains
di: Jelodar, Hamed, et al.
Pubblicazione: (2025)
di: Jelodar, Hamed, et al.
Pubblicazione: (2025)
AI data transparency: an exploration through the lens of AI incidents
di: Worth, Sophia, et al.
Pubblicazione: (2024)
di: Worth, Sophia, et al.
Pubblicazione: (2024)
Generating Skyline Datasets for Data Science Models
di: Wang, Mengying, et al.
Pubblicazione: (2025)
di: Wang, Mengying, et al.
Pubblicazione: (2025)
Distinctiveness Maximization in Datasets Assemblage
di: Wang, Tingting, et al.
Pubblicazione: (2024)
di: Wang, Tingting, et al.
Pubblicazione: (2024)
A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives
di: Li, Pengyue, et al.
Pubblicazione: (2025)
di: Li, Pengyue, et al.
Pubblicazione: (2025)
Dataset Discovery via Line Charts
di: Ji, Daomin, et al.
Pubblicazione: (2024)
di: Ji, Daomin, et al.
Pubblicazione: (2024)
TableNet A Large-Scale Table Dataset with LLM-Powered Autonomous
di: Zhang, Ruilin, et al.
Pubblicazione: (2026)
di: Zhang, Ruilin, et al.
Pubblicazione: (2026)
LLMClean: Context-Aware Tabular Data Cleaning via LLM-Generated OFDs
di: Biester, Fabian, et al.
Pubblicazione: (2024)
di: Biester, Fabian, et al.
Pubblicazione: (2024)
A Standardized Machine-readable Dataset Documentation Format for Responsible AI
di: Jain, Nitisha, et al.
Pubblicazione: (2024)
di: Jain, Nitisha, et al.
Pubblicazione: (2024)
ChatPD: An LLM-driven Paper-Dataset Networking System
di: Xu, Anjie, et al.
Pubblicazione: (2025)
di: Xu, Anjie, et al.
Pubblicazione: (2025)
Query Based Construction of Chronic Disease Datasets
di: Ngo, Vuong M., et al.
Pubblicazione: (2024)
di: Ngo, Vuong M., et al.
Pubblicazione: (2024)
Croissant: A Metadata Format for ML-Ready Datasets
di: Akhtar, Mubashara, et al.
Pubblicazione: (2024)
di: Akhtar, Mubashara, et al.
Pubblicazione: (2024)
DataLens: Enhancing Dataset Discovery via Network Topologies
di: Ollagnier, Anaïs, et al.
Pubblicazione: (2025)
di: Ollagnier, Anaïs, et al.
Pubblicazione: (2025)
SchemaDB: Structures in Relational Datasets
di: Christopher, Cody James, et al.
Pubblicazione: (2021)
di: Christopher, Cody James, et al.
Pubblicazione: (2021)
Conceptual Schema Inference for Tabular Datasets using Large Language Models
di: Wu, Zhenyu, et al.
Pubblicazione: (2026)
di: Wu, Zhenyu, et al.
Pubblicazione: (2026)
A Unified Approach for Multi-Granularity Search over Spatial Datasets
di: Yang, Wenzhe, et al.
Pubblicazione: (2024)
di: Yang, Wenzhe, et al.
Pubblicazione: (2024)
Jelly-Patch: a Fast Format for Recording Changes in RDF Datasets
di: Sowinski, Piotr, et al.
Pubblicazione: (2025)
di: Sowinski, Piotr, et al.
Pubblicazione: (2025)
From RDF Graph Validation to RDF Dataset Validation with SHACL-DS
di: Dao, Davan Chiem, et al.
Pubblicazione: (2025)
di: Dao, Davan Chiem, et al.
Pubblicazione: (2025)
The FormAI Dataset: Generative AI in Software Security Through the Lens of Formal Verification
di: Tihanyi, Norbert, et al.
Pubblicazione: (2023)
di: Tihanyi, Norbert, et al.
Pubblicazione: (2023)
PBE Meets LLM: When Few Examples Aren't Few-Shot Enough
di: Zhang, Shuning, et al.
Pubblicazione: (2025)
di: Zhang, Shuning, et al.
Pubblicazione: (2025)
OSM+: Billion-Level OpenStreetMap Dataset for City-wide Experiments
di: Zheng, Guanjie, et al.
Pubblicazione: (2025)
di: Zheng, Guanjie, et al.
Pubblicazione: (2025)
MOCAS: A Multimodal Dataset for Objective Cognitive Workload Assessment on Simultaneous Tasks
di: Jo, Wonse, et al.
Pubblicazione: (2022)
di: Jo, Wonse, et al.
Pubblicazione: (2022)
Computing the Non-Dominated Flexible Skyline in Vertically Distributed Datasets with No Random Access
di: Martinenghi, Davide
Pubblicazione: (2024)
di: Martinenghi, Davide
Pubblicazione: (2024)
Joinable Search over Multi-source Spatial Datasets: Overlap, Coverage, and Efficiency
di: Yang, Wenzhe, et al.
Pubblicazione: (2023)
di: Yang, Wenzhe, et al.
Pubblicazione: (2023)
Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets
di: Attrach, Rafi Al, et al.
Pubblicazione: (2026)
di: Attrach, Rafi Al, et al.
Pubblicazione: (2026)
An Intelligent Innovation Dataset on Scientific Research Outcomes
di: Wu, Xinran, et al.
Pubblicazione: (2024)
di: Wu, Xinran, et al.
Pubblicazione: (2024)
Within-Dataset Disclosure Risk for Differential Privacy
di: Zhu, Zhiru, et al.
Pubblicazione: (2023)
di: Zhu, Zhiru, et al.
Pubblicazione: (2023)
Global Dataset of Solar Power Plants: Multidimensional Integration and Analysis
di: Mantilla-Guerra, Anibal, et al.
Pubblicazione: (2026)
di: Mantilla-Guerra, Anibal, et al.
Pubblicazione: (2026)
StraTyper: Automated Semantic Type Discovery and Multi-Type Annotation for Dataset Collections
di: Koutras, Christos, et al.
Pubblicazione: (2026)
di: Koutras, Christos, et al.
Pubblicazione: (2026)
Evaluating Data Quality Tools: Measurement Capabilities and LLM Integration
di: Rehberger, Tobias, et al.
Pubblicazione: (2026)
di: Rehberger, Tobias, et al.
Pubblicazione: (2026)
RF-Behavior: A Multimodal Radio-Frequency Dataset for Human Behavior and Emotion Analysis
di: Zuo, Si, et al.
Pubblicazione: (2025)
di: Zuo, Si, et al.
Pubblicazione: (2025)
Context-Enriched Natural Language Descriptions of Vessel Trajectories
di: Patroumpas, Kostas, et al.
Pubblicazione: (2026)
di: Patroumpas, Kostas, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Keywords are not always the key: A metadata field analysis for natural language search on open data portals
di: Gan, Lisa-Yao, et al.
Pubblicazione: (2025) -
Less is More: Efficient Time Series Dataset Condensation via Two-fold Modal Matching--Extended Version
di: Miao, Hao, et al.
Pubblicazione: (2024) -
AutoDDG: Automated Dataset Description Generation using Large Language Models
di: Zhang, Haoxiang, et al.
Pubblicazione: (2025) -
Smaller and More Flexible Cuckoo Filters
di: Schmitz, Johanna Elena, et al.
Pubblicazione: (2025) -
LAKEGEN: A LLM-based Tabular Corpus Generator for Evaluating Dataset Discovery in Data Lakes
di: Dai, Zhenwei, et al.
Pubblicazione: (2025)