Measuring Dataset Diversity from a Geometric Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Ba, Yang, Abolhasani, Mohammad Sadeq, Mancenido, Michelle V, Pan, Rong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predict Training Data Quality via Its Geometry in Metric Space
by: Ba, Yang, et al.
Published: (2025)
by: Ba, Yang, et al.
Published: (2025)
Data Diversity as Implicit Regularization: How Does Diversity Shape the Weight Space of Deep Neural Networks?
by: Ba, Yang, et al.
Published: (2024)
by: Ba, Yang, et al.
Published: (2024)
Leveraging LLM for Automated Ontology Extraction and Knowledge Graph Generation
by: Abolhasani, Mohammad Sadeq, et al.
Published: (2024)
by: Abolhasani, Mohammad Sadeq, et al.
Published: (2024)
Beyond Predefined Schemas: TRACE-KG for Context-Enriched Knowledge Graphs from Complex Documents
by: Abolhasani, Mohammad Sadeq, et al.
Published: (2026)
by: Abolhasani, Mohammad Sadeq, et al.
Published: (2026)
Fill In The Gaps: Model Calibration and Generalization with Synthetic Data
by: Ba, Yang, et al.
Published: (2024)
by: Ba, Yang, et al.
Published: (2024)
Data Quality in Crowdsourcing and Spamming Behavior Detection
by: Ba, Yang, et al.
Published: (2024)
by: Ba, Yang, et al.
Published: (2024)
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets
by: Zakizadeh, Mahdi, et al.
Published: (2025)
by: Zakizadeh, Mahdi, et al.
Published: (2025)
Measuring Diversity in Synthetic Datasets
by: Zhu, Yuchang, et al.
Published: (2025)
by: Zhu, Yuchang, et al.
Published: (2025)
MemArchitect: A Policy Driven Memory Governance Layer
by: Kumar, Lingavasan Suresh, et al.
Published: (2026)
by: Kumar, Lingavasan Suresh, et al.
Published: (2026)
Mastering Diverse Domains through World Models
by: Hafner, Danijar, et al.
Published: (2023)
by: Hafner, Danijar, et al.
Published: (2023)
Diversity-Aware Reverse Kullback-Leibler Divergence for Large Language Model Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026)
by: Luong, Hoang-Chau, et al.
Published: (2026)
Measuring Diversity of Game Scenarios
by: Li, Yuchen, et al.
Published: (2024)
by: Li, Yuchen, et al.
Published: (2024)
Temporal convolutional and fusional transformer model with Bi-LSTM encoder-decoder for multi-time-window remaining useful life prediction
by: Pour, Mohamadreza Akbari, et al.
Published: (2025)
by: Pour, Mohamadreza Akbari, et al.
Published: (2025)
Bridging Conversational and Collaborative Signals for Conversational Recommendation
by: Rabiah, Ahmad Bin, et al.
Published: (2024)
by: Rabiah, Ahmad Bin, et al.
Published: (2024)
Tree-Sliced Wasserstein Distance: A Geometric Perspective
by: Tran, Viet-Hoang, et al.
Published: (2024)
by: Tran, Viet-Hoang, et al.
Published: (2024)
Reasoning in Large Language Models: A Geometric Perspective
by: Cosentino, Romain, et al.
Published: (2024)
by: Cosentino, Romain, et al.
Published: (2024)
WolBanking77: Wolof Banking Speech Intent Classification Dataset
by: Kandji, Abdou Karim, et al.
Published: (2025)
by: Kandji, Abdou Karim, et al.
Published: (2025)
EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL
by: Zhang, Lunjun, et al.
Published: (2026)
by: Zhang, Lunjun, et al.
Published: (2026)
Perspective Dial: Measuring Perspective of Text and Guiding LLM Outputs
by: Kim, Taejin, et al.
Published: (2025)
by: Kim, Taejin, et al.
Published: (2025)
Limited Linguistic Diversity in Embodied AI Datasets
by: Wanna, Selma, et al.
Published: (2026)
by: Wanna, Selma, et al.
Published: (2026)
MCD: Diverse Large-Scale Multi-Campus Dataset for Robot Perception
by: Nguyen, Thien-Minh, et al.
Published: (2024)
by: Nguyen, Thien-Minh, et al.
Published: (2024)
Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues
by: Ba, Zhongjie, et al.
Published: (2026)
by: Ba, Zhongjie, et al.
Published: (2026)
Exploration by Learning Diverse Skills through Successor State Measures
by: Tolguenec, Paul-Antoine Le, et al.
Published: (2024)
by: Tolguenec, Paul-Antoine Le, et al.
Published: (2024)
What do Geometric Hallucination Detection Metrics Actually Measure?
by: Yeats, Eric, et al.
Published: (2026)
by: Yeats, Eric, et al.
Published: (2026)
LLM Active Alignment: A Nash Equilibrium Perspective
by: Wang, Tonghan, et al.
Published: (2026)
by: Wang, Tonghan, et al.
Published: (2026)
UNSEEN: Enhancing Dataset Pruning from a Generalization Perspective
by: Xu, Furui, et al.
Published: (2025)
by: Xu, Furui, et al.
Published: (2025)
Rethinking Purity and Diversity in Multi-Behavior Sequential Recommendation from the Frequency Perspective
by: Han, Yongqiang, et al.
Published: (2025)
by: Han, Yongqiang, et al.
Published: (2025)
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
Beyond Autonomy: A Dynamic Tiered AgentRunner Framework for Governable and Resilient Enterprise AI Execution
by: Pan, Kai, et al.
Published: (2026)
by: Pan, Kai, et al.
Published: (2026)
Causal Discovery by Interventions via Integer Programming
by: Elrefaey, Abdelmonem, et al.
Published: (2024)
by: Elrefaey, Abdelmonem, et al.
Published: (2024)
MEC-IP: Efficient Discovery of Markov Equivalent Classes via Integer Programming
by: Elrefaey, Abdelmonem, et al.
Published: (2024)
by: Elrefaey, Abdelmonem, et al.
Published: (2024)
PromptForge-350k: A Large-Scale Dataset and Contrastive Framework for Prompt-Based AI Image Forgery Localization
by: Wang, Jianpeng, et al.
Published: (2026)
by: Wang, Jianpeng, et al.
Published: (2026)
Measuring What Matters Beyond Text: Evaluating Multimodal Summaries by Quality, Alignment, and Diversity
by: Ali, Abid, et al.
Published: (2026)
by: Ali, Abid, et al.
Published: (2026)
Generating Diverse Synthetic Datasets for Evaluation of Real-life Recommender Systems
by: Malenšek, Miha, et al.
Published: (2024)
by: Malenšek, Miha, et al.
Published: (2024)
Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
On the Roles of LLMs in Planning: Embedding LLMs into Planning Graphs
by: Zhuo, Hankz Hankui, et al.
Published: (2024)
by: Zhuo, Hankz Hankui, et al.
Published: (2024)
Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure
by: Yin, Michelle, et al.
Published: (2026)
by: Yin, Michelle, et al.
Published: (2026)
Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation
by: Zheng, Lulu, et al.
Published: (2026)
by: Zheng, Lulu, et al.
Published: (2026)
We Need to Measure Data Diversity in NLP -- Better and Broader
by: Nguyen, Dong, et al.
Published: (2025)
by: Nguyen, Dong, et al.
Published: (2025)
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
by: Song, Yueqi, et al.
Published: (2025)
by: Song, Yueqi, et al.
Published: (2025)
Similar Items
-
Predict Training Data Quality via Its Geometry in Metric Space
by: Ba, Yang, et al.
Published: (2025) -
Data Diversity as Implicit Regularization: How Does Diversity Shape the Weight Space of Deep Neural Networks?
by: Ba, Yang, et al.
Published: (2024) -
Leveraging LLM for Automated Ontology Extraction and Knowledge Graph Generation
by: Abolhasani, Mohammad Sadeq, et al.
Published: (2024) -
Beyond Predefined Schemas: TRACE-KG for Context-Enriched Knowledge Graphs from Complex Documents
by: Abolhasani, Mohammad Sadeq, et al.
Published: (2026) -
Fill In The Gaps: Model Calibration and Generalization with Synthetic Data
by: Ba, Yang, et al.
Published: (2024)