Efficient data selection employing Semantic Similarity-based Graph Structures for model training
Fuente:
arXiv
Saved in:
| Main Authors: | Petcu, Roxana, Maji, Subhadeep |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Graph Structures to Detect Hallucinations in Large Language Models
by: Nonkes, Noa, et al.
Published: (2024)
by: Nonkes, Noa, et al.
Published: (2024)
BhashaSetu: Cross-Lingual Knowledge Transfer from High-Resource to Extreme Low-Resource Languages
by: Maji, Subhadip, et al.
Published: (2026)
by: Maji, Subhadip, et al.
Published: (2026)
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
by: Burns, Thomas F, et al.
Published: (2025)
by: Burns, Thomas F, et al.
Published: (2025)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
by: Nikitin, Alexander, et al.
Published: (2024)
by: Nikitin, Alexander, et al.
Published: (2024)
Similarity-based Neighbor Selection for Graph LLMs
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
by: Bonthu, Sridevi, et al.
Published: (2025)
by: Bonthu, Sridevi, et al.
Published: (2025)
Semantic Refinement with LLMs for Graph Representations
by: Thapaliya, Safal, et al.
Published: (2025)
by: Thapaliya, Safal, et al.
Published: (2025)
Unsupervised Domain Adaptation with Global and Local Graph Neural Networks in Limited Labeled Data Scenario: Application to Disaster Management
by: Ghosh, Samujjwal, et al.
Published: (2021)
by: Ghosh, Samujjwal, et al.
Published: (2021)
MiMIC: Multi-Modal Indian Earnings Calls Dataset to Predict Stock Prices
by: Ghosh, Sohom, et al.
Published: (2025)
by: Ghosh, Sohom, et al.
Published: (2025)
Post-training for Efficient Communication via Convention Formation
by: Hua, Yilun, et al.
Published: (2025)
by: Hua, Yilun, et al.
Published: (2025)
Efficient Post-training Quantization with FP8 Formats
by: Shen, Haihao, et al.
Published: (2023)
by: Shen, Haihao, et al.
Published: (2023)
Parallel Structures in Pre-training Data Yield In-Context Learning
by: Chen, Yanda, et al.
Published: (2024)
by: Chen, Yanda, et al.
Published: (2024)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
by: Liu, Yichen, et al.
Published: (2022)
by: Liu, Yichen, et al.
Published: (2022)
Self-Improving Pretraining: using post-trained models to pretrain better models
by: Tan, Ellen Xiaoqing, et al.
Published: (2026)
by: Tan, Ellen Xiaoqing, et al.
Published: (2026)
SubSearch: Intermediate Rewards for Unsupervised Guided Reasoning in Complex Retrieval
by: Petcu, Roxana, et al.
Published: (2026)
by: Petcu, Roxana, et al.
Published: (2026)
Negation Neglect: When models fail to learn negations in training
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
Does Pre-trained Language Model Actually Infer Unseen Links in Knowledge Graph Completion?
by: Sakai, Yusuke, et al.
Published: (2023)
by: Sakai, Yusuke, et al.
Published: (2023)
Post-training makes large language models less human-like
by: Binz, Marcel, et al.
Published: (2026)
by: Binz, Marcel, et al.
Published: (2026)
Adjoint sharding for very long context training of state space models
by: Xu, Xingzi, et al.
Published: (2025)
by: Xu, Xingzi, et al.
Published: (2025)
CASE: Efficient Curricular Data Pre-training for Building Assistive Psychology Expert Models
by: Harne, Sarthak, et al.
Published: (2024)
by: Harne, Sarthak, et al.
Published: (2024)
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
by: Xia, Mengzhou, et al.
Published: (2023)
by: Xia, Mengzhou, et al.
Published: (2023)
Efficient Text-Attributed Graph Learning through Selective Annotation and Graph Alignment
by: Xie, Huanyi, et al.
Published: (2025)
by: Xie, Huanyi, et al.
Published: (2025)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
by: Cao, Qi, et al.
Published: (2026)
by: Cao, Qi, et al.
Published: (2026)
PhoneLM:an Efficient and Capable Small Language Model Family through Principled Pre-training
by: Yi, Rongjie, et al.
Published: (2024)
by: Yi, Rongjie, et al.
Published: (2024)
SEMFED: Semantic-Aware Resource-Efficient Federated Learning for Heterogeneous NLP Tasks
by: Hussain, Sajid, et al.
Published: (2025)
by: Hussain, Sajid, et al.
Published: (2025)
Efficient Relational Context Perception for Knowledge Graph Completion
by: Tu, Wenkai, et al.
Published: (2024)
by: Tu, Wenkai, et al.
Published: (2024)
CDGP: Automatic Cloze Distractor Generation based on Pre-trained Language Model
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
Structured Query Construction via Knowledge Graph Embedding
by: Wang, Ruijie, et al.
Published: (2019)
by: Wang, Ruijie, et al.
Published: (2019)
A Survey of Quantized Graph Representation Learning: Connecting Graph Structures with Large Language Models
by: Lin, Qika, et al.
Published: (2025)
by: Lin, Qika, et al.
Published: (2025)
HelpSteer2: Open-source dataset for training top-performing reward models
by: Wang, Zhilin, et al.
Published: (2024)
by: Wang, Zhilin, et al.
Published: (2024)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Multimodal Fusion of EHR in Structures and Semantics: Integrating Clinical Records and Notes with Hypergraph and LLM
by: Cui, Hejie, et al.
Published: (2024)
by: Cui, Hejie, et al.
Published: (2024)
One Swallow Does Not Make a Summer: Understanding Semantic Structures in Embedding Spaces
by: Sun, Yandong, et al.
Published: (2025)
by: Sun, Yandong, et al.
Published: (2025)
Efficient End-to-end Language Model Fine-tuning on Graphs
by: Xue, Rui, et al.
Published: (2023)
by: Xue, Rui, et al.
Published: (2023)
DoPAMine: Domain-specific Pre-training Adaptation from seed-guided data Mining
by: Arannil, Vinayak, et al.
Published: (2024)
by: Arannil, Vinayak, et al.
Published: (2024)
A Survey on Recent Advances in Conversational Data Generation
by: Soudani, Heydar, et al.
Published: (2024)
by: Soudani, Heydar, et al.
Published: (2024)
A New Method for Cross-Lingual-based Semantic Role Labeling
by: Ebrahimi, Mohammad, et al.
Published: (2024)
by: Ebrahimi, Mohammad, et al.
Published: (2024)
Mitigate Negative Transfer with Similarity Heuristic Lifelong Prompt Tuning
by: Wu, Chenyuan, et al.
Published: (2024)
by: Wu, Chenyuan, et al.
Published: (2024)
Earley-Driven Dynamic Pruning for Efficient Structured Decoding
by: Sun, Xintong, et al.
Published: (2025)
by: Sun, Xintong, et al.
Published: (2025)
Model-based Subsampling for Knowledge Graph Completion
by: Feng, Xincan, et al.
Published: (2023)
by: Feng, Xincan, et al.
Published: (2023)
Similar Items
-
Leveraging Graph Structures to Detect Hallucinations in Large Language Models
by: Nonkes, Noa, et al.
Published: (2024) -
BhashaSetu: Cross-Lingual Knowledge Transfer from High-Resource to Extreme Low-Resource Languages
by: Maji, Subhadip, et al.
Published: (2026) -
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
by: Burns, Thomas F, et al.
Published: (2025) -
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
by: Nikitin, Alexander, et al.
Published: (2024) -
Similarity-based Neighbor Selection for Graph LLMs
by: Li, Rui, et al.
Published: (2024)