LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xuemiao, Ren, Can, Tu, Chengying, Weng, Rongxiang, Yan, Hongfei, Wang, Jingang, Cai, Xunliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large-Scale Diverse Synthesis for Mid-Training
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
A Survey on LLM Mid-Training
by: Tu, Chengying, et al.
Published: (2025)
by: Tu, Chengying, et al.
Published: (2025)
FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training
by: Xu, Liangyu, et al.
Published: (2025)
by: Xu, Liangyu, et al.
Published: (2025)
FRAME: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text
by: Xu, Zhihao, et al.
Published: (2026)
by: Xu, Zhihao, et al.
Published: (2026)
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Libra: Assessing and Improving Reward Model by Learning to Think
by: Zhou, Meng, et al.
Published: (2025)
by: Zhou, Meng, et al.
Published: (2025)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
by: Bai, Yang, et al.
Published: (2026)
by: Bai, Yang, et al.
Published: (2026)
Length Desensitization in Direct Preference Optimization
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
BioGraphletQA: Knowledge-Anchored Generation of Complex QA Datasets
by: Jonker, Richard A. A., et al.
Published: (2026)
by: Jonker, Richard A. A., et al.
Published: (2026)
RJUA-QA: A Comprehensive QA Dataset for Urology
by: Lyu, Shiwei, et al.
Published: (2023)
by: Lyu, Shiwei, et al.
Published: (2023)
Enhancing LLMs via High-Knowledge Data Selection
by: Duan, Feiyu, et al.
Published: (2025)
by: Duan, Feiyu, et al.
Published: (2025)
StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children's Story-Based Learning
by: Chen, Jiaju, et al.
Published: (2023)
by: Chen, Jiaju, et al.
Published: (2023)
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
by: Cao, Zhihan, et al.
Published: (2025)
by: Cao, Zhihan, et al.
Published: (2025)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
by: Liu, Jiahao, et al.
Published: (2024)
by: Liu, Jiahao, et al.
Published: (2024)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
by: Zong, Qing, et al.
Published: (2024)
by: Zong, Qing, et al.
Published: (2024)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024)
by: Lai, Viet Dac, et al.
Published: (2024)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
by: Dineen, Jacob, et al.
Published: (2025)
by: Dineen, Jacob, et al.
Published: (2025)
Diversify, Rationalize, and Combine: Ensembling Multiple QA Strategies for Zero-shot Knowledge-based VQA
by: Li, Miaoyu, et al.
Published: (2024)
by: Li, Miaoyu, et al.
Published: (2024)
DebateQA: Evaluating Question Answering on Debatable Knowledge
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance
by: Fan, Yuchun, et al.
Published: (2026)
by: Fan, Yuchun, et al.
Published: (2026)
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
by: Huang, Tiancheng, et al.
Published: (2025)
by: Huang, Tiancheng, et al.
Published: (2025)
PaperHelper: Knowledge-Based LLM QA Paper Reading Assistant
by: Yin, Congrui, et al.
Published: (2025)
by: Yin, Congrui, et al.
Published: (2025)
KET-QA: A Dataset for Knowledge Enhanced Table Question Answering
by: Hu, Mengkang, et al.
Published: (2024)
by: Hu, Mengkang, et al.
Published: (2024)
DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts
by: Lu, Yujing, et al.
Published: (2025)
by: Lu, Yujing, et al.
Published: (2025)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
by: Bonomo, Tommaso, et al.
Published: (2025)
by: Bonomo, Tommaso, et al.
Published: (2025)
ExpertGenQA: Open-ended QA generation in Specialized Domains
by: Shahgir, Haz Sameen, et al.
Published: (2025)
by: Shahgir, Haz Sameen, et al.
Published: (2025)
MedConceptsQA: Open Source Medical Concepts QA Benchmark
by: Shoham, Ofir Ben, et al.
Published: (2024)
by: Shoham, Ofir Ben, et al.
Published: (2024)
Enhancing Food-Domain Question Answering with a Multimodal Knowledge Graph: Hybrid QA Generation and Diversity Analysis
by: B, Srihari K, et al.
Published: (2025)
by: B, Srihari K, et al.
Published: (2025)
Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT
by: Liu, Yesheng, et al.
Published: (2025)
by: Liu, Yesheng, et al.
Published: (2025)
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
by: Ruan, Junhao, et al.
Published: (2026)
by: Ruan, Junhao, et al.
Published: (2026)
FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
by: Kirchenbauer, John, et al.
Published: (2025)
by: Kirchenbauer, John, et al.
Published: (2025)
ChatQA: Surpassing GPT-4 on Conversational QA and RAG
by: Liu, Zihan, et al.
Published: (2024)
by: Liu, Zihan, et al.
Published: (2024)
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
by: Schimanski, Tobias, et al.
Published: (2026)
by: Schimanski, Tobias, et al.
Published: (2026)
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
by: Weck, Benno, et al.
Published: (2026)
by: Weck, Benno, et al.
Published: (2026)
Syn-QA2: Evaluating False Assumptions in Long-tail Questions with Synthetic QA Datasets
by: Daswani, Ashwin, et al.
Published: (2024)
by: Daswani, Ashwin, et al.
Published: (2024)
EffiQA: Efficient Question-Answering with Strategic Multi-Model Collaboration on Knowledge Graphs
by: Dong, Zixuan, et al.
Published: (2024)
by: Dong, Zixuan, et al.
Published: (2024)
PRIV-QA: Privacy-Preserving Question Answering for Cloud Large Language Models
by: Li, Guangwei, et al.
Published: (2025)
by: Li, Guangwei, et al.
Published: (2025)
Similar Items
-
Large-Scale Diverse Synthesis for Mid-Training
by: Zhang, Xuemiao, et al.
Published: (2025) -
Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
by: Zhang, Xuemiao, et al.
Published: (2025) -
A Survey on LLM Mid-Training
by: Tu, Chengying, et al.
Published: (2025) -
FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training
by: Xu, Liangyu, et al.
Published: (2025) -
FRAME: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy
by: Zhang, Xuemiao, et al.
Published: (2025)