Synthesis by Design: Controlled Data Generation via Structural Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Lei, Chen, Sirui, Huang, Yuxuan, Lu, Chaochao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension Ability
by: Han, Yujin, et al.
Published: (2024)
by: Han, Yujin, et al.
Published: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
by: Chen, Junqi, et al.
Published: (2026)
by: Chen, Junqi, et al.
Published: (2026)
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
by: Chen, Sirui, et al.
Published: (2026)
by: Chen, Sirui, et al.
Published: (2026)
DEPO: Dual-Efficiency Preference Optimization for LLM Agents
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search
by: Zhang, Yize, et al.
Published: (2025)
by: Zhang, Yize, et al.
Published: (2025)
Reasoning Factual Knowledge in Structured Data with Large Language Models
by: Huang, Sirui, et al.
Published: (2024)
by: Huang, Sirui, et al.
Published: (2024)
Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance
by: Jiang, Yuxuan, et al.
Published: (2026)
by: Jiang, Yuxuan, et al.
Published: (2026)
KALE: Enhancing Knowledge Manipulation in Large Language Models via Knowledge-aware Learning
by: Lv, Qitan, et al.
Published: (2026)
by: Lv, Qitan, et al.
Published: (2026)
CLEAR: Can Language Models Really Understand Causal Graphs?
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation
by: Wang, Sirui, et al.
Published: (2025)
by: Wang, Sirui, et al.
Published: (2025)
Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems
by: He, Xi, et al.
Published: (2025)
by: He, Xi, et al.
Published: (2025)
DStruct2Design: Data and Benchmarks for Data Structure Driven Generative Floor Plan Design
by: Luo, Zhi Hao, et al.
Published: (2024)
by: Luo, Zhi Hao, et al.
Published: (2024)
Causal Evaluation of Language Models
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
ADAM: An Embodied Causal Agent in Open-World Environments
by: Yu, Shu, et al.
Published: (2024)
by: Yu, Shu, et al.
Published: (2024)
Generalizing From Short to Long: Effective Data Synthesis for Long-Context Instruction Tuning
by: Zhu, Wenhao, et al.
Published: (2025)
by: Zhu, Wenhao, et al.
Published: (2025)
Meaningful Learning: Enhancing Abstract Reasoning in Large Language Models via Generic Fact Guidance
by: Xiong, Kai, et al.
Published: (2024)
by: Xiong, Kai, et al.
Published: (2024)
TARGA: Targeted Synthetic Data Generation for Practical Reasoning over Structured Data
by: Huang, Xiang, et al.
Published: (2024)
by: Huang, Xiang, et al.
Published: (2024)
MDBench: A Synthetic Multi-Document Reasoning Benchmark Generated with Knowledge Guidance
by: Peper, Joseph J., et al.
Published: (2025)
by: Peper, Joseph J., et al.
Published: (2025)
Orca: Enhancing Role-Playing Abilities of Large Language Models by Integrating Personality Traits
by: Huang, Yuxuan
Published: (2024)
by: Huang, Yuxuan
Published: (2024)
Evaluating and Enhancing Large Language Models for Conversational Reasoning on Knowledge Graphs
by: Huang, Yuxuan
Published: (2023)
by: Huang, Yuxuan
Published: (2023)
Towards Compositional Generalization of LLMs via Skill Taxonomy Guided Data Synthesis
by: Wei, Yifan, et al.
Published: (2026)
by: Wei, Yifan, et al.
Published: (2026)
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
by: Ren, Yanwei, et al.
Published: (2026)
by: Ren, Yanwei, et al.
Published: (2026)
TrustUQA: A Trustful Framework for Unified Structured Data Question Answering
by: Zhang, Wen, et al.
Published: (2024)
by: Zhang, Wen, et al.
Published: (2024)
RGD: Multi-LLM Based Agent Debugger via Refinement and Generation Guidance
by: Jin, Haolin, et al.
Published: (2024)
by: Jin, Haolin, et al.
Published: (2024)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
by: Li, Jialu, et al.
Published: (2025)
by: Li, Jialu, et al.
Published: (2025)
CADDesigner: Conceptual CAD Model Generation with a General-Purpose Agent
by: Fan, Fengxiao, et al.
Published: (2025)
by: Fan, Fengxiao, et al.
Published: (2025)
HES-SQL: Hybrid Reasoning for Efficient Text-to-SQL with Structural Skeleton Guidance
by: Qiu, Suming, et al.
Published: (2025)
by: Qiu, Suming, et al.
Published: (2025)
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
by: Dong, Guanting, et al.
Published: (2026)
by: Dong, Guanting, et al.
Published: (2026)
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Decoupled Reasoning with Implicit Fact Tokens (DRIFT): A Dual-Model Framework for Efficient Long-Context Inference
by: Xie, Wenxuan, et al.
Published: (2026)
by: Xie, Wenxuan, et al.
Published: (2026)
TPD: Enhancing Student Language Model Reasoning via Principle Discovery and Guidance
by: Wang, Haorui, et al.
Published: (2024)
by: Wang, Haorui, et al.
Published: (2024)
Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis
by: Yuan, Lin, et al.
Published: (2025)
by: Yuan, Lin, et al.
Published: (2025)
Unleashing Scientific Reasoning for Bio-experimental Protocol Generation via Structured Component-based Reward Mechanism
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
MedGellan: LLM-Generated Medical Guidance to Support Physicians
by: Banerjee, Debodeep, et al.
Published: (2025)
by: Banerjee, Debodeep, et al.
Published: (2025)
The 2nd FutureDial Challenge: Dialog Systems with Retrieval Augmented Generation (FutureDial-RAG)
by: Cai, Yucheng, et al.
Published: (2024)
by: Cai, Yucheng, et al.
Published: (2024)
MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
by: Zhao, Xingjian, et al.
Published: (2025)
by: Zhao, Xingjian, et al.
Published: (2025)
BetterV: Controlled Verilog Generation with Discriminative Guidance
by: Pei, Zehua, et al.
Published: (2024)
by: Pei, Zehua, et al.
Published: (2024)
Automated Item Neutralization for Non-Cognitive Scales: A Large Language Model Approach to Reducing Social-Desirability Bias
by: Wu, Sirui, et al.
Published: (2025)
by: Wu, Sirui, et al.
Published: (2025)
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
by: Huang, Yiming, et al.
Published: (2024)
by: Huang, Yiming, et al.
Published: (2024)
Similar Items
-
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
by: Peng, Bo, et al.
Published: (2026) -
Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension Ability
by: Han, Yujin, et al.
Published: (2024) -
Can Post-Training Transform LLMs into Causal Reasoners?
by: Chen, Junqi, et al.
Published: (2026) -
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
by: Chen, Sirui, et al.
Published: (2026) -
DEPO: Dual-Efficiency Preference Optimization for LLM Agents
by: Chen, Sirui, et al.
Published: (2025)