CorrSynth -- A Correlated Sampling Method for Diverse Dataset Generation from LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Kowshik, Suhas S, Divekar, Abhishek, Malik, Vijit |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
by: Divekar, Abhishek, et al.
Published: (2024)
by: Divekar, Abhishek, et al.
Published: (2024)
PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation
by: Divekar, Abhishek, et al.
Published: (2026)
by: Divekar, Abhishek, et al.
Published: (2026)
MetaSynth: Meta-Prompting-Driven Agentic Scaffolds for Diverse Synthetic Data Generation
by: Riaz, Haris, et al.
Published: (2025)
by: Riaz, Haris, et al.
Published: (2025)
CasualSynth: Generating Structurally Sound Synthetic Data
by: Cheng, Zehua, et al.
Published: (2026)
by: Cheng, Zehua, et al.
Published: (2026)
Transformer Block Coupling and its Correlation with Generalization in LLMs
by: Aubry, Murdock, et al.
Published: (2024)
by: Aubry, Murdock, et al.
Published: (2024)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
DiffSampling: Enhancing Diversity and Accuracy in Neural Text Generation
by: Franceschelli, Giorgio, et al.
Published: (2025)
by: Franceschelli, Giorgio, et al.
Published: (2025)
SynthAgent: Adapting Web Agents with Synthetic Supervision
by: Wang, Zhaoyang, et al.
Published: (2025)
by: Wang, Zhaoyang, et al.
Published: (2025)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
MisSynth: Improving MISSCI Logical Fallacies Classification with Synthetic Data
by: Poliakov, Mykhailo, et al.
Published: (2025)
by: Poliakov, Mykhailo, et al.
Published: (2025)
The Price of Format: Diversity Collapse in LLMs
by: Yun, Longfei, et al.
Published: (2025)
by: Yun, Longfei, et al.
Published: (2025)
Sample-Efficient Alignment for LLMs
by: Liu, Zichen, et al.
Published: (2024)
by: Liu, Zichen, et al.
Published: (2024)
When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges
by: Darshan, Parth, et al.
Published: (2026)
by: Darshan, Parth, et al.
Published: (2026)
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking
by: Kulkarni, Atharva, et al.
Published: (2024)
by: Kulkarni, Atharva, et al.
Published: (2024)
Reframing Data Value for Large Language Models Through the Lens of Plausibility
by: Rammal, Mohamad Rida, et al.
Published: (2024)
by: Rammal, Mohamad Rida, et al.
Published: (2024)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
by: Liu, Wanlong, et al.
Published: (2024)
by: Liu, Wanlong, et al.
Published: (2024)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
by: Kossen, Jannik, et al.
Published: (2024)
by: Kossen, Jannik, et al.
Published: (2024)
Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets
by: Indurthi, Sathish Reddy, et al.
Published: (2024)
by: Indurthi, Sathish Reddy, et al.
Published: (2024)
Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
by: Troshin, Sergey, et al.
Published: (2025)
by: Troshin, Sergey, et al.
Published: (2025)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
by: Laine, Rudolf, et al.
Published: (2024)
by: Laine, Rudolf, et al.
Published: (2024)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
by: Liang, Buyun, et al.
Published: (2025)
by: Liang, Buyun, et al.
Published: (2025)
MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization
by: Wang, Ziqing, et al.
Published: (2026)
by: Wang, Ziqing, et al.
Published: (2026)
FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs
by: Bhattacharya, Debarpan, et al.
Published: (2025)
by: Bhattacharya, Debarpan, et al.
Published: (2025)
APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
by: Liu, Zuxin, et al.
Published: (2024)
by: Liu, Zuxin, et al.
Published: (2024)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
by: Liu, Mingjie, et al.
Published: (2025)
by: Liu, Mingjie, et al.
Published: (2025)
Mitigating Spurious Correlations in NLI via LLM-Synthesized Counterfactuals and Dynamic Balanced Sampling
by: Jaimes, Christopher Román
Published: (2025)
by: Jaimes, Christopher Román
Published: (2025)
Evaluation of Large Language Models via Coupled Token Generation
by: Benz, Nina Corvelo, et al.
Published: (2025)
by: Benz, Nina Corvelo, et al.
Published: (2025)
When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
by: Wang, Shaowen, et al.
Published: (2025)
by: Wang, Shaowen, et al.
Published: (2025)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
by: Sinha, Yash, et al.
Published: (2024)
by: Sinha, Yash, et al.
Published: (2024)
CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities
by: Mao, Yujun, et al.
Published: (2024)
by: Mao, Yujun, et al.
Published: (2024)
Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data
by: Ling, Zhenqing, et al.
Published: (2025)
by: Ling, Zhenqing, et al.
Published: (2025)
QuIM-RAG: Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance
by: Saha, Binita, et al.
Published: (2025)
by: Saha, Binita, et al.
Published: (2025)
EUROPA: A Legal Multilingual Keyphrase Generation Dataset
by: Salaün, Olivier, et al.
Published: (2024)
by: Salaün, Olivier, et al.
Published: (2024)
Instruction Diversity Drives Generalization To Unseen Tasks
by: Zhang, Dylan, et al.
Published: (2024)
by: Zhang, Dylan, et al.
Published: (2024)
Diversity Boosts AI-Generated Text Detection
by: Basani, Advik Raj, et al.
Published: (2025)
by: Basani, Advik Raj, et al.
Published: (2025)
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
by: Hu, Haiquan, et al.
Published: (2025)
by: Hu, Haiquan, et al.
Published: (2025)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
by: Chen, Justin Chih-Yao, et al.
Published: (2023)
by: Chen, Justin Chih-Yao, et al.
Published: (2023)
Similar Items
-
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
by: Divekar, Abhishek, et al.
Published: (2024) -
PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation
by: Divekar, Abhishek, et al.
Published: (2026) -
MetaSynth: Meta-Prompting-Driven Agentic Scaffolds for Diverse Synthetic Data Generation
by: Riaz, Haris, et al.
Published: (2025) -
CasualSynth: Generating Structurally Sound Synthetic Data
by: Cheng, Zehua, et al.
Published: (2026) -
Transformer Block Coupling and its Correlation with Generalization in LLMs
by: Aubry, Murdock, et al.
Published: (2024)