Best Practices and Lessons Learned on Synthetic Data
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Ruibo, Wei, Jerry, Liu, Fangyu, Si, Chenglei, Zhang, Yanzhe, Rao, Jinmeng, Zheng, Steven, Peng, Daiyi, Yang, Diyi, Zhou, Denny, Dai, Andrew M. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
di: Si, Chenglei, et al.
Pubblicazione: (2024)
di: Si, Chenglei, et al.
Pubblicazione: (2024)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
di: Si, Chenglei, et al.
Pubblicazione: (2024)
di: Si, Chenglei, et al.
Pubblicazione: (2024)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
di: Si, Chenglei, et al.
Pubblicazione: (2025)
di: Si, Chenglei, et al.
Pubblicazione: (2025)
Higher Layers Need More LoRA Experts
di: Gao, Chongyang, et al.
Pubblicazione: (2024)
di: Gao, Chongyang, et al.
Pubblicazione: (2024)
Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
di: Chen, Qi, et al.
Pubblicazione: (2025)
di: Chen, Qi, et al.
Pubblicazione: (2025)
Searching for Privacy Risks in LLM Agents via Simulation
di: Zhang, Yanzhe, et al.
Pubblicazione: (2025)
di: Zhang, Yanzhe, et al.
Pubblicazione: (2025)
A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
di: Liu, Zijun, et al.
Pubblicazione: (2023)
di: Liu, Zijun, et al.
Pubblicazione: (2023)
Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
di: Yang, Diji, et al.
Pubblicazione: (2025)
di: Yang, Diji, et al.
Pubblicazione: (2025)
Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping
di: Li, Ryan, et al.
Pubblicazione: (2024)
di: Li, Ryan, et al.
Pubblicazione: (2024)
Attacking Vision-Language Computer Agents via Pop-ups
di: Zhang, Yanzhe, et al.
Pubblicazione: (2024)
di: Zhang, Yanzhe, et al.
Pubblicazione: (2024)
Towards Execution-Grounded Automated AI Research
di: Si, Chenglei, et al.
Pubblicazione: (2026)
di: Si, Chenglei, et al.
Pubblicazione: (2026)
Auditing Gender Presentation Differences in Text-to-Image Models
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
Long-form factuality in large language models
di: Wei, Jerry, et al.
Pubblicazione: (2024)
di: Wei, Jerry, et al.
Pubblicazione: (2024)
Distilling an End-to-End Voice Assistant Without Instruction Training Data
di: Held, William, et al.
Pubblicazione: (2024)
di: Held, William, et al.
Pubblicazione: (2024)
Highly Efficient and Stable Narrow Band Green Emitting Phosphor of Sb 3+ /Ce 3+ Sensitized Cs 2 NaTbCl 6 for WLED
di: Changheng Chen, et al.
Pubblicazione: (2024)
di: Changheng Chen, et al.
Pubblicazione: (2024)
Borderline content and platformised speech governance: Mapping TikTok's moderation controversies in South and Southeast Asia
di: Diyi Liu
Pubblicazione: (2024)
di: Diyi Liu
Pubblicazione: (2024)
Contextual Experience Replay for Self-Improvement of Language Agents
di: Liu, Yitao, et al.
Pubblicazione: (2025)
di: Liu, Yitao, et al.
Pubblicazione: (2025)
Type-Compliant Adaptation Cascades: Adapting Programmatic LM Workflows to Data
di: Lin, Chu-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Chu-Cheng, et al.
Pubblicazione: (2025)
The Best Instruction-Tuning Data are Those That Fit
di: Zhang, Dylan, et al.
Pubblicazione: (2025)
di: Zhang, Dylan, et al.
Pubblicazione: (2025)
Generative Interfaces for Language Models
di: Chen, Jiaqi, et al.
Pubblicazione: (2025)
di: Chen, Jiaqi, et al.
Pubblicazione: (2025)
Real-Time Reasoning Agents in Evolving Environments
di: Wen, Yule, et al.
Pubblicazione: (2025)
di: Wen, Yule, et al.
Pubblicazione: (2025)
Security and Innovation in ERP Systems: Best Practices for AI, OIC, and Automation Integration
di: Sreenivasa Rao Sola
Pubblicazione: (2023)
di: Sreenivasa Rao Sola
Pubblicazione: (2023)
Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
Challenges and Best Practices in Corporate AI Governance:Lessons from the Biopharmaceutical Industry
di: Mökander, Jakob, et al.
Pubblicazione: (2024)
di: Mökander, Jakob, et al.
Pubblicazione: (2024)
Defect‐Engineered Zero‐Dimensional Perovskite Cs 3 LuCl 6 : Tb 3+ Scintillator with Exceptional Thermal Stability for Flexible High‐Temperature X‐Ray Imaging
di: Ruibo Gao, et al.
Pubblicazione: (2026)
di: Ruibo Gao, et al.
Pubblicazione: (2026)
Achieving Single‐Phased Full Visible Spectrum Broadband White Emission in Ag⁺, Bi 3 ⁺, and Sb 3 ⁺ Tri‐Doped Cs₂NaLuCl₆ Double Perovskite Phosphor
di: Changheng Chen, et al.
Pubblicazione: (2025)
di: Changheng Chen, et al.
Pubblicazione: (2025)
SPHERE: An Evaluation Card for Human-AI Systems
di: Ma, Qianou, et al.
Pubblicazione: (2025)
di: Ma, Qianou, et al.
Pubblicazione: (2025)
Selecting the Best Optimizing System
di: Si, Nian, et al.
Pubblicazione: (2022)
di: Si, Nian, et al.
Pubblicazione: (2022)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
Simple synthetic data reduces sycophancy in large language models
di: Wei, Jerry, et al.
Pubblicazione: (2023)
di: Wei, Jerry, et al.
Pubblicazione: (2023)
When to Showcase Automated Production Processes? Disclosing Production Processes Increases Evaluation of Low‐End but Decreases Evaluation of High‐End Products
di: Diyi Liu, et al.
Pubblicazione: (2025)
di: Diyi Liu, et al.
Pubblicazione: (2025)
Contextualized Privacy Defense for LLM Agents
di: Wen, Yule, et al.
Pubblicazione: (2026)
di: Wen, Yule, et al.
Pubblicazione: (2026)
S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models
di: Lei, Fangyu, et al.
Pubblicazione: (2023)
di: Lei, Fangyu, et al.
Pubblicazione: (2023)
Robust Output Regulation of Uncertain Linear Time-Varying Systems
di: Zha, Jinmeng, et al.
Pubblicazione: (2026)
di: Zha, Jinmeng, et al.
Pubblicazione: (2026)
DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation
di: Zhu, Speed, et al.
Pubblicazione: (2025)
di: Zhu, Speed, et al.
Pubblicazione: (2025)
Learning Password Best Practices Through In-Task Instruction
di: Ma, Qian, et al.
Pubblicazione: (2026)
di: Ma, Qian, et al.
Pubblicazione: (2026)
Some Best Practices in Operator Learning
di: Enyeart, Dustin, et al.
Pubblicazione: (2024)
di: Enyeart, Dustin, et al.
Pubblicazione: (2024)
Applying Machine Learning Methods to Laser Acceleration of Protons: Lessons Learned from Synthetic Data
di: Desai, Ronak, et al.
Pubblicazione: (2023)
di: Desai, Ronak, et al.
Pubblicazione: (2023)
Applying Machine‐Learning Methods to Laser Acceleration of Protons: Lessons Learned From Synthetic Data
di: Ronak Desai, et al.
Pubblicazione: (2024)
di: Ronak Desai, et al.
Pubblicazione: (2024)
Relic abundance of dark matter with coannihilation in non-standard cosmological scenarios
di: Liu, Fangyu, et al.
Pubblicazione: (2023)
di: Liu, Fangyu, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
di: Si, Chenglei, et al.
Pubblicazione: (2024) -
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
di: Si, Chenglei, et al.
Pubblicazione: (2024) -
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
di: Si, Chenglei, et al.
Pubblicazione: (2025) -
Higher Layers Need More LoRA Experts
di: Gao, Chongyang, et al.
Pubblicazione: (2024) -
Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
di: Chen, Qi, et al.
Pubblicazione: (2025)