Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Zhangchen, Jiang, Fengqing, Niu, Luyao, Deng, Yuntian, Poovendran, Radha, Choi, Yejin, Lin, Bill Yuchen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stronger Models are NOT Stronger Teachers for Instruction Tuning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
Temporal Sampling for Forgotten Reasoning in LLMs
von: Li, Yuetai, et al.
Veröffentlicht: (2025)
von: Li, Yuetai, et al.
Veröffentlicht: (2025)
Brave: Byzantine-Resilient and Privacy-Preserving Peer-to-Peer Federated Learning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
Small Models Struggle to Learn from Strong Reasoners
von: Li, Yuetai, et al.
Veröffentlicht: (2025)
von: Li, Yuetai, et al.
Veröffentlicht: (2025)
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
von: Feng, Yichen, et al.
Veröffentlicht: (2025)
von: Feng, Yichen, et al.
Veröffentlicht: (2025)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
von: Li, Yuetai, et al.
Veröffentlicht: (2024)
von: Li, Yuetai, et al.
Veröffentlicht: (2024)
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
von: Han, Seungju, et al.
Veröffentlicht: (2024)
von: Han, Seungju, et al.
Veröffentlicht: (2024)
Fault Tolerant Neural Control Barrier Functions for Robotic Systems under Sensor Faults and Attacks
von: Zhang, Hongchao, et al.
Veröffentlicht: (2024)
von: Zhang, Hongchao, et al.
Veröffentlicht: (2024)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
von: Xiang, Zhen, et al.
Veröffentlicht: (2024)
von: Xiang, Zhen, et al.
Veröffentlicht: (2024)
PromptTailor: Multi-turn Intent-Aligned Prompt Synthesis for Lightweight LLMs
von: Xu, Yizhou, et al.
Veröffentlicht: (2025)
von: Xu, Yizhou, et al.
Veröffentlicht: (2025)
WildChat: 1M ChatGPT Interaction Logs in the Wild
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
von: Li, Huihan, et al.
Veröffentlicht: (2024)
von: Li, Huihan, et al.
Veröffentlicht: (2024)
Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
The WidthWall: A Strict Expressivity Hierarchy for Hypergraph Neural Networks
von: Jiang, Fengqing, et al.
Veröffentlicht: (2026)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2026)
Polyhedral Instability Governs Regret in Online Learning
von: Li, Yuetai, et al.
Veröffentlicht: (2026)
von: Li, Yuetai, et al.
Veröffentlicht: (2026)
Distributed Safety-Critical Control of Multi-Agent Systems with Time-Varying Communication Topologies
von: Cheng, Shiyu, et al.
Veröffentlicht: (2026)
von: Cheng, Shiyu, et al.
Veröffentlicht: (2026)
Modeling and Designing Non-Pharmaceutical Interventions in Epidemics: A Submodular Approach
von: Cheng, Shiyu, et al.
Veröffentlicht: (2024)
von: Cheng, Shiyu, et al.
Veröffentlicht: (2024)
Swarm-STL: A Framework for Motion Planning in Large-Scale, Multi-Swarm Systems
von: Cheng, Shiyu, et al.
Veröffentlicht: (2025)
von: Cheng, Shiyu, et al.
Veröffentlicht: (2025)
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2024)
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2024)
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
von: Cao, Bochuan, et al.
Veröffentlicht: (2025)
von: Cao, Bochuan, et al.
Veröffentlicht: (2025)
Fake Alignment: Are LLMs Really Aligned Well?
von: Wang, Yixu, et al.
Veröffentlicht: (2023)
von: Wang, Yixu, et al.
Veröffentlicht: (2023)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
von: Zhao, Wenting, et al.
Veröffentlicht: (2024)
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
von: Deng, Yuntian, et al.
Veröffentlicht: (2024)
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Stronger Models are NOT Stronger Teachers for Instruction Tuning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024) -
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024) -
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024) -
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024) -
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)