Open Source Ecosystems

This is the official repository for "Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing". Magpie generates high-quality alignment data by prompting aligned LLMs with their pre-query templates. Unlike many existing synthetic data generation methods, Magpie doesn't rely on prompt engineering or seed questions for generating synthetic data. Instead, it uses the prompt template of an aligned LLM to generate both the user query and an LLM response.

🤗 Huggingface (Models and Datasets)
🧭 Dataset Navigation
🕸️ Website
📄 Technical Report
🤗 Magpie Demo (Thanks a lot for the implementation from @davanstrien!)
🐦 Chat with Magpie

🐦 News

[2024/09/17] Ship two new models with SOTA performance: 𝙼𝚊𝚐𝚙𝚒𝚎𝙻𝙼-𝙲𝚑𝚊𝚝 (4B & 8B)! See collection here!
[2024/08/19] Three preference optimization datasets, Magpie-Air-DPO-100K-v0.1, Magpie-Pro-DPO-100K-v0.1, and Magpie-Llama-3.1-Pro-DPO-100K-v0.1 are out!
[2024/07/25] Magpie Llama-3.1 dataset is out! 1M from Meta-Llama-3.1-70B-Instruct! More friendly license compared with Llama-3 😃!
[2024/07/21] Magpie Gemma2 dataset is out! 534K from Gemma-2-27b-it!
[2024/07/19] Llama-3-8B-Magpie-Align-v0.3 is out with enhanced Chinese question-answering ability, thanks to our new Chinese instruction dataset!
[2024/07/14] Llama-3-8B-Magpie-Align-v0.2 is out with enhanced reasoning ability, thanks to our new reasoning booster dataset!
[2024/07/04] Magpie Qwen2 dataset is out! 1M from Qwen2 72B and 3M from Qwen2 7B.
[2024/07/03] 🏆 Our open aligned model, Llama-3-8B-Magpie-Align-v0.1 is out! It is 🏆 the best <30B Model in AI2 WildBench Leaderboard! Even better than the official Meta-Llama-3-8B-Instruct model!
[2024/06/24] Magpie Phi 3 dataset is out! 1M from Phi 3 Medium.
[2024/06/12] Magpie Llama-3 dataset is out! 1M from Llama-3 70B and 3M from Llama-3 8B.
[2024/06/12] Magpie technical report is out! Let's make high-quality alignment data open for all!

Magpie Supports

Currently, Magpie has been tested on the Llama-3, Qwen2, Phi 3 and Gemma-2 series. Please submit an issue for more model support.

Model Family	Magpie	Magpie Scripts	Datasets	Size	Note
Llama 3.1	⭕️	8B,70B	70B,405B(Argilla)	1M	Apply a logits processor to surpress markdown format.
Llama 3	✅	8B,70B	8B,70B	3M + 1M
Qwen2	✅	7B,72B,Math 7B	7B,72B	3M + 1M
Phi 3	✅	mini,small,medium	medium	1M
Gemma-2	⭕️	9B,27B	27B	534K	Apply a filter before generating responses.
Qwen2.5	⭕️	3B,7B,14B,32B,72B
Gemma-1.1	⭕️	7B
Llama 2	⭕️	7B,70B
Vicuna	⭕️	7B
Mistral	⭕️	7B
Yi	⭕️	34B
DeepSeek Coder	⭕️	Coder V2 Lite

✅: Works so great!
⭕️: It works! We can get something interesting, but we may need to apply an additional filter and/or a logit processor.
❌: Not work.
❓: Untested.

The navigation of all available Magpie datasets can be found here.

We hope Magpie can contribute to the democratization of AI with enhanced transparency of model alignment processes!

Abstract

Overview

Installation

Build environment

git clone https://github.com/magpie-align/magpie.git
cd magpie
conda create -n magpie python=3.10 -y
conda activate magpie
pip install -r requirements.txt

Get access to Llama-3 models from 🤗 Huggingface

You can apply for Llama-3 model access here. To login in the terminal, enter:

huggingface-cli login

then enter your Huggingface private key beginning with "hf_".

Toy Example

Play with Jupyter Notebook

The toy example can be found in demo.ipynb. Have fun!

Batched Data Generation

We use Llama-3-8B-Instruct as an example to demonstrate the batched data generation process. To run batched generation, you can simply run:

cd scripts
bash magpie.sh

The script will generate both instructions and responses in the data folder. It has been tested on an RTX 4090 24G GPU. If you are using GPUs with less memory, consider implementing quantization.

We also provide scripts for other models in the scripts folder. You can use this navigation to find specific Magpie scripts. Note that for model sizes greater than 8B, you may need 4*A100 GPUs to run the scripts.

Batched Multi-turn Data Generation [Optional]

After generating instruction-response pairs, you can extend them to multi-turn conversations. To do so, simply run the following command:

bash magpie-multi-turn.sh ***_ins_res.json

where ***_ins_res.json is the single-turn instruction-response pairs generated in the previous step.

Dataset Filtering

1. Tagging

To tag the generated instruction-response pairs, you can run:

cd scripts
bash unitag.sh ***_ins_res.json all

This script will automatically generate quality, difficulty, task category, safety, reward, and language for the generated dataset. You can also generate one tag at a time. For example, if you just want to generate the safety label using device 0, you can run:

cd scripts
bash unitag.sh ***_ins_res.json safety 0

2. Data Concatenation and Converting

You may generate datasets with different generation configurations. We provide a Jupyter notebook here for concatenating all datasets and converting them to ShareGPT format, which is fully supported by Axolotl for fine-tuning.

3. Removing Repetition

Once you have a full dataset converted to ShareGPT format, you can calculate the minimum neighbor distance of each instruction and remove repetitions. To do so, run:

cd exp
python gen_dis.py --input_file ***_sharegpt.jsonl

where ***_sharegpt.jsonl is the dataset path obtained in the previous step. The Python script will take care of building the FAISS index and calculating the minimum distance.

4. Design and Apply Your Filter

We provide a Jupyter notebook here for simple filtering. You can adjust the filtering parameters to design and apply your own filter based on your needs.

Fine-tuning

Please take a look at the recipes directory for instructions and our Magpie model recipes.

Citation

If you find the model, data, or code useful, please cite our paper 🤩:

@article{xu2024magpie,
  title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
  author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin},
  journal={ArXiv},
  year={2024},
  volume={abs/2406.08464},
  url={https://api.semanticscholar.org/CorpusID:270391432}
}

Star History

Badges

Extracted from project README

Related Projects

Baichuan-7B

A large-scale 7B pretraining language model developed by BaiChuan-Inc.

14 Jun 2023 5,670

LLaVA

[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and b...

17 Apr 2023 19,659

GPT-4-LLM

Instruction Tuning with GPT-4

06 Apr 2023 3,923

LLamaTuner

Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Fal...

25 May 2023 568

MedicalGPT

MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型，实现了包括增量预训...

02 Jun 2023 2,446

LLaVA-NeXT-Image-Llama3-Lora

LLaVA-NeXT-Image-Llama3-Lora, Modified from https://github.com/arielnlee/LLaVA-1.6-ft

24 Jun 2024 37

Chinese-Vicuna

Chinese-Vicuna: A Chinese Instruction-following LLaMA-based Model —— 一个中文低资源的llama+lora方案，结构参考alpaca

23 Mar 2023 4,144

safe-rlhf

Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback

15 May 2023 1,137

Mol-Instructions

[ICLR 2024] Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language M...

12 Apr 2023 237

data-juicer

A one-stop data processing system to make data higher-quality, juicier, and more digestible for (...

01 Aug 2023 2,315

airllm

AirLLM 70B inference with single 4GB GPU

12 Jun 2023 4,536

KoAlpaca

KoAlpaca: 한국어 명령어를 이해하는 오픈소스 언어모델

18 Mar 2023 1,460

Multimodal-GPT

26 Apr 2023 1,467

KnowLM

An Open-sourced Knowledgable Large Language Model Framework.

01 Apr 2023 980

textgen

TextGen: Implementation of Text Generation models, include LLaMA, BLOOM, GPT2, BART, T5, SongNet ...

07 Apr 2021 926