HPAI-BSC
/

Meta-Llama-3.1-8B-Instruct-Egida-DPO

Model card Files Files and versions Community

danihinjos commited on Mar 4

Commit

a2697e2

·

verified ·

1 Parent(s): a1362bc

Update README.md

Files changed (1) hide show

README.md +16 -0

README.md CHANGED Viewed

@@ -16,6 +16,22 @@ This is a fine-tuned Llama-3.1-8B-Instruct model on the [Egida-DPO-Llama-3.1-8B-
 The [Egida](https://huggingface.co/datasets/HPAI-BSC/Egida/viewer/Egida?views%5B%5D=egida_full) dataset is a collection of adversarial prompts that are thought to ellicit unsafe behaviors from language models. Specifically for this case, the Egida train split is used to run inference on Llama-3.1-70B-Instruct. Unsafe answers are selected, and paired with safe answers to create a customized DPO
 dataset for this model. This results in a DPO dataset composed by triplets < ”question”, ”chosen answer”, ”discarded answer” > which contain questions that elicit unsafe responses by this target model, as well as the unsafe responses produced by it.
 ## Training Details
 - **Hardware:** NVIDIA H100 64 GB GPUs

 The [Egida](https://huggingface.co/datasets/HPAI-BSC/Egida/viewer/Egida?views%5B%5D=egida_full) dataset is a collection of adversarial prompts that are thought to ellicit unsafe behaviors from language models. Specifically for this case, the Egida train split is used to run inference on Llama-3.1-70B-Instruct. Unsafe answers are selected, and paired with safe answers to create a customized DPO
 dataset for this model. This results in a DPO dataset composed by triplets < ”question”, ”chosen answer”, ”discarded answer” > which contain questions that elicit unsafe responses by this target model, as well as the unsafe responses produced by it.
+## Performance
+### Safety Performance (Attack Success Ratio)
+|                              | Egida (test) ↓ | DELPHI ↓ | Alert-Base ↓ | Alert-Adv ↓ |
+|------------------------------|:--------------:|:--------:|:------------:|:-----------:|
+| Meta-Llama-3.1-8B-Instruct   |     0.347      |  0.160   |    0.446     |    0.039    |
+| Meta-Llama-3.1-8B-Egida-DPO  |     0.038      |  0.025   |    0.038     |    0.014    |
+### General Purpose Performance
+|                              | OpenLLM Leaderboard (Average) ↑ | MMLU (ROUGE1) ↑ |
+|------------------------------|:---------------------:|:---------------:|
+| Meta-Llama-3.1-8B-Instruct   |         0.453         |      0.646      |
+| Meta-Llama-3.1-8B-Egida-DPO  |         0.453         |      0.643      |
 ## Training Details
 - **Hardware:** NVIDIA H100 64 GB GPUs