zephyr_0.1_a8.0

The extrapolated (ExPO) model based on chujiezheng/zephyr_0.1 and alignment-handbook/zephyr-7b-sft-full, as in the "Weak-to-Strong Extrapolation Expedites Alignment" paper.

Specifically, we obtain this model by extrapolating from the weights of the SFT and DPO/RLHF checkpoints, achieving superior alignment with human preference.

Downloads last month: 4

Safetensors

Model size

7.24B params

Tensor type

BF16

Inference Providers NEW

Text Generation

This model is not currently available via any of the supported third-party Inference Providers, and the model is not deployed on the HF Inference API.

Collection including chujiezheng/zephyr_0.1_a8.0

Model Checkpoints in the ExPO Paper

Collection

15 items • Updated May 19, 2024 • 2