HumanTuring-DPO-8B-Instruct

This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.

Informations

Thanks for 1k downloads!

This model is trained by my smolturing model [which is a smollumi model trained with turing test dataset] with this dataset:

Trained by unsloth in Google Colab

GGUF

Model size

8.03B params

Architecture

llama

4-bit

Inference Providers NEW

This model is not currently available via any of the supported Inference Providers.

The model cannot be deployed to the HF Inference API: The model has no pipeline_tag.

Base model

Finetuned

Quantized

Quantized

(2)

this model