open-r1
/

OlympicCoder-32B

Text Generation

text-generation-inference

Model card Files Files and versions

lewtun HF Staff commited on Mar 13

Commit

3d7ab14

·

verified ·

1 Parent(s): d53a81c

Update README.md

Files changed (1) hide show

README.md +6 -0

README.md CHANGED Viewed

@@ -13,6 +13,9 @@ pipeline_tag: text-generation
 OlympicCoder-32B is a code mode that achieves very strong performance on competitive coding benchmarks such as LiveCodeBench andthe 2024 International Olympiad in Informatics.
 ## Model description
 - **Model type:** A 32B parameter model fine-tuned on a decontaminated version of the codeforces dataset.
@@ -52,6 +55,9 @@ print(outputs[0]["generated_text"])
 #<think>Okay, I need to write a Python program that calculates the 10th Fibonacci number. Hmm, the Fibonacci sequence starts with 0 and 1. Each subsequent number is the sum of the two preceding ones. So the sequence goes: 0, 1, 1, 2, 3, 5, 8, 13, 21, 34, and so on. ...
 ```
 ## Training procedure
 ### Training hyper-parameters

 OlympicCoder-32B is a code mode that achieves very strong performance on competitive coding benchmarks such as LiveCodeBench andthe 2024 International Olympiad in Informatics.
+* Repository: https://github.com/huggingface/open-r1
+* Blog post: https://huggingface.co/blog/open-r1/update-3
 ## Model description
 - **Model type:** A 32B parameter model fine-tuned on a decontaminated version of the codeforces dataset.
 #<think>Okay, I need to write a Python program that calculates the 10th Fibonacci number. Hmm, the Fibonacci sequence starts with 0 and 1. Each subsequent number is the sum of the two preceding ones. So the sequence goes: 0, 1, 1, 2, 3, 5, 8, 13, 21, 34, and so on. ...
 ```
+> [!IMPORTANT]
+> To ensure that the model consistently outputs a long chain-of-thought, we have edited the chat template to prefill the first assistant turn with a `<think>` token. As a result, the outputs from this model will not show the opening `<think>` token if you use the model's `generate()` method.  To apply reinforcement learning with a format reward, either prepend the `<think>` token to the model's completions or amend the chat template to remove the prefill. Check out our [blog post](https://huggingface.co/blog/open-r1/update-3#lesson-4-prefill-with-think-to-consistently-enable-long-cot) for more details.
 ## Training procedure
 ### Training hyper-parameters