I found this paper to be thought-provoking: "Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling" by Bansal, Hosseini, Agarwal, Tran, and Kazemi.
https://arxiv.org/abs/2408.16737
The direct implication is that smaller models could be used to create cost-effective synthetic datasets. And on that note, in the Gemma terms of use, Google explicitly claims no rights on outputs generated from those models, which means one is free to synthgen from the Gemma line. Meta's Llama 3 licence forbids synthetic generation of outputs if used to improve other models. Relevant Mistral, Qwen, and Yi models under the Apache 2.0 license are unrestricted for this purpose.

2 replies

liked a model 7 months ago

anthracite-org/magnum-v3-34b

Text Generation • Updated Oct 23, 2024 • 109 • 29

liked 2 models 8 months ago

grimjim/Mistral-Nemo-Instruct-2407-12B-6.4bpw-exl2

Text Generation • Updated Aug 25, 2024 • 41 • 5

ZeusLabs/L3-Aethora-15B-V2

Text Generation • Updated Jul 24, 2024 • 152 • 41