AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl 0.4B • Updated Apr 30 • 2
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_propsft_propprefix_nokl 0.4B • Updated Apr 30 • 2
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-184_eval-dataset Viewer • Updated May 1 • 6.45k • 15
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-26_eval-dataset Viewer • Updated May 1 • 6.45k • 12
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-78_eval-dataset Viewer • Updated May 1 • 6.45k • 14
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-52_eval-dataset Viewer • Updated May 1 • 6.45k • 16
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-104_eval-dataset Viewer • Updated May 1 • 6.45k • 22
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-255_eval-dataset Viewer • Updated May 1 • 6.45k • 8
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft_prefix_nokl_checkpoint-255_eval-dataset Viewer • Updated Apr 30 • 6.45k • 5
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_checkpoint-255_eval-dataset Viewer • Updated Apr 30 • 6.45k • 10
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_checkpoint-104_eval-dataset Viewer • Updated Apr 30 • 6.45k • 18
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_checkpoint-52_eval-dataset Viewer • Updated Apr 30 • 6.45k • 4