Warm-started Checkpoints Collection A collection of three models trained on the Nemotron Post Training Dataset for reasoning tasks with IVON ⢠4 items ⢠Updated 3 days ago
Parameter Exploration for RLVR via Variational Learning Paper ⢠2608.09805 ⢠Published 17 days ago ⢠6
Parameter Exploration for RLVR via Variational Learning Paper ⢠2608.09805 ⢠Published 17 days ago ⢠6
3PO Models Collection 3PO family methods trained on DapoMath-17k using Olmo3-IVON-SFT-7B and Qwen2.5Math-IVON-SFT-7B ⢠10 items ⢠Updated 15 days ago
Parameter Exploration for RLVR via Variational Learning Paper ⢠2608.09805 ⢠Published 17 days ago ⢠6