slm-500m-rlaif2
500m legal SLM: base -> QA SFT -> instruction SFT -> RLAIF v2. A Bradley-Terry reward model (warm-started from the instruction model) was trained on ~4k AI-written preference pairs spanning closed-book QA AND instruction-following failure modes; the policy was then improved by best-of-N reward-weighted SFT.
- Downloads last month
- 11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support