Inference Providers
Active filters: grpo
jeremierostan/smol-question0
Updated
kapfy78/qwen-2.5-3b-r1-countdown
Text Generation
• 1B • Updated • 16
• susumuota/Qwen2.5-1.5B-Open-R1-GRPO
Text Generation
• 2B • Updated • 10
Text Generation
• Updated • 20
• 3
mradermacher/prem-1B-SQL-grpo-GGUF
1B • Updated • 72
mradermacher/Qwen-CoC-GRPO-GGUF
2B • Updated • 132
khuang2/qwen-2.5-3b-r1-countdown_vdebug
Text Generation
• 3B • Updated • 5
gnokit/Qwen2.5-Reasoning-3B-Instruct
3B • Updated • 9
Text Generation
• Updated • 24
chirag12/Qwen2-0.5B-GRPO-test
Updated
bushou/Qwen2.5-1.5B-Open-R1-GRPO
Text Generation
• 2B • Updated • 6
umarigan/llama-3.2-8B-R1-Tr
Text Generation
• 8B • Updated • 7
• 1
valerielucro/Qwen2-0.5B-GRPO_dummy
Text Generation
• 2.43M • Updated • 7
fhai50032/Qwen2.5-GRPO-7B
Text Generation
• Updated • 38
hyunseoki/llama3.2-1b-Open-R1-GRPO-test0
Text Generation
• 1B • Updated • 14
• 1
saswatach/Qwen2-0.5B-GRPO-test
Updated
pmking27/SaishamMathLM-3B-R1
Text Generation
• Updated • 4
BleachNick/Llama-3.2-1B-Instruct-GRPO-45k_RAGv1.5
Text Generation
• 1B • Updated • 12
• saswatach/qwen-r1-aha-moment
Updated
januverma/QwenMath0.5B_GRPO
Text Generation
• 0.5B • Updated • 6
• 1
chinmaydk99/Qwen2.5-0.5b-GRPO-math
Text Generation
• 0.5B • Updated • 11
• 1
kenhktsui/Qwen-0.5B-GRPO-gsm8k-count-wait-cap-cross-correct
Text Generation
• 0.5B • Updated • 6
• suayptalha/ThinkerLlama-8B-v1
Text Generation
• 8B • Updated • 7
• 3
Trojanssafsdg/phi4_merged_16bit
Text Generation
• 15B • Updated • 7
totalyielddot/llama3.1_reasoning
Text Generation
• 8B • Updated • 4
mradermacher/Qwen2.5-1.5B-R1-GRPO-GGUF
2B • Updated • 92
jayasuryajsk/Qwen2.5-3B-reasoner
Text Generation
• 3B • Updated • 7
pashocles/qwen-2.5-3b-r1-countdown
Text Generation
• 3B • Updated • 6
mesbahuddin1989/Qwen-0.5B-GRPO
Text Generation
• 0.5B • Updated • 5
hooman650/MedQwen3B-Reasoner
Text Generation
• 3B • Updated • 313
• 13