Inference Providers
Active filters: grpo
mradermacher/DeepSeek-R1-Qwen-2.5-1.5b-Latest-Unstructured-To-Structured-GGUF
2B • Updated • 1.42k
• 2
hyunw3/qwen-2.5-0.5b-r1-countdown_lr5e-6
Text Generation
• 0.5B • Updated • 12
khuang2/qwen-2.5-3b-r1-countdown
Text Generation
• 3B • Updated • 6
• 2
spinech/qwen2.5-3b-r1-arc-train-thinker
Text Generation
• 3B • Updated • 4
• 1
Dongwei/DeepSeek-R1-Distill-Qwen-7B-GRPO_Math_lowlr
Text Generation
• 8B • Updated • 5
Dongwei/Qwen-2.5-7B_Math_smalllr
Text Generation
• 8B • Updated • 10
Dongwei/Qwen2.5-1.5B-Open-R1-GRPO_Math_smalllr
Text Generation
• 2B • Updated • 10
Dongwei/DeepSeek-R1-Distill-Qwen-1.5B-GRPO_Math_smalllr
Text Generation
• 2B • Updated • 7
mradermacher/Qwen2.5-1.5B-Thinking-v1.1-GGUF
2B • Updated • 78
• 2
mradermacher/Qwen2.5-1.5B-Thinking-v1.1-i1-GGUF
2B • Updated • 136
• 1
Dongwei/Qwen-2.5-7B_Base_Math_smalllr
Text Generation
• 8B • Updated • 24
• 6
jeremierostan/qwen-guiding-question
Updated
May811/Qwen2.5-1.5B-Open-R1-GRPO
Text Generation
• 2B • Updated • 6
spinech/qwen2.5-3b-r1-arc-train-thinker-2
Text Generation
• 3B • Updated • 3
Dongwei/Qwen-2.5-7B_Base_Math_smallestlr
Text Generation
• 8B • Updated • 10
Dongwei/Qwen-2.5-7B_Base_Math_smallestlr_newdata
Text Generation
• 8B • Updated • 9
sohyunan/gemma-2-2b-it_controller-grpo
Text Generation
• 3B • Updated • 12
zzhang1987/Qwen2.5-VL-3B-Instruct-Open-R1-Distill
Image-Text-to-Text
• 4B • Updated • 9
rzhao17/qwen-2.5-3b-r1-countdown
Text Generation
• 3B • Updated • 5
Novaciano/Q5KM-Charcard-RP-1B-GRPO_MiniThinky-GGUF
Text Generation
• 1B • Updated • 9
• 2
schwamaths/Qwen2.5-1.5B-Open-R1-GRPO
Text Generation
• 2B • Updated • 5
Chris126/qwen-r1-aha-moment
Updated
Text Generation
• 0.1B • Updated • 10
ibndias/Qwen2.5-1.5B-Open-R1-GRPO1st
Text Generation
• 2B • Updated • 6
jdqqjr/Qwen2.5-0.5B-Open-R1-GRPO
Text Generation
• 0.6B • Updated • 5
khuang2/qwen-2.5-3b-r1-countdown-offline_query_gen
Text Generation
• 3B • Updated • 5
mradermacher/qwen-2.5-3b-r1-countdown-GGUF
3B • Updated • 53
• 1
mradermacher/prem-1B-grpo-GGUF
Reinforcement Learning
• 1B • Updated • 37
mradermacher/qwen2.5-3b-r1-arc-train-thinker-GGUF
3B • Updated • 51
• 1
khuang2/qwen-2.5-3b-r1-countdown-offline_query_gen_solvable_only
Text Generation
• 3B • Updated • 5