Inference Providers
Active filters: grpo
mradermacher/DeepSeek-R1-Qwen-2.5-1.5b-Latest-Unstructured-To-Structured-GGUF
2B • Updated • 1.3k
• 2
advaithc/Litespark-1.5B-IFT-Math-Openrs
Text Generation
• 2B • Updated • 35
• 1
Lazarus-Ai/ReAligned-Qwen3.5-27B
Text Generation
• 27B • Updated • 35
• • 3
mindlab-research/Macaron-A2UI-Tall
Text Generation
• Updated • 20
• 5
rparkr/LFM2.5-1.2B-Instruct-Coding
Text Generation
• 4B • Updated • 2.31k
• • 30
vadimbelsky/medgemma-4b-esi-triage-grpo-v1
Text Generation
• 4B • Updated • 16
• 1
mradermacher/Qwable-9B-Claude-Fable-5-StraTA-GGUF
Reinforcement Learning
• 9B • Updated • 329
• 2
mradermacher/Atomight-V2.5-1.7B-C1-GGUF
2B • Updated • 857
• 1
lokahq/Trinity-Mini-AI-Scientist
Text Generation
• Updated • 13
• 1
MaliDDD/ds-medqa-9b-grpo-specific
Text Generation
• 9B • Updated • 49
• 1
Spreadsheet-RL/Spreadsheet-RL-8B
Text Generation
• 8B • Updated • 59
• 1
TengfeiLiuCoder/RefCaptioner
Image-Text-to-Text
• 9B • Updated • 49
• 1
WuuuuJH/ChronusOmni-Purned-7.98b-recover
Video-Text-to-Text
• Updated • 1
Chun121/Qwen3-4B-RPG-Roleplay-V2
Text Generation
• 4B • Updated • 17.5k
• 65
Text Generation
• 0.1B • Updated • 15
8B • Updated • 8
sergiopaniego/Qwen2-0.5B-GRPO-test
Updated
Novaciano/ESP-NSFW-GRPO-1B-Sin_Censura-GGUF
1B • Updated • 104
• 6
nbd22/Llama-3.1-8B-Instruct-GRPO-gsm8k-ft-lora
Updated
sergiopaniego/Qwen2-0.5B-GRPO
Updated
philschmid/qwen-2.5-3b-r1-countdown
Text Generation
• 3B • Updated • 12
• 8
spinech/qwen-2.5-3b-r1-countdown
Text Generation
• 3B • Updated • 13
Dongwei/Qwen2.5-1.5B-Open-R1-GRPO
Text Generation
• 2B • Updated • 8
• 1
spinech/qwen2.5-3b-r1-rearc-stage1
Text Generation
• 3B • Updated • 8
Dongwei/DeepSeek-R1-Distill-Qwen-7B-GRPO
Text Generation
• 8B • Updated • 15
• 1
MasterControlAIML/DeepSeek-R1-Strategy-Qwen-2.5-1.5b-Unstructured-To-Structured
Text Generation
• 2B • Updated • 30
• 5
mradermacher/DeepSeek-R1-Strategy-Qwen-2.5-1.5b-Unstructured-To-Structured-GGUF
2B • Updated • 88
• 2