NousResearch/DeepHermes-Egregore-v2-RLAIF-8b-Atropos-GGUF Reinforcement Learning • 8B • Updated May 5, 2025 • 67 • 2
NousResearch/DeepHermes-Egregore-v1-RLAIF-8b-Atropos-GGUF Reinforcement Learning • 8B • Updated May 5, 2025 • 50 • 3
NousResearch/DeepHermes-AscensionMaze-RLAIF-8b-Atropos-GGUF Reinforcement Learning • 127k • Updated May 10, 2025 • 122 • 7
mradermacher/SEOcrate-4B_grpo_new_01-GGUF Reinforcement Learning • 4B • Updated Jul 11, 2025 • 79 • 1
mradermacher/SEOcrate-4B_grpo_new_01-i1-GGUF Reinforcement Learning • 4B • Updated Jul 11, 2025 • 163
Nellyw888/VeriReason-Qwen2.5-7b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • 8B • Updated May 31, 2025 • 32 • 4
vegeta03/DeepHermes-ToolCalling-Specialist-Atropos-Q8_0-GGUF Reinforcement Learning • 8B • Updated May 9, 2025 • 3
ajagota71/pythia-70m-detox-irl-rlhf-test-facebook-filter Reinforcement Learning • 70.4M • Updated May 11, 2025 • 4
ajagota71/pythia-70m-detox-raw-logits-test2 Reinforcement Learning • 70.4M • Updated May 11, 2025 • 4
ajagota71/pythia-160m-detox-raw-logits-test2 Reinforcement Learning • 0.2B • Updated May 11, 2025 • 4
ajagota71/pythia-70m-detox-irl-rlhf-seed-42 Reinforcement Learning • 70.4M • Updated May 11, 2025 • 4
ajagota71/pythia-70m-detox-irl-rlhf-seed-200 Reinforcement Learning • 70.4M • Updated May 11, 2025 • 4
ajagota71/pythia-410m-detox-irl-rlhf-seed-42 Reinforcement Learning • 0.4B • Updated May 11, 2025 • 4
ajagota71/pythia-410m-detox-irl-rlhf-seed-100 Reinforcement Learning • 0.4B • Updated May 11, 2025 • 5
ajagota71/pythia-410m-detox-irl-rlhf-seed-200 Reinforcement Learning • 0.4B • Updated May 11, 2025 • 2
ajagota71/pythia-410m-detox-irl-rlhf-seed-300 Reinforcement Learning • 0.4B • Updated May 12, 2025 • 4
ajagota71/pythia-410m-detox-irl-rlhf-seed-400 Reinforcement Learning • 0.4B • Updated May 12, 2025 • 3
Nellyw888/VeriReason-codeLlama-7b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • 7B • Updated May 31, 2025 • 48 • 2
Nellyw888/VeriReason-Qwen2.5-3b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • 3B • Updated May 31, 2025 • 499
Nellyw888/VeriReason-Qwen2.5-1.5b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • 2B • Updated May 20, 2025 • 35 • 1
ajagota71/pythia-70m-fb-detox-checkpoint-epoch-20 Reinforcement Learning • 70.4M • Updated May 16, 2025 • 2
ajagota71/pythia-70m-fb-detox-checkpoint-epoch-40 Reinforcement Learning • 70.4M • Updated May 16, 2025 • 3
ajagota71/pythia-70m-fb-detox-checkpoint-epoch-60 Reinforcement Learning • 70.4M • Updated May 16, 2025 • 3
ajagota71/pythia-70m-fb-detox-checkpoint-epoch-80 Reinforcement Learning • 70.4M • Updated May 16, 2025 • 3
ajagota71/pythia-70m-fb-detox-checkpoint-epoch-100 Reinforcement Learning • 70.4M • Updated May 16, 2025 • 3