Nellyw888/VeriReason-Qwen2.5-7b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • Updated 9 days ago • 1.36k • 1
vegeta03/DeepHermes-Egregore-v2-RLAIF-8b-Atropos-Q8_0-GGUF Reinforcement Learning • Updated 20 days ago • 26
vegeta03/DeepHermes-AscensionMaze-RLAIF-8b-Atropos-Q8_0-GGUF Reinforcement Learning • Updated 20 days ago • 23
vegeta03/DeepHermes-ToolCalling-Specialist-Atropos-Q8_0-GGUF Reinforcement Learning • Updated 20 days ago • 23
ajagota71/pythia-70m-detox-irl-rlhf-test-facebook-filter Reinforcement Learning • Updated 18 days ago • 2
Nellyw888/VeriReason-codeLlama-7b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • Updated 9 days ago • 1.37k
Nellyw888/VeriReason-Qwen2.5-3b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • Updated 9 days ago • 24
Nellyw888/VeriReason-Qwen2.5-1.5b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • Updated 9 days ago • 21 • 1