view article Article Navigating the RLHF Landscape: From Policy Gradients to PPO, GAE, and DPO for LLM Alignment NormalUhr • Feb 11, 2025 • 135
DatPySci/Self-Distillation-Qwen2.5-7B-Instruct-medreason-prefix-instruct-tokenized Viewer • Updated Jul 31 • 32.7k • 17
DatPySci/Self-Distillation-Qwen2.5-7B-Instruct-medreason-prefix-instruct-tokenized Viewer • Updated Jul 31 • 32.7k • 17
DatPySci/Self-Distillation-Qwen3-4B-Thinking-2507-medreason-prefix-instruct-tokenized Viewer • Updated Jul 31 • 32.7k • 24
DatPySci/Self-Distillation-Qwen3-4B-Thinking-2507-medreason-prefix-instruct-tokenized Viewer • Updated Jul 31 • 32.7k • 24
DatPySci/Self-Distillation-Qwen3-8B-medreason-prefix-instruct-tokenized Viewer • Updated Jul 31 • 32.7k • 15
DatPySci/Self-Distillation-Qwen3-8B-medreason-prefix-instruct-tokenized Viewer • Updated Jul 31 • 32.7k • 15
DatPySci/Self-Distillation-Qwen3-4B-Thinking-2507-medreason-tokenized Viewer • Updated Jul 21 • 32.7k • 30
DatPySci/Self-Distillation-Qwen3-4B-Thinking-2507-medreason-tokenized Viewer • Updated Jul 21 • 32.7k • 30