Reward Hacking MO Checkpoints and Rollouts Artefacts from "(Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL". KL 0.0 hacks with faithful CoT, 0.2 with unfaithful. ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2 Text Generation • Updated Jul 1 ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2 Text Generation • Updated Jul 1 ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Viewer • Updated Jul 1 • 25.7k • 360 ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Viewer • Updated Jul 1 • 25.8k • 359
ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Viewer • Updated Jul 1 • 25.7k • 360
ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Viewer • Updated Jul 1 • 25.8k • 359
Lie Detection Did you lie? Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms ai-safety-institute/lie-detection-rollouts Viewer • Updated Jun 25 • 2.48M • 4.31k Lie Detection Model Organisms Collection Model organisms trained to reason about lying in CoT, then lie in text output. • 17 items • Updated Jun 10 Apollo-Style Deception Probes Collection Lie detection probes trained following the approach of Detecting Strategic Deception Using Linear Probes. • 53 items • Updated Jun 25 Targeted Apollo Deception Probes Collection Lie detection probes trained following the approach of 'Building Better Deception Probes Using Targeted Instruction Pairs' • 53 items • Updated Jun 25
Lie Detection Model Organisms Collection Model organisms trained to reason about lying in CoT, then lie in text output. • 17 items • Updated Jun 10
Apollo-Style Deception Probes Collection Lie detection probes trained following the approach of Detecting Strategic Deception Using Linear Probes. • 53 items • Updated Jun 25
Targeted Apollo Deception Probes Collection Lie detection probes trained following the approach of 'Building Better Deception Probes Using Targeted Instruction Pairs' • 53 items • Updated Jun 25
Reward Hacking MO Checkpoints and Rollouts Artefacts from "(Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL". KL 0.0 hacks with faithful CoT, 0.2 with unfaithful. ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2 Text Generation • Updated Jul 1 ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2 Text Generation • Updated Jul 1 ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Viewer • Updated Jul 1 • 25.7k • 360 ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Viewer • Updated Jul 1 • 25.8k • 359
ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Viewer • Updated Jul 1 • 25.7k • 360
ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Viewer • Updated Jul 1 • 25.8k • 359
Lie Detection Did you lie? Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms ai-safety-institute/lie-detection-rollouts Viewer • Updated Jun 25 • 2.48M • 4.31k Lie Detection Model Organisms Collection Model organisms trained to reason about lying in CoT, then lie in text output. • 17 items • Updated Jun 10 Apollo-Style Deception Probes Collection Lie detection probes trained following the approach of Detecting Strategic Deception Using Linear Probes. • 53 items • Updated Jun 25 Targeted Apollo Deception Probes Collection Lie detection probes trained following the approach of 'Building Better Deception Probes Using Targeted Instruction Pairs' • 53 items • Updated Jun 25
Lie Detection Model Organisms Collection Model organisms trained to reason about lying in CoT, then lie in text output. • 17 items • Updated Jun 10
Apollo-Style Deception Probes Collection Lie detection probes trained following the approach of Detecting Strategic Deception Using Linear Probes. • 53 items • Updated Jun 25
Targeted Apollo Deception Probes Collection Lie detection probes trained following the approach of 'Building Better Deception Probes Using Targeted Instruction Pairs' • 53 items • Updated Jun 25
ai-safety-institute/dyl-honest-mlabonne-gemma-3-27b-it-abliterated__aletheias-quest-collusion-model-organis-205e9d95 Updated Jun 25
ai-safety-institute/dyl-honest-google-gemma-3-27b-it__aletheias-quest-hidden-goal-model-organism-gemma3-27b-v1 Updated Jun 25
ai-safety-institute/dyl-honest-qwen-qwen3.5-27b__aletheias-quest-botc-latest-checkpoint Updated Jun 25
ai-safety-institute/dyl-truthful-mlabonne-gemma-3-27b-it-abliterated__aletheias-quest-collusion-model-organ-9ead2124 Updated Jun 25
ai-safety-institute/dyl-mlabonne-gemma-3-27b-it-abliterated__aletheias-quest-collusion-model-organism-gemma3-27b-v1 Updated Jun 25
ai-safety-institute/uq-mlabonne-gemma-3-27b-it-abliterated__aletheias-quest-collusion-model-organism-gemma3-27b-v1 Updated Jun 25
ai-safety-institute/uq-google-gemma-3-27b-it__aletheias-quest-hidden-goal-model-organism-gemma3-27b-v1 Updated Jun 25
ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Viewer • Updated Jul 1 • 25.8k • 359
ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Viewer • Updated Jul 1 • 25.7k • 360
ai-safety-institute/glm_5_2_fp8_ab_hallucinates_citations_rollouts Viewer • Updated Jun 29 • 6.1k • 70
ai-safety-institute/glm_5_2_fp8_ab_contextual_optimism_rollouts Viewer • Updated Jun 29 • 6.11k • 136