UnfaithRL/OLMo-2-0425-1B-hint_verbalization_reward_relaxed_v7-512 Reinforcement Learning • 1B • Updated Jul 8 • 35
UnfaithRL/OLMo-2-0425-1B-hint_verbalization_reward_relaxed_v7-1024 Reinforcement Learning • 1B • Updated Jul 8 • 25
UnfaithRL/OLMo-2-0425-1B-hint_verbalization_reward_relaxed_v7-2048 Reinforcement Learning • 1B • Updated Jul 8 • 29
UnfaithRL/Qwen2.5-0.5B-hint_verbalization_reward_relaxed_v7-512 Reinforcement Learning • 0.5B • Updated Jul 8 • 41
UnfaithRL/Qwen2.5-0.5B-hint_verbalization_reward_relaxed_v7-1024 Reinforcement Learning • 0.5B • Updated Jul 8 • 24
UnfaithRL/Qwen2.5-0.5B-hint_verbalization_reward_relaxed_v7-2048 Reinforcement Learning • 0.5B • Updated Jul 8 • 28
UnfaithRL/OLMo-2-0425-1B-hint_verbalization_reward_strict_v9-512 Reinforcement Learning • 1B • Updated Jul 8 • 9
UnfaithRL/OLMo-2-0425-1B-hint_verbalization_reward_strict_v9-1024 Reinforcement Learning • 1B • Updated Jul 8 • 9
UnfaithRL/OLMo-2-0425-1B-hint_verbalization_reward_strict_v9-2048 Reinforcement Learning • 1B • Updated Jul 8 • 8
UnfaithRL/Qwen2.5-0.5B-hint_verbalization_reward_strict_v9-512 Reinforcement Learning • 0.5B • Updated Jul 8 • 8
UnfaithRL/Qwen2.5-0.5B-hint_verbalization_reward_strict_v9-1024 Reinforcement Learning • 0.5B • Updated Jul 8 • 5
UnfaithRL/Qwen2.5-0.5B-hint_verbalization_reward_strict_v9-2048 Reinforcement Learning • 0.5B • Updated Jul 8 • 6
UnfaithRL/OLMo-2-0425-1B-Instruct-hint_verbalization_reward_strict_v2-512 Reinforcement Learning • 1B • Updated Jul 8 • 33
UnfaithRL/OLMo-2-0425-1B-Instruct-hint_verbalization_reward_strict_v2-1024 Reinforcement Learning • 1B • Updated Jul 8 • 28
UnfaithRL/OLMo-2-0425-1B-Instruct-hint_verbalization_reward_strict_v2-2048 Reinforcement Learning • 1B • Updated Jul 8 • 31
UnfaithRL/Qwen2.5-0.5B-Instruct-hint_verbalization_reward_strict_v2-512 Reinforcement Learning • 0.5B • Updated Jul 8 • 42
UnfaithRL/Qwen2.5-0.5B-Instruct-hint_verbalization_reward_strict_v2-1024 Reinforcement Learning • 0.5B • Updated Jul 8 • 26
UnfaithRL/Qwen2.5-0.5B-Instruct-hint_verbalization_reward_strict_v2-2048 Reinforcement Learning • 0.5B • Updated Jul 8 • 28
UnfaithRL/OLMo-2-0425-1B-hint_following_reward_faithful_prompt-512 Reinforcement Learning • 1B • Updated Jul 8 • 8
UnfaithRL/OLMo-2-0425-1B-hint_following_reward_faithful_prompt-1024 Reinforcement Learning • 1B • Updated Jul 8 • 6
UnfaithRL/OLMo-2-0425-1B-hint_following_reward_faithful_prompt-2048 Reinforcement Learning • 1B • Updated Jul 8 • 7
UnfaithRL/Qwen2.5-0.5B-hint_following_reward_faithful_prompt-512 Reinforcement Learning • 0.5B • Updated Jul 8 • 15
UnfaithRL/Qwen2.5-0.5B-hint_following_reward_faithful_prompt-1024 Reinforcement Learning • 0.5B • Updated Jul 8 • 5
UnfaithRL/Qwen2.5-0.5B-hint_following_reward_faithful_prompt-2048 Reinforcement Learning • 0.5B • Updated Jul 8 • 7