PolyAlign Chinese HDPO MLP Critic: llama32_3b

This repository contains the frozen-encoder, bucket-conditioned MLP critic bundle used to score Chinese HDPO preference pairs for llama32_3b.

The policy-training HDPO JSONs were produced locally under:

data/chinese/hdpo_prepared/llama32_3b/llamafactory/

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support