Text Generation
Transformers
Safetensors
English
qwen3
conversational
text-generation-inference
dsouzaJithesh commited on
Commit
d8521c7
·
verified ·
1 Parent(s): 9282e10

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +0 -7
README.md CHANGED
@@ -76,13 +76,6 @@ Below is a comparison between the **Base Qwen3-0.6B** model and **This DPO-Train
76
  | TruthfulQA MC2 | acc | 0.4275 | 0.4335 | +0.0060 |
77
  | Winogrande | acc | 0.5604 | 0.5627 | +0.0023 |
78
 
79
- ### 📝 Summary
80
-
81
- - The DPO model shows **small but consistent improvements** across reasoning benchmarks.
82
- - **TruthfulQA improves**, indicating better factuality and reduced hallucination.
83
- - No regressions observed — core reasoning abilities remain stable.
84
- - These results match expectations for preference-based DPO training using HelpSteer2.
85
-
86
 
87
  ## Model Card Authors
88
 
 
76
  | TruthfulQA MC2 | acc | 0.4275 | 0.4335 | +0.0060 |
77
  | Winogrande | acc | 0.5604 | 0.5627 | +0.0023 |
78
 
 
 
 
 
 
 
 
79
 
80
  ## Model Card Authors
81