finetuned_byT5large_pnb_joint

This model is a fine-tuned version of google/byt5-large on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 2.3437
  • Chrf++: 4.1915
  • Gen Len: 189.0588

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 5e-05
  • train_batch_size: 4
  • eval_batch_size: 1
  • seed: 42
  • gradient_accumulation_steps: 8
  • total_train_batch_size: 32
  • optimizer: Use OptimizerNames.ADAMW_TORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • num_epochs: 100

Training results

Training Loss Epoch Step Validation Loss Chrf++ Gen Len
41.4608 1.4706 4 2.4036 5.2353 188.3529
41.4608 2.9412 8 2.4034 5.2089 188.9412
41.4608 4.0 12 2.4040 5.6686 188.8824
41.4608 5.4706 16 2.4033 4.9653 188.9412
41.4608 6.9412 20 2.4033 4.6819 188.8235
41.4608 8.0 24 2.4026 4.4912 188.7647
41.4608 9.4706 28 2.4011 4.4052 188.7059
41.4608 10.9412 32 2.4008 4.5225 189.1765
41.4608 12.0 36 2.3990 5.1372 189.0588
41.4608 13.4706 40 2.3954 4.3585 189.0
41.4608 14.9412 44 2.3907 4.7137 188.2941
41.4608 16.0 48 2.3887 4.8434 188.9412
41.4608 17.4706 52 2.3836 4.7941 188.7647
41.4608 18.9412 56 2.3819 4.7571 189.0
41.4608 20.0 60 2.3742 5.3375 188.8235
41.4608 21.4706 64 2.3583 3.9694 186.2941
41.4608 22.9412 68 2.3528 3.7939 188.8235
41.4608 24.0 72 2.3437 4.7144 188.0
41.4608 25.4706 76 2.3437 4.1915 189.0588

Framework versions

  • Transformers 5.12.1
  • Pytorch 2.6.0+cu124
  • Datasets 5.0.0
  • Tokenizers 0.22.2
Downloads last month
17
Safetensors
Model size
1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Tsedeniya/finetuned_byT5large_pnb_joint

Finetuned
(30)
this model