Token Classification
Transformers
Safetensors
English
bert
ner

Applied NER Stage 3 โ€” BERT Tiny

An eight-label English token classifier fine-tuned from google/bert_uncased_L-2_H-128_A-2. Repository: THemidli/applied-ner-stage3-bert-tiny.

Results

Exact entity-level seqeval metrics:

Split Precision Recall F1 Token accuracy
Train 0.9581 0.9726 0.9653 0.9957
Test 0.4215 0.5273 0.4685 0.8269
Label Precision Recall F1 Support
PERSON 0.461 0.641 0.536 195
ORGANIZATION 0.200 0.218 0.208 147
LOCATION 0.432 0.552 0.485 143
TIMEDATE 0.792 0.844 0.817 167
PRODUCT 0.171 0.189 0.180 127
WORKOFART 0.128 0.247 0.168 97
JOB 0.699 0.798 0.745 99
AMOUNT 0.556 0.625 0.588 104

On 40 fresh, manually gold-labeled wild probes, exact span F1 was 0.4785 (precision 0.4505, recall 0.5102).

Training

  • Dataset: THemidli/applied-ner-stage2-expanded
  • Seed: 20260802
  • Hardware: Apple MPS (macOS-27.0-arm64-arm-64bit)
  • Runtime: 17.780 seconds
  • Records/chunks: 641/665 train; 159/165 test
  • Maximum length: 256; fast-tokenizer overflow chunks, no overlapping stride
  • Hyperparameters: {"epochs": 20, "eval_batch_size": 64, "learning_rate": 0.0005, "scheduler": "linear", "train_batch_size": 32, "warmup_steps": 42, "weight_decay": 0.01}
  • No validation split and no test-driven checkpoint selection

Footprint and CPU benchmark

  • Parameters: 4,371,601 (17.49 MB tensor storage)
  • Saved artifact: 18.21 MB
  • Model-load RSS delta: 31.85 MB
  • End-to-end inference RSS delta: 41.30 MB
  • CPU throughput: 11459.9 examples/s at batch 32 with 8 threads
  • Mean latency: 0.0873 ms/example at that batch size

The benchmark covers tokenizer plus PyTorch CPU forward pass over 40 short probes, repeated 50 times. It is workload- and hardware-specific, not single-request latency.

Labels

PERSON, ORGANIZATION, LOCATION, TIMEDATE, PRODUCT, WORKOFART, JOB, AMOUNT using BIO encoding.

Limitations

This is a 4.37M-parameter uncased two-layer BERT trained on a small, heterogeneous dataset. It is a compact baseline, not a production privacy system. Rare works/products, company-versus-product context, exact boundaries, and subword-heavy names remain weak. The 40-probe wild set is diagnostic, not a population benchmark.

Downloads last month
-
Safetensors
Model size
4.37M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for THemidli/applied-ner-stage3-bert-tiny

Finetuned
(126)
this model

Dataset used to train THemidli/applied-ner-stage3-bert-tiny