rudrakshrakeshzodage's picture
Upload folder using huggingface_hub
f9a1172 verified
|
Raw
History Blame Contribute Delete
3.26 kB
metadata
license: apache-2.0
base_model: distil-whisper/distil-large-v3
base_model_relation: quantized
library_name: transformers
pipeline_tag: automatic-speech-recognition
tags:
  - automatic-speech-recognition
  - whisper
  - distil-whisper
  - quantization
  - int8
  - 4-bit
  - nf4
  - bitsandbytes
  - ctranslate2
  - faster-whisper

Distil-Whisper large-v3 β€” PyTorch Q4 (NF4) Quantized (GPU - bitsandbytes)

This model is a quantized/optimized version of the baseline Distil-Whisper (distil-large-v3) model. It is part of a benchmarked suite of quantized models evaluated on local hardware.

This model is part of a suite of optimized/quantized versions of the base model. Other variants in this direction:

Base model: distil-whisper/distil-large-v3


πŸ“Š Quantization & Performance Benchmark Results

Below is the comparative performance table generated empirically using 8 CPU threads / cores:

Backend Precision Device Model Size (GB) Mean Latency (s) Throughput (Words/s) RTF WER (%) Peak RAM (GB) Peak VRAM (GB)
PYTORCH FP32 CUDA 3.024 0.427 44.81 0.064 5.29% 3.62 2.98
PYTORCH FP16 CUDA 1.512 0.189 100.25 0.028 5.29% 2.12 1.46
PYTORCH INT8 CUDA 0.756 0.293 64.00 0.042 5.29% 1.92 0.85
PYTORCH Q4 CUDA 0.378 0.377 50.63 0.056 5.29% 1.93 0.60
GGML FP16 CPU 1.409 17.269 1.14 2.633 5.29% 4.14 0.00
GGML INT8 CPU 1.409 9.271 2.18 1.441 5.29% 1.50 0.00
GGML Q5 CPU 1.409 16.735 1.18 2.566 5.29% 2.28 0.00

πŸ’Ύ Model Size Reduction Comparison

Model Size

πŸ“ˆ Transcription Speed & Latency Comparison

Latency

πŸš€ Words Transcribed per Second (Throughput)

Throughput


πŸš€ Usage

import torch
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, BitsAndBytesConfig

model_id = "rudrakshrakeshzodage/distil-whisper-large-v3-pytorch-q4"

quant_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_use_double_quant=True,
    bnb_4bit_compute_dtype=torch.float16
)

model = AutoModelForSpeechSeq2Seq.from_pretrained(
    model_id,
    quantization_config=quant_config,
    device_map="cuda"
)
processor = AutoProcessor.from_pretrained(model_id)

πŸ“œ Citation & Credits

Quantization research and benchmarking by RudrakshRakeshZodage.