whisper-large-v3-turbo-onnx-int8

This is a INT8 Quantized ONNX version of openai/whisper-large-v3-turbo.

Model Details

Size Comparison

Version Size
Base ONNX (FP32) 4198.51 MB
FP16 ONNX 2099.54 MB
INT8 Quantized ONNX 1459.58 MB
Compression 2.88x

Usage

from optimum.onnxruntime import ORTModelForSpeechSeq2Seq
from transformers import AutoProcessor

model = ORTModelForSpeechSeq2Seq.from_pretrained("kostasang/whisper-large-v3-turbo-onnx-int8")
processor = AutoProcessor.from_pretrained("kostasang/whisper-large-v3-turbo-onnx-int8")

Quantization Details

This model was quantized using dynamic INT8 quantization with ONNX Runtime.

  • Quantized Operators: MatMul, Attention, Gemm
  • Target Architecture: arm64
Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kostasang/whisper-large-v3-turbo-onnx-int8

Quantized
(223)
this model