How to use from the
Use from the
ColPali library
# No code snippets available yet for this library.

# To use this model, check the repository files and the library's documentation.

# Want to help? PRs adding snippets are welcome at:
# https://github.com/huggingface/huggingface.js



Jina AI: Your Search Foundation, Supercharged!

ONNX conversion of Jina AI's jina-embeddings-v4 (text-matching task).

Jina Embeddings v4 β€” text-matching β†’ ONNX (sub-part decomposition)

Original Model | Blog | Technical Report | API

Model overview

Source checkpoint is jinaai/jina-embeddings-v4-vllm-text-matching β€” a stock Qwen2.5-VL-3B with the text-matching task LoRA merged into the base weights (no custom adapter code, single full checkpoint). It is one of three per-task variants of jina-embeddings-v4; this repo hosts the ONNX conversion of the text-matching one.

text-matching is the symmetric similarity task: both sides of a pair use the same Query: prefix (unlike retrieval/code, which are asymmetric Query: / Passage:). Use it for sentence-similarity, STS, deduplication, and clustering.

This is a single-vector model: the embedding is a masked mean-pool over the last hidden state, L2-normalized, 2048-d, with Matryoshka truncation to 128/256/512/1024/2048.

Decomposition

The model is split into three ONNX sub-parts (same pattern as the other recipes in this repo), so the heavy backbone is stored once and reused by both the text and image paths:

Sub-part Input β†’ Output Notes
vision.onnx pixel_values [N,1176] β†’ image_features [N,2048] task-agnostic (vision tower has no LoRA); grid baked at build resolution
embeddings.onnx input_ids [B,S] (+ image_features [N,2048]) β†’ inputs_embeds [B,S,2048] token embeds; image features scattered into <image_pad> positions
backbone.onnx inputs_embeds, attention_mask, position_ids [3,B,S] β†’ last_hidden [B,S,2048] text-matching LoRA is merged here; MROPE position_ids host-computed

Compose at inference (all ONNX; the driver only wires sessions and pools):

text  : embeddings(ids)                     β†’ backbone β†’ mean-pool(attn_mask)   β†’ L2norm
image : vision(px) β†’ embeddings(ids, feats) β†’ backbone β†’ mean-pool(vision-span) β†’ L2norm

Pooling and Matryoshka truncation happen in the driver (nothing baked), so one build serves every output dimension.

Why host-computed position_ids?

The model uses MROPE (mrope_section [16,24,24]), which onnxruntime-genai's ModelBuilder cannot emit. position_ids [3,B,S] are therefore computed on the host and fed in: cumulative positions for text, and get_rope_index(...) over the image grid for image inputs. The graph stays clean.

Prompts

Both sides of a comparison use the same Query: prefix (symmetric). The image prompt is the fixed template used across all tasks. manifest.json records this per build under prompts ({"query": "Query:", "document": "Query:", "symmetric": true}).

text : "Query: <your text>"
image: "<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|>Describe the image.<|im_end|>\n"

Files

Each build directory is self-contained (sub-parts + image_meta.npz + tokenizer/processor assets + manifest.json, which records the source hf_model id, precision, and any quantized sub-parts):

Dir Precision total
fp16 fp16 7.0 GB
fp32 fp32 14 GB
int8 fp16, backbone int8 ~4.6 GB
int4 fp16, backbone int4 3.3 GB

int8 (fp16 graph + int8 backbone) is the recommended quantized option β€” smallest build that still clears the 0.999 fidelity bar.

cuda_* directories, if present and empty, are placeholders. This environment's PyTorch/ORT are CPU builds, so GPU builds produce nothing there. The exported ONNX is execution-provider agnostic β€” the same files run on CUDAExecutionProvider via onnxruntime-gpu with no rebuild and no device flag.

Fidelity vs full PyTorch

Composed ONNX chain vs the full Qwen2_5_VLForConditionalGeneration (pooled-embedding cosine, worst of 3 text samples + 1 image):

Build worst cosine verdict
fp32 0.999984 βœ…
fp16 0.999987 βœ…
int8 (backbone) 0.999436 βœ… (β‰₯0.999)
int4 (backbone) 0.912048 ❌ not for production

(Reference model loaded in fp16 for the eval; eval.py dedupes repeated paths, so listing a dir twice runs it once.)

int8 is the quantization sweet spot β€” ~35 % smaller than fp16 with negligible cosine drift. int4 is too coarse for an embedding model (the pooled/normalized vector amplifies 4-bit weight error into ~8 % drift, which wrecks similarity ranking) β€” build.py emits it with a warning, not a failure.

Reproducing / using

CPU only; runs in the repo's uv project env (transformers 5.x, torchvision for the image processor). The pipeline is three task-agnostic scripts sharing common.py β€” point --model at the text-matching source:

# build sub-parts β€” one --precision flag: fp16 (default) | fp32 | int8 | int4
uv run build.py --model vllm-text-matching --output onnx/fp16                  # fp16
uv run build.py --model vllm-text-matching --output onnx/fp32 --precision fp32
uv run build.py --model vllm-text-matching --output onnx/int8 --precision int8 # fp16 graph + int8 backbone
uv run build.py --model vllm-text-matching --output onnx/int4 --precision int4 # lossy (see above)

# eval β€” accepts multiple build dirs (positional), dedupes repeats, auto-detects each one's precision
uv run eval.py --model vllm-text-matching onnx/fp16 onnx/fp32 onnx/int8 onnx/int4

# inference (no PyTorch load) β€” text-matching uses the Query: prefix for BOTH texts
uv run inference.py --onnx-dir onnx/fp16 --text "The impacts of climate change on coastal cities"
uv run inference.py --onnx-dir onnx/fp16 --text "..." --truncate-dim 256
uv run inference.py --onnx-dir onnx/fp16 --image doc.png

int8/int4 build the fp16 graph then weight-quantize the backbone in place (block-wise MatMulNBits; vision/embeddings stay fp16). build.py runs a composed self-sanity check: fp16 / fp32 / int8 must hit cosine β‰₯ 0.999 or the build fails, while int4 only warns. The vision sub-part is identical across all tasks (no LoRA), so a single vision.onnx can be shared to save disk.

Same scripts serve the other tasks β€” --model vllm-retrieval or --model vllm-text-code β€” the only difference being the prompt convention (those are asymmetric Query: / Passage:).

Minimal ONNX Runtime example (text)

import json, numpy as np, onnxruntime as ort
from pathlib import Path
from transformers import AutoTokenizer

d = Path("onnx/fp16")
man = json.loads((d / "manifest.json").read_text())
npdt = np.float16 if man["precision"] == "fp16" else np.float32
tok = AutoTokenizer.from_pretrained(str(d))

def sess(name):  # log level raised to silence the harmless constant-fold notice
    so = ort.SessionOptions(); so.log_severity_level = 3
    return ort.InferenceSession(str(d / name), so, providers=["CPUExecutionProvider"])

emb_s, back_s = sess("embeddings.onnx"), sess("backbone.onnx")

# text-matching: both texts use the SAME "Query:" prefix
enc = tok(["Query: The impacts of climate change on coastal cities"], return_tensors="np", padding="longest")
ids, am = enc["input_ids"], enc["attention_mask"]
pos = np.clip(np.cumsum(am, -1) - 1, 0, None)[None].repeat(3, 0)         # MROPE (text)
e = emb_s.run(None, {"input_ids": ids, "image_features": np.zeros((0, 2048), npdt)})[0]
h = back_s.run(None, {"inputs_embeds": e, "attention_mask": am, "position_ids": pos})[0]

pooled = (h * am[..., None]).sum(1) / am.sum(1, keepdims=True)           # masked mean-pool
emb = pooled / np.linalg.norm(pooled, axis=-1, keepdims=True)           # L2-norm β†’ [1, 2048]
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for onnx-community/jina-embeddings-v4-vllm-text-matching

Quantized
(2)
this model

Paper for onnx-community/jina-embeddings-v4-vllm-text-matching