Instructions to use Vladimirlv/ru-promptriever-qwen3-4b-ru-only with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Vladimirlv/ru-promptriever-qwen3-4b-ru-only with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="Vladimirlv/ru-promptriever-qwen3-4b-ru-only")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Vladimirlv/ru-promptriever-qwen3-4b-ru-only", device_map="auto") - Notebooks
- Google Colab
- Kaggle
ru-Promptriever-Qwen3-4B-ru-only
Overview
Standard dense retrieval models score query–passage pairs using a single semantic similarity signal, giving users no control over what "relevant" means beyond keyword choice. Promptriever (Weller et al., 2024) introduced per-instance natural language instructions that dynamically redefine relevance — a capability previously limited to generative LLMs.
ru-Promptriever extends this paradigm to Russian:
- Architecture: Qwen3-based causal LM fine-tuned as a bi-encoder with LoRA + GradCache
- Pooling: last-token (EOS) pooling, same as the original Promptriever
- Key training signal: instruction negatives — passages that are topically relevant to the query but violate the instruction constraint
Model Family
| Model | Parameters | Description | Link |
|---|---|---|---|
| ru-Promptriever-4B | 4B | Final model — best results | link |
| ru-Promptriever-4B-pretrained | 4B | Base pretrained on synthetic data only | link |
| ru-Promptriever-4B-ru-only | 4B | Continued training on Russian-only data | this model |
| ru-Promptriever-1.7B | 1.7B | Scaling experiment | link |
| ru-Promptriever-0.6B | 0.6B | Scaling experiment | link |
This Model
This is the Russian-only continued training variant. Starting from ru-Promptriever-4B-pretrained (trained on synthetic data only), it was further fine-tuned on a mix of:
- Russian real data — MIRACL + MrTyDi (~11k pairs with human-annotated relevance)
- Russian synthetic data — instruction-augmented (
11k) and standard pairs (6k)
Adding real retrieval data (not just synthetic) improved both nDCG and p-MRR compared to the pretrained model. However, the final model with additional English instruction-following data achieves even higher p-MRR.
Note: For best performance, use the final model instead.
Evaluation Results
mFollowIR-RU
Russian split of mFollowIR — multilingual instruction-following retrieval using TREC NeuCLIR narratives as instructions.
p-MRR (Pairwise Mean Reciprocal Rank, ×100) is the primary instruction-following metric — higher means the model correctly adjusts rankings when instructions change. nDCG@20 measures standard retrieval quality.
| Model | nDCG@20 | p-MRR |
|---|---|---|
| BM25 | 0.452 | +0.67 |
| mE5-large | 0.428 | −2.03 |
| BGE-M3 | 0.479 | −4.15 |
| Promptriever-8B | 0.532 | +12.21 |
| Qwen3-Embedding-4B | 0.549 | +8.10 |
| ru-Promptriever-0.6B | 0.231 | −4.35 |
| ru-Promptriever-1.7B | 0.444 | +14.65 |
| ru-Promptriever-4B-pretrained | 0.461 | +15.26 |
| ru-Promptriever-4B-ru-only (this model) | 0.512 | +16.80 |
| ru-Promptriever-4B | 0.512 | +18.57 |
Usage
Basic Retrieval (no instruction)
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch.nn.functional as F
model_name = "Vladimirlv/ru-promptriever-qwen3-4b-ru-only"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model.eval()
def encode(texts: list[str], max_length: int = 512) -> torch.Tensor:
"""Encode texts using last-token (EOS) pooling."""
inputs = tokenizer(
texts,
padding=True,
truncation=True,
max_length=max_length,
return_tensors="pt",
).to(model.device)
with torch.no_grad():
# Bypass lm_head to get post-norm hidden states
original_lm_head = model.lm_head
model.lm_head = torch.nn.Identity()
outputs = model(**inputs, use_cache=False, return_dict=True)
model.lm_head = original_lm_head
# EOS pooling: take embedding at last non-padding token
seq_len = inputs["attention_mask"].sum(dim=1) - 1
embeddings = outputs.logits[torch.arange(len(texts)), seq_len]
return F.normalize(embeddings, p=2, dim=1)
query = "Когда была основана Москва?"
passages = [
"Москва была основана в 1147 году князем Юрием Долгоруким.",
"Санкт-Петербург был основан Петром I в 1703 году.",
]
q_emb = encode([query])
p_emb = encode(passages)
scores = (q_emb @ p_emb.T).squeeze()
print(scores) # tensor([0.82, 0.61])
Instruction-Following Retrieval
# Append the instruction directly to the query (same format as training)
instruction = "Найди документ, в котором упоминается конкретная дата основания города."
instructed_query = f"{query} {instruction}"
q_emb = encode([instructed_query])
p_emb = encode(passages)
scores = (q_emb @ p_emb.T).squeeze()
# The model adjusts rankings based on the instruction
Using with sentence-transformers
This model is not compatible with sentence-transformers out of the box due to the custom EOS pooling. Use the snippet above directly with transformers.
Model Details
| Property | Value |
|---|---|
| Base model | ru-Promptriever-4B-pretrained (continued training) |
| Architecture | CausalLM bi-encoder (EOS pooling) |
| Fine-tuning method | LoRA (rank-32, α=64, all linear layers) |
| Training data | ~28k rows (11k Russian real + 11k synthetic instructed + 6k synthetic standard) |
| Effective batch size | 128 (8 per device × 4 accum × 4 GPUs) |
| Loss | InfoNCE contrastive (temperature=0.01) |
| Learning rate | 5e-5 |
| Epochs | 2 |
| Max query length | 512 tokens |
| Max passage length | 256 tokens |
Training Data
This model was trained on a mix of:
- Russian real data (~11k) — from MIRACL and MrTyDi retrieval datasets with human-annotated relevance judgments
- Russian synthetic data (~17k) — instruction-augmented and standard pairs from ru-promptriever-dataset
Key properties:
- Instruction negatives: passages that are topically relevant but violate the instruction (3 per instructed query)
- Paired rows: each source query has both a standard row and an instructed row to prevent catastrophic forgetting
Intended Use
- Instruction-following dense retrieval in Russian: RAG pipelines, search systems, and scenarios requiring fine-grained query control via natural language
- Research on multilingual instruction-following retrieval and bi-encoder training
- Benchmarking alongside mE5, BGE-M3, and Promptriever-style models
Out-of-Scope
- General-purpose text embedding (use mE5-large or BGE-M3 if no instruction-following is needed)
- Commercial applications (see License below)
Limitations
- MS MARCO origin: The training corpus derives from English web passages machine-translated to Russian. A portion of passages retain translation artifacts despite LLM-based rewriting.
- Standard retrieval trade-off: Instruction-following training slightly reduces standard retrieval quality compared to encoder-only models (mE5-large, BGE-M3).
- Noisy synthetic data: Instructions and negatives were generated and validated automatically by an LLM; a small fraction of imperfect examples may remain.
- Russian only: This model variant was trained and evaluated exclusively on Russian data.
License
This model is released under CC BY-NC 4.0 (Creative Commons Attribution–NonCommercial 4.0 International).
The non-commercial restriction is inherited from the upstream MS MARCO license (Microsoft Research License — non-commercial use only), which governs the training corpus.
Citation
If you use this model, please cite the original Promptriever paper:
@article{weller2024promptriever,
title = {Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models},
author = {Weller, Orion and Van Durme, Benjamin and Lawrie, Dawn and
Paranjape, Ashwin and Zhang, Yuhao and Hessel, Jack},
journal = {arXiv preprint arXiv:2409.11136},
year = {2024}
}
Model tree for Vladimirlv/ru-promptriever-qwen3-4b-ru-only
Base model
Qwen/Qwen3-4B-Base