Configuration Parsing Warning:In config.json: "quantization_config.modules_to_not_convert" must be an array
DeepSeek-R1-Distill-Qwen-7B-AWQ book-RAG submission
This is a training-free IOL-AI submission using a 4-bit AWQ conversion of
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B. It retrieves only from the public,
string-only extraction of Vlad A. Neacșu's Linguistics Olympiad: Training
guide. Private curated training data are not present or required at runtime.
The checkpoint uses Qwen2ForCausalLM, so it is compatible with the challenge's
Transformers 4.44.1 and AutoAWQ 0.2.7 runtime. The included weights are from
casperhansen/deepseek-r1-distill-qwen-7b-awq, revision
305e6f12907dc78ae61a1f0bb7a19faa2b25e8a3, which is an AWQ conversion of the
official DeepSeek model.
Challenge execution
python script.py
The script reads /tmp/data/test.csv and writes submission.csv. Its columns
are id,pred,explanation; pred is a JSON-encoded list of answer strings.
Explanation generation is disabled by default, leaving that optional column
blank. It can be restored with:
python script.py --explanations on
For local data:
python script.py --input path/to/test.csv --output submission.csv
python script.py --self-test
After the model loads successfully, the output file is initialized and then atomically rewritten after every completed row. If evaluation reaches its time limit, predictions already completed remain in a valid submission file.
DeepSeek reasoning and generation
DeepSeek's solver instructions, retrieved context, and current problem are
placed in one user message rather than a separate system message. The assistant
prompt is prefixed with <think> to engage the distilled reasoning behavior.
Reasoning stops when </think> appears or at its configured cap. The runtime
then starts a separate FINAL ANSWERS: stage, guaranteeing that reasoning
cannot consume the answer budget. Only that answer block is serialized into
pred.
Defaults follow DeepSeek's recommended sampling values while retaining a bounded challenge-time output:
temperature=0.6top_p=0.95- sampling enabled
max_reasoning_tokens=4096max_answer_tokens=1024- one answer-only retry with
answer_retry_tokens=512
The model context is capped at 32,768 tokens. The script loads AWQ directly in
FP16, enables the KV cache, and applies an inference-only last-token lm_head
hook to avoid materializing full-prompt FP32 vocabulary logits on the T4.
Book-only retrieval
The default prompt includes one general book method and two worked book
examples. Retrieval combines dependency-free BM25, character 3–5-gram TF-IDF,
task-family inference, and book metadata boosts. It uses only the current
problem's context + query; answers and private curated datasets are never
indexed.
Settings are in rag_resources/config.json. Environment overrides are available
for IOL_TOP_METHODS, IOL_TOP_EXAMPLES, IOL_CHAR_TFIDF_WEIGHT,
IOL_RAG_MAX_CHARS, IOL_MAX_REASONING_TOKENS,
IOL_MAX_ANSWER_TOKENS, IOL_ANSWER_RETRY_TOKENS,
IOL_EXPLANATION_MAX_NEW_TOKENS,
IOL_ENABLE_EXPLANATIONS, and IOL_SEED.
Sources and licenses
The model and conversion are marked MIT. The book-derived resources retain
their own license and attribution in rag_resources/ATTRIBUTION.md.
- Downloads last month
- 75
Model tree for pei39/deepseek_r1_distill_qwen_7b_awq_rag
Base model
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B