raw-250m

UNTRAINED MODEL — RAW INITIALIZED WEIGHTS

This repository contains a raw foundation-model starting point.

The weights are randomly initialized and have not been pretrained. The model will not generate meaningful text until a pretraining stage is completed.

This release is intended for:

  • architecture experiments
  • pretraining research
  • transformer testing
  • educational purposes

Architecture

Field Value
Parameters 248.91M
Layers 10
Hidden size 1280
Attention heads 20
KV heads 5
Head dimension 64
FFN 3776 SwiGLU
Context length 4096 tokens
Vocabulary 49157
Tied embeddings True
RoPE theta 500000
Precision bfloat16

Initialization

  • Method: gpt2_scaled
  • Seed: 42
  • Deterministic initialization

Memory

  • Weights: 497.81 MB
  • KV cache: 12.80 KB/token
  • AdamW training state estimate: 3.98 GB

Usage

This model is not pretrained.

Example loading code:

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "Quantum47Vision/raw-250m"
)

tokenizer = AutoTokenizer.from_pretrained(
    "Quantum47Vision/raw-250m"
)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support