CSD (ViT-L/14) β€” HuggingFace format

Unofficial transformers-format port of the CSD style model from "Measuring Style Similarity in Diffusion Models" (Somepalli, Gupta, Gupta, Shrivastava, Goldstein, Feizi; 2024). Loads via trust_remote_code with no clip / open_clip runtime dependency β€” the OpenAI CLIP ViT-L/14 vision tower is vendored into modeling_csd.py and the released CSD weights are stored as model.safetensors.

Not an official release. Original code: https://github.com/learn2phoenix/CSD (MIT). Official checkpoint mirror: https://huggingface.co/tomg-group-umd/CSD-ViT-L (CC-BY-4.0). This repo repackages that checkpoint for AutoModel.from_pretrained.

What it is

A CLIP ViT-L/14 vision backbone (projection removed) whose pre-projection feature (1024-d) is mapped by a learned style head and a content head to 768-d descriptors, each L2-normalized. Style similarity between two images is the cosine of their style embeddings.

Usage

import torch
from PIL import Image
from transformers import AutoModel, AutoImageProcessor

model = AutoModel.from_pretrained("bigshanedogg/CSD", trust_remote_code=True).eval()
proc = AutoImageProcessor.from_pretrained("bigshanedogg/CSD", trust_remote_code=True)

px = proc(images=Image.open("a.png"), return_tensors="pt")["pixel_values"]
out = model(pixel_values=px)
style = out.embeddings            # (1, 768), L2-normalized style descriptor
content = out.content_embeddings  # (1, 768), L2-normalized content descriptor

The image processor resizes the short side to 224 (BICUBIC), center-crops 224, and applies the CLIP mean/std β€” matching the upstream CSD preprocessing.

Licensing

  • Port (modeling/config/processing): MIT β€” Copyright (c) 2026 bigshanedogg.
  • CSD original code: MIT β€” Copyright (c) 2023 the CSD authors (https://github.com/learn2phoenix/CSD).
  • Released CSD weights (model.safetensors, from tomg-group-umd/CSD-ViT-L): CC-BY-4.0 β€” attribute Somepalli et al. / University of Maryland.
  • Vendored ViT tower: MIT β€” Copyright (c) 2021 OpenAI (https://github.com/openai/CLIP), MODIFIED.

See LICENSE for the full notices.

Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support