Instructions to use bigshanedogg/CSD with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bigshanedogg/CSD with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-feature-extraction", model="bigshanedogg/CSD", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("bigshanedogg/CSD", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
CSD (ViT-L/14) β HuggingFace format
Unofficial transformers-format port of the CSD style model from "Measuring Style
Similarity in Diffusion Models" (Somepalli, Gupta, Gupta, Shrivastava, Goldstein, Feizi;
2024). Loads via trust_remote_code with no clip / open_clip runtime dependency β
the OpenAI CLIP ViT-L/14 vision tower is vendored into modeling_csd.py and the released
CSD weights are stored as model.safetensors.
Not an official release. Original code: https://github.com/learn2phoenix/CSD (MIT). Official checkpoint mirror: https://huggingface.co/tomg-group-umd/CSD-ViT-L (CC-BY-4.0). This repo repackages that checkpoint for
AutoModel.from_pretrained.
What it is
A CLIP ViT-L/14 vision backbone (projection removed) whose pre-projection feature (1024-d) is mapped by a learned style head and a content head to 768-d descriptors, each L2-normalized. Style similarity between two images is the cosine of their style embeddings.
Usage
import torch
from PIL import Image
from transformers import AutoModel, AutoImageProcessor
model = AutoModel.from_pretrained("bigshanedogg/CSD", trust_remote_code=True).eval()
proc = AutoImageProcessor.from_pretrained("bigshanedogg/CSD", trust_remote_code=True)
px = proc(images=Image.open("a.png"), return_tensors="pt")["pixel_values"]
out = model(pixel_values=px)
style = out.embeddings # (1, 768), L2-normalized style descriptor
content = out.content_embeddings # (1, 768), L2-normalized content descriptor
The image processor resizes the short side to 224 (BICUBIC), center-crops 224, and applies the CLIP mean/std β matching the upstream CSD preprocessing.
Licensing
- Port (modeling/config/processing): MIT β Copyright (c) 2026 bigshanedogg.
- CSD original code: MIT β Copyright (c) 2023 the CSD authors (https://github.com/learn2phoenix/CSD).
- Released CSD weights (
model.safetensors, fromtomg-group-umd/CSD-ViT-L): CC-BY-4.0 β attribute Somepalli et al. / University of Maryland. - Vendored ViT tower: MIT β Copyright (c) 2021 OpenAI (https://github.com/openai/CLIP), MODIFIED.
See LICENSE for the full notices.
- Downloads last month
- -