Instructions to use ArtmeScienceLab/Garments2Look-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ArtmeScienceLab/Garments2Look-LoRA with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image-Edit-2509", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("ArtmeScienceLab/Garments2Look-LoRA") prompt = "Turn this cat into a dog" input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png") image = pipe(image=input_image, prompt=prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image-Edit-2509", dtype=torch.bfloat16, device_map="cuda")
pipe.load_lora_weights("ArtmeScienceLab/Garments2Look-LoRA")
prompt = "Turn this cat into a dog"
input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png")
image = pipe(image=input_image, prompt=prompt).images[0]Garments2Look-LoRA
Task-specific DiffSynth LoRA adapters for Qwen-Image-Edit-2509 and Qwen-Image-2.1, trained for outfit-level virtual try-on with clothing and accessories.
This model repository includes the four adapters, a compatible Qwen inference runtime, one unified inference script, and a paired test example. The runtime is a trimmed DiffSynth 2.1.8 implementation used for these results. The script below supports both base models; no GitHub update is required to run it.
Checkpoints
| Base model | Task | File | Training |
|---|---|---|---|
| Qwen-Image-Edit-2509 | Inpainting | Qwen-Image-Edit-2509-LoRA-2-refer-20k-inpainting-epoch-1.safetensors |
20,000 samples; 2 completed epochs; rank 32 |
| Qwen-Image-Edit-2509 | Editing | Qwen-Image-Edit-2509-LoRA-2-refer-20k-editing-epoch-1.safetensors |
20,000 samples; 2 completed epochs; rank 32 |
| Qwen-Image-2.1 | Inpainting | Qwen-Image-2.1-LoRA-2-refer-97068-inpainting-epoch-0.safetensors |
97,068 samples; 1 completed epoch; final step 48,534; rank 32 |
| Qwen-Image-2.1 | Editing | Qwen-Image-2.1-LoRA-2-refer-97068-editing-epoch-0.safetensors |
97,068 samples; 1 completed epoch; final step 48,534; rank 32 |
Epoch numbers are zero-indexed: epoch-1 means two completed epochs, and epoch-0 means one completed epoch. Both 2.1 files are the final step-48534 checkpoints, not step-48000. These are adapters, not full models. Use the matching base model and task. File checksums and training metadata are in manifest.json.
Inputs
Both tasks use the same two-reference interface:
| Input | Inpainting | Editing |
|---|---|---|
| Figure 1 | Person with clothing regions already masked gray (128) | Source person wearing an existing outfit |
| Figure 2 | OOTD collage of target products | OOTD collage of target products |
| Prompt | Numbered items, styling, and layering order | Numbered items, styling, and layering order |
No separate mask argument is passed to inference. Match the numbered items to the collage. For dataset-based inpainting, use the updated v3 masks with dilation. The dataset has 98,012 total records: 97,068 training and 944 test records, with v1.1 annotations and editing source images. See the dataset card for preparation.
Installation and download
Tested with Python 3.10, PyTorch 2.7.1 / CUDA 12.8, Transformers 5.16.1, and H200 GPUs. Run all commands below from the same working directory. Download only the base model you intend to use.
conda create -n g2l-lora python=3.10 -y
conda activate g2l-lora
python -m pip install torch==2.7.1 torchvision==0.22.1 --index-url https://download.pytorch.org/whl/cu128
python -m pip install huggingface_hub
hf download ArtmeScienceLab/Garments2Look-LoRA --local-dir models/Garments2Look-LoRA
python -m zipfile -e models/Garments2Look-LoRA/qwen-runtime.zip models/Garments2Look-LoRA
python -m pip install -e models/Garments2Look-LoRA/qwen-runtime
python -m pip install transformers==5.16.1
hf download Qwen/Qwen-Image-Edit-2509 --local-dir models/Qwen-Image-Edit-2509
hf download Qwen/Qwen-Image-2.1 --local-dir models/Qwen-Image-2.1
The bundled script imports the extracted runtime. Adapters use the DiffSynth format; compatibility with other LoRA loaders has not been verified. The Hub diffusers library tag enables download tracking for the root-level safetensors files; it does not indicate verified Diffusers loader compatibility. Set CUDA_VISIBLE_DEVICES to choose a GPU.
Inpainting inference
Qwen-Image-Edit-2509
python models/Garments2Look-LoRA/inference.py --model-version 2509 --task inpainting \
--model-dir models/Qwen-Image-Edit-2509 \
--lora models/Garments2Look-LoRA/Qwen-Image-Edit-2509-LoRA-2-refer-20k-inpainting-epoch-1.safetensors \
--origin models/Garments2Look-LoRA/examples/qwen2509/input-inpainting.png \
--ootd models/Garments2Look-LoRA/examples/qwen2509/ootd.png \
--prompt-file models/Garments2Look-LoRA/examples/qwen2509/prompt-inpainting.txt \
--output output/qwen2509-inpainting.png --seed 0 --steps 40 --cfg-scale 4
Qwen-Image-2.1
python models/Garments2Look-LoRA/inference.py --model-version 2.1 --task inpainting \
--model-dir models/Qwen-Image-2.1 \
--lora models/Garments2Look-LoRA/Qwen-Image-2.1-LoRA-2-refer-97068-inpainting-epoch-0.safetensors \
--origin models/Garments2Look-LoRA/examples/qwen21/input-inpainting.png \
--ootd models/Garments2Look-LoRA/examples/qwen21/ootd.png \
--prompt-file models/Garments2Look-LoRA/examples/qwen21/prompt-inpainting.txt \
--output output/qwen21-inpainting.png --seed 0 --steps 40 --cfg-scale 1
Editing inference
Qwen-Image-Edit-2509
python models/Garments2Look-LoRA/inference.py --model-version 2509 --task editing \
--model-dir models/Qwen-Image-Edit-2509 \
--lora models/Garments2Look-LoRA/Qwen-Image-Edit-2509-LoRA-2-refer-20k-editing-epoch-1.safetensors \
--origin models/Garments2Look-LoRA/examples/qwen2509/input-editing.png \
--ootd models/Garments2Look-LoRA/examples/qwen2509/ootd.png \
--prompt-file models/Garments2Look-LoRA/examples/qwen2509/prompt-editing.txt \
--output output/qwen2509-editing.png --seed 0 --steps 40 --cfg-scale 4
Qwen-Image-2.1
python models/Garments2Look-LoRA/inference.py --model-version 2.1 --task editing \
--model-dir models/Qwen-Image-2.1 \
--lora models/Garments2Look-LoRA/Qwen-Image-2.1-LoRA-2-refer-97068-editing-epoch-0.safetensors \
--origin models/Garments2Look-LoRA/examples/qwen21/input-editing.png \
--ootd models/Garments2Look-LoRA/examples/qwen21/ootd.png \
--prompt-file models/Garments2Look-LoRA/examples/qwen21/prompt-editing.txt \
--output output/qwen21-editing.png --seed 0 --steps 40 --cfg-scale 1
Default seed is 0, with 40 steps. If CFG is omitted, the script uses 4 for 2509 and 1 for 2.1. Images use a 1,048,576-pixel budget, aligned to multiples of 16 for 2509 and 32 for 2.1. Each invocation saves its PNG and a JSON parameter record. For custom outfits, replace Figure 1, the collage, and the full prompt.
Paired example: P00958796_b1
This five-item test outfit includes a top, jacket, shorts, sandals, and bag. Styling specifies a tucked-in top and an unbuttoned jacket; layering is top β jacket. Both task prompts are supplied in the example folders.
Qwen-Image-Edit-2509 LoRA
Qwen-Image-2.1 LoRA
Columns: OOTD, inpainting input, inpainting output, editing input, editing output. Full task prompts are printed below each figure. Both models use seed 0 and 40 steps, with CFG 4 for 2509 and CFG 1 for 2.1. The 2.1 images use the final adapters released above.
Full prompt (both tasks):
Keep the woman's identity, pose, background in Figure 1 unchanged, wearing the outfit in Figure 2, include (1) a top (tucked-in), (2) a jacket (unbuttoned), (3) shorts, (4) sandals, (5) a bag. Layering Order: (1) -> (2).
This example was selected after comparing six candidate outfits at fixed seed 0, four outputs per candidate. It illustrates a selected successful case, not aggregate test performance. The two base models were trained on different data counts and epoch counts. Bag shape, garment hems, and pose may still differ. Selection provenance is in examples/selection.json; each output has a parameter JSON, and target images are included.
Training details
| Setting | 2509 | 2.1 |
|---|---|---|
| Training examples | 20,000 | Full 97,068 training split |
| Completed epochs | 2 | 1 |
| LoRA rank | 32 | 32 |
| Learning rate | 1e-4 | 1e-4 |
| Pixel budget | 1,048,576 | 1,048,576 |
| Precision | BF16 | BF16 |
| Gradient checkpointing | Enabled | Enabled |
| Final step | Epoch-based checkpoint | 48,534 |
The 2.1 runs used two H200 GPUs and global batch size 2. Training took approximately 58h36m (inpainting) and 58h50m (editing). These adapters are later releases than the paper's rebuttal-stage models; the selected examples do not establish aggregate improvements over those models.
For dataset preparation and the public 2509 training workflow, see the code repository. The bundled runtime and script here reproduce inference for both models; this package does not include a complete 2.1 training launcher.
Repository layout
Garments2Look-LoRA/
βββ README.md
βββ manifest.json
βββ inference.py
βββ qwen-runtime.zip
βββ examples/
β βββ selection.json
β βββ qwen2509/ # inputs, prompts, target, outputs, parameters, comparison
β βββ qwen21/ # same outfit; final 2.1 outputs
βββ Qwen-Image-Edit-2509-LoRA-2-refer-20k-*-epoch-1.safetensors
βββ Qwen-Image-2.1-LoRA-2-refer-97068-*-epoch-0.safetensors
Limitations and license
Fine accessory details, garment fidelity, styling, layering, and pose preservation vary with input and seed. The showcased example is selected, not an aggregate evaluation. Full base models are required, and their respective licenses apply separately. The bundled DiffSynth code retains its upstream license; an explicit adapter license has not yet been specified.
Citation
If you use our dataset or models in your research, please consider citing our paper:
@inproceedings{cvpr2026garments2look,
title={Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories},
author={Hu, Junyao and Cheng, Zhongwei and Wong, Waikeung and Zou, Xingxing},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2026}
}
- Downloads last month
- 19
Model tree for ArtmeScienceLab/Garments2Look-LoRA
Base model
Qwen/Qwen-Image-2.1

# Gated model: Login with a HF token with gated access permission hf auth login