Instructions to use lingshu-medical-mllm/Lingshu-32B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lingshu-medical-mllm/Lingshu-32B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="lingshu-medical-mllm/Lingshu-32B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("lingshu-medical-mllm/Lingshu-32B") model = AutoModelForMultimodalLM.from_pretrained("lingshu-medical-mllm/Lingshu-32B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use lingshu-medical-mllm/Lingshu-32B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "lingshu-medical-mllm/Lingshu-32B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lingshu-medical-mllm/Lingshu-32B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/lingshu-medical-mllm/Lingshu-32B
- SGLang
How to use lingshu-medical-mllm/Lingshu-32B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "lingshu-medical-mllm/Lingshu-32B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lingshu-medical-mllm/Lingshu-32B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "lingshu-medical-mllm/Lingshu-32B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lingshu-medical-mllm/Lingshu-32B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use lingshu-medical-mllm/Lingshu-32B with Docker Model Runner:
docker model run hf.co/lingshu-medical-mllm/Lingshu-32B
Some wrong when load the 32B model
I am encountering an issue when attempting to load the 32B model. While loading the 7B model works without any problems, trying to load the 32B model results in the following error:
safetensors_rust.SafetensorError: Error while deserializing header: InvalidHeaderDeserialization
I have already re-downloaded the official model files from Hugging Face twice to rule out any download corruption, but the error persists. This leads me to suspect that the model upload might be incomplete or corrupted on the source side.
Could anyone please help me troubleshoot this issue or confirm whether the 32B model files are intact?
Thank you in advance for your support.
I am encountering an issue when attempting to load the 32B model. While loading the 7B model works without any problems, trying to load the 32B model results in the following error:
safetensors_rust.SafetensorError: Error while deserializing header: InvalidHeaderDeserialization
I have already re-downloaded the official model files from Hugging Face twice to rule out any download corruption, but the error persists. This leads me to suspect that the model upload might be incomplete or corrupted on the source side.Could anyone please help me troubleshoot this issue or confirm whether the 32B model files are intact?
Thank you in advance for your support.
what is your transformer version? Some transformers versions are not compatible with our model. You may need to use transformers<=4.52.1