Spaces:
Paused
title: Bonsai {1, 1.58}-bit GPU
emoji: 🌿
colorFrom: green
colorTo: blue
sdk: docker
app_port: 7860
suggested_hardware: l40sx1
pinned: true
short_description: Run {1, 1.58}-bit Bonsai LLMs on GPUs
models:
- prism-ml/Bonsai-8B-gguf
- prism-ml/Bonsai-4B-gguf
- prism-ml/Bonsai-1.7B-gguf
- prism-ml/Ternary-Bonsai-8B-gguf
- prism-ml/Ternary-Bonsai-4B-gguf
- prism-ml/Ternary-Bonsai-1.7B-gguf
Bonsai Demo
Interactive demo for Bonsai — the first commercially viable 1-bit LLMs, by PrismML.
Bonsai models run at true 1-bit precision — every weight is a single bit. An 8B model fits in 1.15 GB, a 1.7B model in just 240 MB. Small enough to run in a browser, on a phone, or on any laptop — while remaining competitive with full-precision models on benchmarks.
Ternary-Bonsai is the 1.58-bit sibling series — each weight is one of {−1, 0, +1}. It trades a bit of size for a quality bump over the pure 1-bit models. An 8B Ternary-Bonsai fits in 2.03 GB, a 1.7B in just 430 MB.
This demo will be available for a limited time. Enjoy it while it lasts!
Highlights
Bonsai-8B fits in 1.15 GB (14x smaller than FP16) and generates at ~330 tok/s on an L40S (6.3x faster than FP16). Scores 70.5 average across 6 benchmark tasks, competitive with full-precision 8B models.
Ternary-Bonsai-8B (1.58-bit, 2.03 GB) is also available in the chat UI — pick it from the model dropdown for a quality bump at a modest size increase over the 1-bit models.
Tool calling & MCP: the chat UI supports OpenAI-style tool calling and comes with Brave Search and ArXiv MCP servers pre-wired — open the MCP tab in the chat settings to use them.
Models
1-bit Bonsai
| Model | Size | GGUF | MLX |
|---|---|---|---|
| Bonsai-8B | 1.15 GB | prism-ml/Bonsai-8B-gguf | prism-ml/Bonsai-8B-mlx-1bit |
| Bonsai-4B | 570 MB | prism-ml/Bonsai-4B-gguf | prism-ml/Bonsai-4B-mlx-1bit |
| Bonsai-1.7B | 240 MB | prism-ml/Bonsai-1.7B-gguf | prism-ml/Bonsai-1.7B-mlx-1bit |
1.58-bit Ternary-Bonsai
| Model | Size | GGUF |
|---|---|---|
| Ternary-Bonsai-8B | 2.03 GB | prism-ml/Ternary-Bonsai-8B-gguf |
| Ternary-Bonsai-4B | 1.00 GB | prism-ml/Ternary-Bonsai-4B-gguf |
| Ternary-Bonsai-1.7B | 430 MB | prism-ml/Ternary-Bonsai-1.7B-gguf |
Resources
- 1-bit Bonsai Whitepaper
- Google Colab notebook
- GitHub Demo
- Discord community
- Prism ML website
- Run Bonsai locally in the browser (WebGPU)
Privacy
- We do not log any messages. Chat content is never stored on the server.
- This demo uses the built-in llama-server UI, which saves your conversation history in your browser's local storage only. Clearing your browser cache will erase it.
- That said, please do not submit sensitive, private, or confidential information in your messages.
Fair Use
We've allocated multiple GPUs to keep this demo responsive, but resources are shared across all users. Under heavy load you may experience slower responses or brief queuing. Please be mindful of usage and avoid sending large bursts of automated requests so everyone can enjoy the demo.