Instructions to use JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8") model = AutoModelForCausalLM.from_pretrained("JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8
- SGLang
How to use JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8 with Docker Model Runner:
docker model run hf.co/JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8
MegaBeam-Mistral-7B-512k-FP8-W8A8
FP8 (W8A8) quantized version of aws-prototyping/MegaBeam-Mistral-7B-512k.
This checkpoint quantizes both weights and activations to FP8, and additionally uses an FP8 KV cache, cutting memory footprint and improving throughput on Hopper (SM90) and newer GPUs while preserving the model's 512K-token context window.
Model details
| Base model | aws-prototyping/MegaBeam-Mistral-7B-512k |
| Architecture | MistralForCausalLM (7B, 32 layers, GQA with 8 KV heads) |
| Max context | 524,288 tokens (rope_theta = 7.5e7) |
| Quantization | FP8 W8A8 — weights and input activations |
| Strategy | Per-tensor, static, symmetric (minmax observer) |
| KV cache | FP8 (per-tensor, static, symmetric) |
| Ignored modules | lm_head (kept in higher precision) |
| Format | compressed-tensors (float-quantized, v0.13.0) |
| Produced with | llm-compressor |
The full quantization recipe is included in recipe.yaml.
Usage
vLLM (recommended)
vllm serve JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8 \
--kv-cache-dtype fp8 \
--max-model-len 524288
from vllm import LLM, SamplingParams
llm = LLM(model="JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8", kv_cache_dtype="fp8")
out = llm.generate("Summarize the following document:\n...", SamplingParams(max_tokens=256))
print(out[0].outputs[0].text)
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
messages = [{"role": "user", "content": "Hello!"}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
print(tokenizer.decode(model.generate(inputs, max_new_tokens=128)[0]))
Note: FP8 W8A8 inference requires an FP8-capable GPU (NVIDIA Hopper / Ada Lovelace or newer, compute capability ≥ 8.9) and a recent
compressed-tensorsruntime.
Quantization reproduction
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
recipe = QuantizationModifier(
targets="Linear",
scheme="FP8", # W8A8 per-tensor float
ignore=["lm_head"],
kv_cache_scheme={"num_bits": 8, "type": "float", "symmetric": True, "strategy": "tensor"},
)
oneshot(model="aws-prototyping/MegaBeam-Mistral-7B-512k", recipe=recipe)
(See recipe.yaml for the exact configuration used.)
License
Inherits the Apache 2.0 license of the base model. Please also review the base model's card for intended use and limitations.
- Downloads last month
- 884
Model tree for JongYeop/MegaBeam-Mistral-7B-512k-FP8-W8A8
Base model
aws-prototyping/MegaBeam-Mistral-7B-512k