Deploying a 1-Bit Bonsai-27B Model with PrismML llama.cpp and OpenAI-Compatible Local Inference Workflows
In this tutorial, we deploy the 1-bit Bonsai-27B language model using the PrismML fork of llama.cpp, which provides the specialized CUDA kernels required to decode the model’s Q1_0_g128 GGUF quantization format. We begin by validating the GPU runtime, installing the required Python dependencies, compiling the CUDA-enabled inference binaries, and downloading the compressed model weights from Hugging Face. We then test the model through llama-cli, launch an OpenAI-compatible local inference server, and interact with it through a reusable Python client that supports standard completions, streamed responses, multi-turn conversations, and code generation. We also examine optional configurations for throughput benchmarking, quantized key-value caching, long-context inference, speculative decoding, and multimodal extensions. Copy CodeCopiedUse a different Browser import os import sys import time import json import shutil import subprocess import multiprocessing WORK_DIR = “/content” REPO_URL = “https://github.com/PrismML-Eng/llama.cpp” REPO_DIR = os.path.join(WORK_DIR, “llama.cpp”) BUILD_DIR = os.path.join(REPO_DIR, “build”) BIN_DIR = os.path.join(BUILD_DIR, “bin”) HF_REPO = “prism-ml/Bonsai-27B-gguf” MODEL_FILE = “Bonsai-27B-Q1_0.gguf” MODEL_PATH = os.path.join(WORK_DIR, MODEL_FILE) SERVER_HOST = “127.0.0.1” SERVER_PORT = 8080 SERVER_URL = f”http://{SERVER_HOST}:{SERVER_PORT}” GEN_PARAMS = {“temperature”: 0.7, “top_p”: 0.95, “top_k”: 20} CTX_SIZE = 8192 N_GPU_LAYERS = 99 USE_KV_Q4 = False def sh(cmd, check=True, **kw): “””Run a shell command, streaming output to the notebook.””” print(f”n$ {cmd}”) return subprocess.run(cmd, shell=True, check=check, **kw) print(“=” * 70) print(“[1/7] Checking environment”) print(“=” * 70) gpu = subprocess.run(“nvidia-smi –query-gpu=name,memory.total –format=csv,noheader”, shell=True, capture_output=True, text=True) if gpu.returncode != 0: sys.exit(“No GPU detected. In Colab: Runtime -> Change runtime type -> GPU (T4).”) print(f”GPU detected: {gpu.stdout.strip()}”) print(“Bonsai-27B needs only ~5.2 GB peak at 4K context — any Colab GPU works.”) sh(“pip -q install huggingface_hub requests”) We configure the Colab workspace, model repository, server endpoint, inference parameters, context size, and GPU offloading settings required throughout the tutorial. We define a reusable shell-command function and verify that the runtime exposes a compatible NVIDIA GPU before continuing. We then install the Hugging Face Hub and HTTP client dependencies needed for model retrieval and API communication. Copy CodeCopiedUse a different Browser print(“=” * 70) print(“[2/7] Building PrismML llama.cpp fork with CUDA (cached after 1st run)”) print(“=” * 70) if not os.path.isdir(REPO_DIR): sh(f”git clone –depth 1 {REPO_URL} {REPO_DIR}”) else: print(“Repo already cloned — skipping.”) cli_bin = os.path.join(BIN_DIR, “llama-cli”) server_bin = os.path.join(BIN_DIR, “llama-server”) bench_bin = os.path.join(BIN_DIR, “llama-bench”) if not (os.path.exists(cli_bin) and os.path.exists(server_bin)): jobs = multiprocessing.cpu_count() sh(f”cmake -S {REPO_DIR} -B {BUILD_DIR} -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release”) sh(f”cmake –build {BUILD_DIR} -j{jobs} –target llama-cli llama-server llama-bench”) else: print(“Binaries already built — skipping.”) We clone the PrismML fork of llama.cpp, which provides the specialized kernels required for the model’s Q1_0_g128 quantization format. We configure a CUDA-enabled release build with CMake and compile the command-line, server, and benchmarking executables. We also reuse previously generated binaries when they already exist, reducing repeated setup time in the same Colab session. Copy CodeCopiedUse a different Browser print(“=” * 70) print(“[3/7] Downloading weights from Hugging Face”) print(“=” * 70) from huggingface_hub import hf_hub_download if not os.path.exists(MODEL_PATH): downloaded = hf_hub_download(repo_id=HF_REPO, filename=MODEL_FILE, local_dir=WORK_DIR) print(f”Downloaded to: {downloaded}”) else: print(“Model already on disk — skipping.”) print(f”Model size on disk: {os.path.getsize(MODEL_PATH) / 1e9:.2f} GB”) We connect to the Hugging Face Hub and download the Bonsai-27B GGUF model into the Colab workspace. We skip the transfer when the model file is already available locally, allowing subsequent runs to proceed more efficiently. We then calculate and display the deployed model size to confirm that the compressed weights are stored correctly. Copy CodeCopiedUse a different Browser print(“=” * 70) print(“[4/7] Smoke test with llama-cli”) print(“=” * 70) sh( f'{cli_bin} -m {MODEL_PATH} ‘ f’-p “Explain in two sentences why 1-bit quantization saves memory.” ‘ f’-n 128 -ngl {N_GPU_LAYERS} ‘ f’–temp {GEN_PARAMS[“temperature”]} ‘ f’–top-p {GEN_PARAMS[“top_p”]} –top-k {GEN_PARAMS[“top_k”]} ‘ f’-no-cnv 2>/dev/null’, check=False, ) print(“=” * 70) print(“[5/7] Starting llama-server (OpenAI-compatible API)”) print(“=” * 70) import requests kv_flags = “-ctk q4_0 -ctv q4_0” if USE_KV_Q4 else “” server_cmd = ( f”{server_bin} -m {MODEL_PATH} ” f”–host {SERVER_HOST} –port {SERVER_PORT} ” f”-ngl {N_GPU_LAYERS} -c {CTX_SIZE} {kv_flags}” ) print(f”$ {server_cmd} (background)”) server_log = open(os.path.join(WORK_DIR, “server.log”), “w”) server_proc = subprocess.Popen(server_cmd, shell=True, stdout=server_log, stderr=server_log) for _ in range(120): try: if requests.get(f”{SERVER_URL}/health”, timeout=2).status_code == 200: print(“Server is up.”) break except requests.exceptions.RequestException: pass time.sleep(2) else: server_proc.kill() sys.exit(“Server failed to start — check /content/server.log”) We perform a command-line smoke test to verify that the compiled runtime can load the quantized model and generate a valid response. We then start llama-server with full GPU layer offloading, the selected context window, and optional quantized KV-cache settings. We repeatedly query the health endpoint until the OpenAI-compatible inference service becomes ready for client requests. Copy CodeCopiedUse a different Browser print(“=” * 70) print(“[6/7] Talking to Bonsai-27B via the OpenAI-compatible API”) print(“=” * 70) def chat(messages, stream=False, max_tokens=512, **overrides): “””Minimal OpenAI-compatible chat client for the local llama-server.””” payload = { “model”: “bonsai-27b”, “messages”: messages, “max_tokens”: max_tokens, “stream”: stream, **GEN_PARAMS, **overrides, } if not stream: r = requests.post(f”{SERVER_URL}/v1/chat/completions”, json=payload) r.raise_for_status() return r.json()[“choices”][0][“message”][“content”] r = requests.post(f”{SERVER_URL}/v1/chat/completions”, json=payload, stream=True) r.raise_for_status() full = [] for line in r.iter_lines(): if not line or not line.startswith(b”data: “): continue chunk = line[len(b”data: “):] if chunk == b”[DONE]”: break delta = json.loads(chunk)[“choices”][0][“delta”].get(“content”, “”) full.append(delta) print(delta, end=””, flush=True) print() return “”.join(full) SYSTEM = {“role”: “system”, “content”: “You are a helpful assistant”} print(“n— 6a: basic completion —“) answer = chat([SYSTEM, {“role”: “user”, “content”: “What is the capital of France? One sentence.”}]) print(answer) print(“n— 6b: math reasoning, streamed token-by-token —“) chat([SYSTEM, {“role”: “user”, “content”: “A train travels 120 km at 80 km/h, then 90 km at ” “60 km/h. What is its average speed for the whole ” “trip? Show your reasoning briefly.”}], stream=True, max_tokens=700) print(“n— 6c: multi-turn chat —“) history = [SYSTEM] for user_msg in [“My name is Priya and I love graph algorithms.”, “Suggest one project idea that combines my interest with LLMs.”, “What was my name again?”]: history.append({“role”: “user”, “content”: user_msg}) reply = chat(history, max_tokens=300) history.append({“role”: “assistant”, “content”: reply}) print(f”nUSER: {user_msg}nBONSAI: {reply}”) print(“n— 6d: code generation —“) print(chat([SYSTEM, {“role”: “user”, “content”: “Write a Python function that returns the n-th ” “Fibonacci number using memoization. Code only.”}], max_tokens=400)) We define a reusable Python chat client that sends OpenAI-compatible requests to the locally hosted Bonsai-27B server. We support both standard


