For AI agents (and humans) helping someone set up this demo. Goal: pick the right flags for the user's hardware and use case. The behavior notes below come from real testing; measure on the user's own hardware before promising performance (the timings object in every API response has the numbers).
The 27B generation is a step change over the earlier 8B/4B/1.7B demos:
--jinja) and mlx_lm.server emit native OpenAI tool_calls, verified with full tool round-trips.reasoning_content) and can be budgeted (--reasoning-budget N) or picked per chat in the web UI.BONSAI_FAMILY / BONSAI_MODEL | Weights | Notes |
|---|---|---|
bonsai2 / 27B (default) | GGUF 5.9-7.2 GB + mmproj 0.63 GB; MLX 2-bit 8.6 GB | The current generation and what plain ./setup.sh installs. Two GGUF bands, both requiring this demo's binaries: PTQ1_0 (1.75 bpw, 5.9 GB, densely packed trits, smallest) and PQ2_0 (2.13 bpw, 7.2 GB, faster prompt processing, what the scripts download). There is no mainline-compatible band; see the Q2_0 warning below. The MLX pack carries its own Hadamard-aware loader in runtime/ plus the vision tower, and runs on stock mlx-vlm from .venv-vlm, not the fork in .venv |
ternary / 27B (previous generation) | GGUF ~6.7-7.1 GB + mmproj 0.9 GB; MLX 2-bit ~7.9 GB | Superseded by bonsai2. Two GGUF formats since the mainline rebase (prism-b10658+): PQ2_0 (group 128, 6.66 GiB, smallest/fastest where supported: CUDA, Metal, CPU, ROCm) and official Q2_0 group 64 (Ternary-Bonsai-27B-Q2_g64.gguf, 7.05 GiB, adds Vulkan/SYCL); smaller sizes use *-Q2_0_g64.gguf naming. The scripts pick per backend. Legacy *-Q2_0.gguf files (no g64) only load on old prism-v5 releases; new binaries refuse them with an error |
bonsai / 27B | GGUF Q1_0 ~3.5 GB + mmproj 0.9 GB; MLX 1-bit ~4.8 GB | Smallest and fastest; fits on a modern iPhone without offloading |
8B / 4B / 1.7B (both families) | smaller | Text-only, no tools wiring, legacy tested flag set |
Every Bonsai 2 band stores its weights in a rotated basis and needs the activation transform that only this demo's binaries have. PQ2_0 and PTQ1_0 are rejected outright by mainline as unknown types, which is the safe failure. Q2_0 is the dangerous one: mainline loads it without a warning and produces gibberish, because Q2_0 is a type upstream already knows and qwen35 is a supported architecture, so nothing in the file makes it refuse.
For that reason the Bonsai 2 Q2_0 band is not in the model repo. It lives on its own at Ternary-Bonsai-2-27B-gguf-dev, named Ternary-Bonsai-2-27B-Q2_0-prism-fork-required.gguf, published for testing and for the work to upstream the Hadamard changes. Do not point a user at it for running the model. It moves into the main repo once mainline can run it.
If someone reports gibberish from Bonsai 2, check which binary they used before anything else.
All three 27B families have the same capabilities (vision, tools, thinking, long context) — they differ in size and speed. Context: 262,144 tokens max; FP16 KV cache is 64 KiB/token (~6.3 GiB at 100K), so 100K context fits on many consumer devices — full peak-memory table in the README's Context Size section. All 27B repos: https://huggingface.co/collections/prism-ml/bonsai-27b. All model repos are public; no token needed.
All extra args pass straight through the start scripts, e.g. ./scripts/start_llama_server.sh --image-max-tokens 1024 -ub 1024.
| Knob | What it does | Trade-off |
|---|---|---|
BONSAI_SPECULATIVE=1 (env, start_llama_server.sh only) | Experimental. Loads the paired DSpark drafter for speculative decoding (--spec-type draft-dspark). Post-migration (prism-b10658+) the drafter is the converted *dspark-dflash* sidecar (~0.6 GiB; embedding/lm_head shared with the target — one-time conversion documented in SPECULATIVE.md). Measured on an L40S: ternary 27B 1.8-2.4x decode (2.06x blended), 1-bit 27B 1.4-1.75x (1.60x blended). Off by default. | The drafter path is stable and fast on CUDA. On Apple Silicon (Metal) it only pays off for ternary code/math (~1.2x) and is a net slowdown on chat/reasoning and for the 1-bit family, so do not recommend it on Macs. Disables cross-request prompt-cache reuse (every turn re-prefills) and forces single-slot (-np 1); worse for multi-turn and the agentic Open WebUI path, which is why it is server-only and opt-in. The prebuilt binaries include both the dspark-capable llama-server and the CLI one-shot llama-speculative-simple. Details: SPECULATIVE.md. |
BONSAI_KV4=1 (env, start_llama_server.sh only) | Experimental. Q4_0 (4-bit) KV cache, ~3.5x smaller KV memory. Optional quality booster: ./scripts/make_kv_bias.sh builds a model-specific mean-centering bias (tiny calibration corpus is enough; users can pass their own text) that the server picks up automatically. | Memory tool, not a speed tool: decode is slightly slower than F16 KV. The 27B's hybrid attention already keeps KV small, so only reach for this at very long contexts on tight machines. The bias is calibrated with K-rotation off and the script/server handle the matching flags automatically; llama.cpp backend only. Details: KV-CACHE.md. |
--reasoning-budget N | Caps thinking at N tokens (default -1 = unlimited) | Middle ground; pair with --reasoning-budget-message |
--image-max-tokens N | Downscales images to ~N vision tokens (1 token ~ 32x32 px). Model allows ~4096 (=4.2 MP) | The scripts default this to 1024 on Metal / Vulkan / CPU and leave CUDA/ROCm uncapped; override with BONSAI_IMAGE_MAX_TOKENS (0 = uncapped). Loses fine detail (small text / OCR) on large images; images under the cap are unaffected |
BONSAI_NGL=N (env) | Overrides GPU layer offload (auto-detect defaults to all layers when CUDA/ROCm/Vulkan/Metal tooling is present; 0 = CPU-only) | The auto-detect keys on installed tooling, not GPU capability: a CPU box with only an integrated GPU and Vulkan drivers will offload to that iGPU. Capable iGPUs (e.g. Strix Halo) genuinely benefit, weak ones are better on CPU; recommend BONSAI_NGL=0 when decode is slower than expected on iGPU-only machines |
-ub N | Prefill microbatch (default 512) | Sometimes faster prefill on Metal at 1024; measure |
BONSAI_CTX=N (env) | Overrides the default context. The scripts default to a RAM-tiered size (8192 up to 131072 by machine memory) for predictable memory use; 0 (or unset) resolves to that same tiered size — it is not passed to llama.cpp as -c 0 (which would use the model's full 262k training context and OOM constrained machines). Pass an explicit number to force full context | Bigger context = more KV memory at 64 KiB/token FP16; pair huge contexts with BONSAI_KV4=1 |
BONSAI_MMPROJ_CPU=1 (env, 27B) | Adds --no-mmproj-offload so the vision projector stays in system RAM instead of VRAM. Off by default | Frees ~0.9 GiB of VRAM for KV/context on tight cards; cost is a slower image prefill (CPU vision encode, tens of ms to seconds/image). Token generation and text-only requests are unaffected, so it's a good trade when images are occasional |
--parallel N | Server slots (default 4) | More slots = more concurrent users, same total context pool |
BRAVE_API_KEY (env, Open WebUI) | Makes a Brave Search MCP server available next to the preconfigured Hugging Face + DeepWiki ones | Needs npm i -g @brave/brave-search-mcp-server; skipped otherwise. Limited to web/news/summarizer (~2.9k tokens); its full tool set is ~29k (brave_place_search alone ~20k). Override with BONSAI_BRAVE_TOOLS. |
BONSAI_CODE_INTERPRETER=0 (env, Open WebUI) | Disables the server-side Jupyter code interpreter (plots via matplotlib, data via pandas/numpy, market data via yfinance). On by default. | Off falls back to browser Pyodide: plots still work, but no yfinance/network. The Jupyter stack (.venv-jupyter) is built by setup.sh. |
Thinking is extracted into reasoning_content by default (collapsible in UIs).
Everything below is per-conversation, in Settings, no server restart:
thinking_budget_tokens per request (0 = off, N = cap, -1 = unlimited).--webui-config-file scripts/webui-config.json, which seeds Hugging Face Hub and DeepWiki as admin defaults (verified served via /props.webui_settings). Config enabled:true only makes them available in the message-box MCP selector; tool schemas are sent only for chats that turn a server on. The new-chat screen's toggle becomes the user's default for future chats (stored in the browser) — that is the knob that trades the prompt tokens below for always-on tools. On a fast GPU always-on is cheap (one-time cost per server run). Add more servers in Settings -> MCP Client or in the JSON (add "useProxy": true if one rejects browser CORS — llama-server ships a /cors-proxy). agenticMaxTurns (default 10) bounds the tool loop.
MCP prompt cost — every enabled server's tool schemas are rendered into the system prompt of each chat, so they add prefill before the first token. Measured from the live servers (2026-07; will drift as they update their tools):
| MCP server | Tools | Approx. prompt tokens |
|---|---|---|
| Hugging Face | 8 | ~2,600 (biggest: hf_fs ~670, hub_repo_details ~490, hub_repo_search ~440) |
| DeepWiki | 3 | ~400 |
| Brave Search | 3 (web/news/summarizer) | ~2,900 (full 8-tool set is ~29k; brave_place_search alone ~20k - demo limits it via --enabled-tools) |
The cost is paid once per server run, not per chat: the schemas are a stable prompt prefix and llama-server reuses the cached prefix across new chats as long as the enabled set (and order) does not change. On slow hardware prefer a small fixed subset over toggling servers between chats. Same logic in the Open WebUI demo: the servers are connected but not attached to the model - pick them per chat from the tool menu, or add server:mcp:<id> to the model's toolIds to make one permanent.
scripts/webui-config.json (systemMessage)./props).Note: UI settings live in the browser (localStorage), so they are per-user/per-machine — server-side pre-seeding of tools is only available in the Open WebUI demo. This also means browser state outlives config changes: if every NEW chat prefills thousands of tokens, the user toggled an MCP server on from the new-chat screen (that saves it as their default for future chats) - toggle it off there, or clear site data for localhost:8080; an incognito window shows what a fresh browser would get.
Agent: tell the user about the image-token cap. If they are on Metal / Vulkan / CPU, say that large images are downscaled to ~1024 vision tokens by default (fast, but fine detail in big images is lost) and ask which they prefer:
BONSAI_IMAGE_MAX_TOKENS=0 — full image detail (best for OCR / screenshots / small text), noticeably slower per large image on consumer hardware. On CUDA/ROCm there is nothing to decide — images run uncapped by default.Full guide with entry examples: TOOLS.md (repo root). The essentials:
start_openwebui.sh is the pattern).mcpServers JSON-string in scripts/webui-config.json ("enabled":true = listed in the per-chat selector, chats still opt in individually; "useProxy":true for CORS-rejecting servers).TOOL_SERVER_CONNECTIONS in scripts/start_openwebui.sh and restart. Do NOT use the admin panel - ENABLE_PERSISTENT_CONFIG=false means panel edits are lost on restart; the script is the source of truth. For bearer tokens set "auth_type":"bearer" and inject the secret at runtime from an environment variable or a gitignored file (mirror the .brave_key handling below); never write a literal token into the tracked script. To attach a server to every chat, add server:mcp:<id> to the model toolIds in scripts/openwebui/seed_openwebui.py (mind the prompt-token table above).Optional web search (Brave) - needs a Brave Search API key; the key stays local (never committed). Install the bridge once: npm i -g @brave/brave-search-mcp-server.
.brave_key file (or BRAVE_API_KEY env) and run start_openwebui.sh - it auto-starts the bridge on 127.0.0.1:8001 and adds the brave MCP (per-chat opt-in). Tell the user this is how they get web search.BRAVE_API_KEY=... brave-search-mcp-server --transport http --host 127.0.0.1 --port 8001) and add http://127.0.0.1:8001/mcp in Settings -> MCP Client. See TOOLS.md../scripts/start_llama_server.sh --reasoning-budget 2048 keeps most of the quality while bounding latency; the web UI's Reasoning-effort picker (Off ... Max) does the same per chat. (Default stays uncapped; these are user choices, not shipped defaults.)--mmproj-offload).start_mlx_server.sh uses the stock-mlx .venv-vlm that setup.sh creates; BONSAI_MLX_VLM=0 opts out). The binary 27B MLX should support vision the same way (the vision tower is full precision in both packs) — it just hasn't been wired through / verified in these scripts yet.error compiling source / command-buffer status 5, set GGML_METAL_TENSOR_DISABLE=1 (README Appendix — FAQ has details). Keep -ngl on GPU; don't reach for BONSAI_NGL=0.# server capabilities and effective context/slots
curl -s http://localhost:8080/props | python3 -m json.tool | head -30
# timing any request: read the "timings" object in the response
# (prompt_ms = encode+prefill, predicted_per_second = generation speed)
Tool calling: send an OpenAI tools array; expect finish_reason: "tool_calls". Vision: send an image_url content part (data URI works). Both verified on both 27B families.