官方AGENTS调参指南.md

Agent guide — tuning the Bonsai demo

For AI agents (and humans) helping someone set up this demo. Goal: pick the right flags for the user's hardware and use case. The behavior notes below come from real testing; measure on the user's own hardware before promising performance (the timings object in every API response has the numbers).

Why the 27B models (what to show off)

The 27B generation is a step change over the earlier 8B/4B/1.7B demos:

The models

BONSAI_FAMILY / BONSAI_MODELWeightsNotes
bonsai2 / 27B (default)GGUF 5.9-7.2 GB + mmproj 0.63 GB; MLX 2-bit 8.6 GBThe current generation and what plain ./setup.sh installs. Two GGUF bands, both requiring this demo's binaries: PTQ1_0 (1.75 bpw, 5.9 GB, densely packed trits, smallest) and PQ2_0 (2.13 bpw, 7.2 GB, faster prompt processing, what the scripts download). There is no mainline-compatible band; see the Q2_0 warning below. The MLX pack carries its own Hadamard-aware loader in runtime/ plus the vision tower, and runs on stock mlx-vlm from .venv-vlm, not the fork in .venv
ternary / 27B (previous generation)GGUF ~6.7-7.1 GB + mmproj 0.9 GB; MLX 2-bit ~7.9 GBSuperseded by bonsai2. Two GGUF formats since the mainline rebase (prism-b10658+): PQ2_0 (group 128, 6.66 GiB, smallest/fastest where supported: CUDA, Metal, CPU, ROCm) and official Q2_0 group 64 (Ternary-Bonsai-27B-Q2_g64.gguf, 7.05 GiB, adds Vulkan/SYCL); smaller sizes use *-Q2_0_g64.gguf naming. The scripts pick per backend. Legacy *-Q2_0.gguf files (no g64) only load on old prism-v5 releases; new binaries refuse them with an error
bonsai / 27BGGUF Q1_0 ~3.5 GB + mmproj 0.9 GB; MLX 1-bit ~4.8 GBSmallest and fastest; fits on a modern iPhone without offloading
8B / 4B / 1.7B (both families)smallerText-only, no tools wiring, legacy tested flag set

Do not run Bonsai 2 on stock llama.cpp

Every Bonsai 2 band stores its weights in a rotated basis and needs the activation transform that only this demo's binaries have. PQ2_0 and PTQ1_0 are rejected outright by mainline as unknown types, which is the safe failure. Q2_0 is the dangerous one: mainline loads it without a warning and produces gibberish, because Q2_0 is a type upstream already knows and qwen35 is a supported architecture, so nothing in the file makes it refuse.

For that reason the Bonsai 2 Q2_0 band is not in the model repo. It lives on its own at Ternary-Bonsai-2-27B-gguf-dev, named Ternary-Bonsai-2-27B-Q2_0-prism-fork-required.gguf, published for testing and for the work to upstream the Hadamard changes. Do not point a user at it for running the model. It moves into the main repo once mainline can run it.

If someone reports gibberish from Bonsai 2, check which binary they used before anything else.

All three 27B families have the same capabilities (vision, tools, thinking, long context) — they differ in size and speed. Context: 262,144 tokens max; FP16 KV cache is 64 KiB/token (~6.3 GiB at 100K), so 100K context fits on many consumer devices — full peak-memory table in the README's Context Size section. All 27B repos: https://huggingface.co/collections/prism-ml/bonsai-27b. All model repos are public; no token needed.

Knobs that matter (27B)

All extra args pass straight through the start scripts, e.g. ./scripts/start_llama_server.sh --image-max-tokens 1024 -ub 1024.

KnobWhat it doesTrade-off
BONSAI_SPECULATIVE=1 (env, start_llama_server.sh only)Experimental. Loads the paired DSpark drafter for speculative decoding (--spec-type draft-dspark). Post-migration (prism-b10658+) the drafter is the converted *dspark-dflash* sidecar (~0.6 GiB; embedding/lm_head shared with the target — one-time conversion documented in SPECULATIVE.md). Measured on an L40S: ternary 27B 1.8-2.4x decode (2.06x blended), 1-bit 27B 1.4-1.75x (1.60x blended). Off by default.The drafter path is stable and fast on CUDA. On Apple Silicon (Metal) it only pays off for ternary code/math (~1.2x) and is a net slowdown on chat/reasoning and for the 1-bit family, so do not recommend it on Macs. Disables cross-request prompt-cache reuse (every turn re-prefills) and forces single-slot (-np 1); worse for multi-turn and the agentic Open WebUI path, which is why it is server-only and opt-in. The prebuilt binaries include both the dspark-capable llama-server and the CLI one-shot llama-speculative-simple. Details: SPECULATIVE.md.
BONSAI_KV4=1 (env, start_llama_server.sh only)Experimental. Q4_0 (4-bit) KV cache, ~3.5x smaller KV memory. Optional quality booster: ./scripts/make_kv_bias.sh builds a model-specific mean-centering bias (tiny calibration corpus is enough; users can pass their own text) that the server picks up automatically.Memory tool, not a speed tool: decode is slightly slower than F16 KV. The 27B's hybrid attention already keeps KV small, so only reach for this at very long contexts on tight machines. The bias is calibrated with K-rotation off and the script/server handle the matching flags automatically; llama.cpp backend only. Details: KV-CACHE.md.
--reasoning-budget NCaps thinking at N tokens (default -1 = unlimited)Middle ground; pair with --reasoning-budget-message
--image-max-tokens NDownscales images to ~N vision tokens (1 token ~ 32x32 px). Model allows ~4096 (=4.2 MP)The scripts default this to 1024 on Metal / Vulkan / CPU and leave CUDA/ROCm uncapped; override with BONSAI_IMAGE_MAX_TOKENS (0 = uncapped). Loses fine detail (small text / OCR) on large images; images under the cap are unaffected
BONSAI_NGL=N (env)Overrides GPU layer offload (auto-detect defaults to all layers when CUDA/ROCm/Vulkan/Metal tooling is present; 0 = CPU-only)The auto-detect keys on installed tooling, not GPU capability: a CPU box with only an integrated GPU and Vulkan drivers will offload to that iGPU. Capable iGPUs (e.g. Strix Halo) genuinely benefit, weak ones are better on CPU; recommend BONSAI_NGL=0 when decode is slower than expected on iGPU-only machines
-ub NPrefill microbatch (default 512)Sometimes faster prefill on Metal at 1024; measure
BONSAI_CTX=N (env)Overrides the default context. The scripts default to a RAM-tiered size (8192 up to 131072 by machine memory) for predictable memory use; 0 (or unset) resolves to that same tiered size — it is not passed to llama.cpp as -c 0 (which would use the model's full 262k training context and OOM constrained machines). Pass an explicit number to force full contextBigger context = more KV memory at 64 KiB/token FP16; pair huge contexts with BONSAI_KV4=1
BONSAI_MMPROJ_CPU=1 (env, 27B)Adds --no-mmproj-offload so the vision projector stays in system RAM instead of VRAM. Off by defaultFrees ~0.9 GiB of VRAM for KV/context on tight cards; cost is a slower image prefill (CPU vision encode, tens of ms to seconds/image). Token generation and text-only requests are unaffected, so it's a good trade when images are occasional
--parallel NServer slots (default 4)More slots = more concurrent users, same total context pool
BRAVE_API_KEY (env, Open WebUI)Makes a Brave Search MCP server available next to the preconfigured Hugging Face + DeepWiki onesNeeds npm i -g @brave/brave-search-mcp-server; skipped otherwise. Limited to web/news/summarizer (~2.9k tokens); its full tool set is ~29k (brave_place_search alone ~20k). Override with BONSAI_BRAVE_TOOLS.
BONSAI_CODE_INTERPRETER=0 (env, Open WebUI)Disables the server-side Jupyter code interpreter (plots via matplotlib, data via pandas/numpy, market data via yfinance). On by default.Off falls back to browser Pyodide: plots still work, but no yfinance/network. The Jupyter stack (.venv-jupyter) is built by setup.sh.

Thinking is extracted into reasoning_content by default (collapsible in UIs).

Built-in web UI (llama-server) — what it can do without any setup

Everything below is per-conversation, in Settings, no server restart:

Note: UI settings live in the browser (localStorage), so they are per-user/per-machine — server-side pre-seeding of tools is only available in the Open WebUI demo. This also means browser state outlives config changes: if every NEW chat prefills thousands of tokens, the user toggled an MCP server on from the new-chat screen (that saves it as their default for future chats) - toggle it off there, or clear site data for localhost:8080; an incognito window shows what a fresh browser would get.

Agent: tell the user about the image-token cap. If they are on Metal / Vulkan / CPU, say that large images are downscaled to ~1024 vision tokens by default (fast, but fine detail in big images is lost) and ask which they prefer:

Adding MCP servers

Full guide with entry examples: TOOLS.md (repo root). The essentials:

Optional web search (Brave) - needs a Brave Search API key; the key stays local (never committed). Install the bridge once: npm i -g @brave/brave-search-mcp-server.

Behavior notes (from testing on Apple Silicon)

Quick verification commands

# server capabilities and effective context/slots
curl -s http://localhost:8080/props | python3 -m json.tool | head -30

# timing any request: read the "timings" object in the response
# (prompt_ms = encode+prefill, predicted_per_second = generation speed)

Tool calling: send an OpenAI tools array; expect finish_reason: "tool_calls". Vision: send an image_url content part (data URI works). Both verified on both 27B families.

下载此文件