sg_reasoning_effort.md

reasoning_effort: the template default (xhigh) against medium, measured on the RTX 3060 12GB

What the template does (read from the GGUF, the template dump)

tokenizer.chat_template of Ternary-Bonsai-2-27B-PTQ1_0.gguf (8,952 characters) contains, when thinking is enabled (enable_thinking undefined or true):

{%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}
{%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %} ... raise_exception ...
{%- if resolved_reasoning_effort == 'xhigh' %}
    {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}
{%- elif resolved_reasoning_effort == 'low' %}
    {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}

medium sets no instruction. The server's /apply-template confirms it : the default prompt for a one-line user message starts with <|im_start|>system\nReasoning effort is set to xhigh. Please think carefully ...<|im_end|>, the medium prompt has no system message and still opens <think>, enable_thinking: false opens and closes <think> at once.

Every way to change it on this fork's server (common/arg.cpp, tools/server/server-common.cpp, tools/server/server-chat.cpp)

  1. --reasoning-effort LEVEL server flag (env LLAMA_ARG_REASONING_EFFORT), default keeps the template default. Present on the prebuilt prism-b10685 llama-server and on the bonsai2 bundle binary (both --help outputs checked).
  2. --chat-template-kwargs '{"reasoning_effort":"medium"}' server flag.
  3. Per request: the OpenAI field "reasoning_effort": "medium" on /v1/chat/completions (none disables reasoning), or "chat_template_kwargs": {"reasoning_effort": "medium"}; the Responses-style reasoning: {effort} is mapped to the same field.
  4. --chat-template-file with an edited template.

Related: --reasoning-budget N caps thinking tokens server-side, --reasoning-budget-message injects a line before the forced close. Neither was measured here.

Measurement

Original file, the bonsai2 bundle build (8971d7b) llama-server, -ngl 99 -fa on -c 65536 -np 1 -ctk q4_0 -ctv q4_0 --jinja, greedy (temperature 0, top_k 1), /v1/chat/completions, two runs per cell, two client caps (max_tokens 4096, the common client default, and 16384). Thinking tokens = server reasoning_content counted with /tokenize; content tokens likewise; generated = usage.completion_tokens. Complete = SVG contains <svg and </svg>; HTML contains <html and </html>; Python parses with ast and has at least 100 non-empty lines. Runs 1 and 2 were identical in every cell (greedy).

Tasks: SVG = the prompt from professorpalmer's bench/reason_ab.py with the subject he lists first, "Draw a five-tier pagoda as a single self-contained SVG. No markdown, no explanation, inline SVG only. Temperature-0 style: clean geometric shapes, white background."; HTML = a single self-contained to-do list page with localStorage; Python = a key-value store CLI of at least 100 lines with a self-test.

capefforttaskfinishgenerated tokensthinking tokenscontent tokenscompletewall s (run 1, run 2)
4096default (xhigh)svg_pagodalength409640960no, empty107.6, 107.6
4096default (xhigh)html_todolength409640960no, empty107.7, 107.6
4096default (xhigh)python_kvlength409640960no, empty107.7, 107.6
4096mediumsvg_pagodastop22381610625yes57.7, 57.7
4096mediumhtml_todostop1758221733yes45.2, 45.2
4096mediumpython_kvlength409623811713no, cut inside the code107.6, 107.6
16384default (xhigh)svg_pagodalength16384163840no, empty477.2, 476.9
16384default (xhigh)html_todostop1298878395145yes368.0, 367.7
16384default (xhigh)python_kvstop1193695362397yes335.0, 335.2
16384mediumsvg_pagodastop22381610625yes57.7, 57.5
16384mediumhtml_todostop1758221733yes45.2, 45.0
16384mediumpython_kvstop467923812295yes123.5, 123.4

Reading

Decision: medium measurably helps. The serve scripts in serve/ and in the prebuilt bundle get --reasoning-effort medium; the README says what it does and what it costs. The tarball already on Hugging Face still carries the old scripts until it is re-uploaded. --reasoning-budget was not measured and is not set.

下载此文件