1# Snow Globe local AI
2
3The first model is Qwen3.8-27B, using Unsloth's UD-Q4_K_M GGUF and F16 vision
4projector. Both files and the llama.cpp CUDA image are pinned. The model runs
5entirely on the RTX 3090, with one inference slot and a 131,072-token context.
6Codex compacts at 96,000 tokens to leave room for tool results and output.
7The KV cache uses Q8; the single slot leaves more VRAM for vision and runtime
8working memory. Additional requests queue so simultaneous generation cannot
9reduce the active request's decode rate. An 8 GiB host-memory prompt cache
10preserves evicted contexts between coding, reviewer, and helper requests.
11The container has a 16 GiB RAM limit for the runtime and checkpoints.
12Medium reasoning is enabled for coding, using Qwen's recommended thinking-mode
13sampling parameters. This adds
14reasoning tokens before answers; the earlier latency baseline disabled thinking.
15
16Models live in `/srv/clover/Media/AI/LLM/Qwen3.8-27B`, also visible on the Mac as
17`/Volumes/clover/Media/AI/LLM/Qwen3.8-27B`. Download once on Zenith:
18
19```sh
20python3 service/local-ai/download.py
21```
22
23The downloader resumes partial transfers, verifies size and SHA-256, and never
24replaces an existing file that fails verification. No model download or GPU
25service runs in rehearsal VMs. `STUDIO_LOCAL_AI=true` enables the physical host.
26
27## API and deployment
28
29The authenticated API exposes `/v1/responses`, `/v1/messages`,
30`/v1/chat/completions`, and `/v1/models`. `/health` is public for deployment
31checks. Use `Authorization: Bearer <key>` or Anthropic's `x-api-key` header.
32`api_key` is a generated Nomad service secret; don't put it in Git or shell
33arguments.
34
35The production endpoint is `https://ai.paperclover.net`. The
36`STUDIO_LOCAL_AI=true` host environment in `nixos/zenith.nix` enables it.
37The 3090 has room for one model instance. Stop any GPU preview before deploying
38production; download the wrappers again when their endpoint changes.
39
40`gateway.py` uses Python's standard library and starts llama.cpp on a container
41loopback port. Anthropic system-text blocks are combined into one system message;
42Chat Completions pass through unchanged. For
43Responses, it collects developer/system instructions at the start of the
44prompt, maps tool namespaces to flat names, restores returned namespace
45identities, and represents raw custom tools as JSON functions. The raw tool
46grammar remains in the tool description. Built-in hosted tools such as OpenAI
47web search aren't implemented; unsupported tool types return an explicit error.
48Codex's hosted web search is disabled in the wrapper.
49
50## Coding clients
51
52Download `snow-codex` or `snow-claude` from the dashboard's **AI / MCP → AI**
53tab. They are standalone Bash wrappers around the standard installed Codex and
54Claude Code clients; client updates require no wrapper rebuild. Make each
55download executable and place it on your PATH. The first run asks for the API
56key from the same tab and saves it with mode 0600 under
57`${XDG_CONFIG_HOME:-~/.config}/snowglobe-ai/api-key`. No checkout or patched
58client is required. The downloads contain endpoint and model settings, never
59the API key. The `ai` group opens AI / MCP and grants API-key access and downloads;
60infrastructure administrators retain access. Role previews can view the page
61and downloads, but cannot access the API key.
62
63`dashboard/ai/launch.sh` owns the wrapper template. The download endpoint
64incorporates the live service hostname and canonical model catalog.
65Codex starts with `--approve-for-me`, retaining its workspace sandbox and
66automatic approval reviewer; Claude starts with `--permission-mode auto`,
67running classification against the same local endpoint. Both use medium
68reasoning. Claude's credential store is isolated from ordinary Claude login,
69so using Snow Globe does not log the user's subscription out.
70
71API and event-stream waits allow eight hours. Claude's byte-stream watchdog
72has a 30-minute maximum, kept alive by ten-second server SSE pings. Stock
73Codex 0.160.0 retains a hardcoded 90-second approval-review deadline that
74provider settings cannot extend. The earlier patched build survived a
7595-second forced review delay, but the shipped wrappers intentionally use
76standard clients as requested. Queueing can therefore still cause a stock
77Codex approval review to time out.
78
79Override the endpoint or key with `SNOWGLOBE_AI_URL` or `SNOWGLOBE_AI_KEY`.
80Explicit client arguments can override wrapper defaults. Existing MCP servers,
81hooks, and ordinary commands keep their normal network access. Model inference
82and approval classification use Snow Globe; cloud plugins, Apps, automatic
83memories, and nonessential client telemetry are disabled for this profile.
84
85## Verification
86
87```sh
88python3 -m unittest discover -s service/local-ai -p 'test_*.py' -v
89python3 tools/verify-local-ai.py --url https://YOUR-ENDPOINT
90```
91
92The downloadable wrappers were verified against production on 2026-10-05
93with stock Codex 0.160.0 and Claude Code 2.1.289. Both clients fixed the median
94fixture and passed all eight unchanged tests, taking 189.3 and 209.4 seconds
95while sharing the single slot. Stock Codex's actual automatic approval review
96authorized an outside-workspace fixture write in 37.5 seconds. The temporary
97file was removed. Download tests also exercised first-run key entry through a
98terminal, mode 0600 storage, argument preservation, auth isolation, and catalog
99cleanup. Unsigned, cross-origin, and view-as requests were denied; the returned
100key authenticated successfully with the live model.
101
102The second command launches both actual clients concurrently in disposable
103fixtures. Each must fix two median bugs, preserve input, pass four tests, and
104leave the tests unchanged. It also exercises an actual Codex approval review and removes its temporary
105file. JSONL evidence is saved in the printed temporary directory. An optional
106`--approval-delay 95` probe demonstrates the stock reviewer deadline; it is
107expected to fail with stock Codex.
108
109The deployed 128K configuration was checked on 2026-10-05 with a synthetic
110122,265-token request. It recovered facts from the beginning, middle, and end
111through a namespaced tool call. A 122,384-token continuation consumed the tool
112result correctly and reused 122,332 cached tokens. A subsequent 122,409-token
113request reused 122,359 tokens and generated 384 tokens at 24.21 tokens/second,
114with a maximum streamed token gap of 45 ms. Thinking was disabled only for
115these retrieval and throughput probes. The server measured 144.4 seconds of
116cold prefill at 846.8 tokens/second. The cached stream's first token took 26.1
117seconds while coding clients shared the slot; that elapsed time includes
118queueing. The allocation used 20.87 GiB VRAM and peaked at 9.79 GiB host RAM,
119with zero OOM events, cgroup limit hits, or restarts throughout verification.
120
121Both actual clients then passed the eight unchanged coding fixture tests with
122medium reasoning enabled. Codex took 285.4 seconds and Claude 343.5 seconds
123while competing with the context probe. This checks real harness compatibility;
124the simple retrieval probe does not establish coding quality at 122K tokens.
125The actual approval-review check passed in 131.0 seconds, including a deliberate
12695-second hold beyond stock Codex's deadline. Its authorized write and read-back
127succeeded, and the temporary file was removed.
128
129On 2026-10-05, a synthetic 60,756-token Responses request recovered distinct
130facts from the beginning, middle, and end through a namespaced tool call. A
13160,875-token continuation consumed the tool result correctly, with 60,823
132cached tokens. A subsequent 60,900-token request generated 384 tokens at
13330.75 tokens/second; the largest streamed token gap was 44 ms. These checks
134used the earlier non-thinking preset. They validate context handling and one
135simple retrieval task, not general coding quality at that length.
136
137Cold processing of roughly 60K input tokens took about a minute. End-to-end
138latency also included waits behind other active requests. The live Codex
139session separately reached about 48K input tokens at 33–34 generated
140tokens/second, with recent-prefix prefill below one second. Its compaction
141requests completed successfully. The earlier 64K limit and 48K compaction threshold
142kept space for tool results, reasoning, and output without enlarging the cold
143prefill workload further.
144
145The cache is bounded RAM storage, with no timed expiry or persistence across
146server restarts. In an earlier cross-context probe, 15,912 of 15,916 tokens were
147restored after two competing contexts, returning in 1.05 seconds versus 23.99
148seconds without the host cache. Distinct large contexts can still evict one
149another. The Anthropic endpoint reports `cache_read_input_tokens` and remaining
150uncached `input_tokens`; Responses reports `input_tokens_details.cached_tokens`.
151There are no API cache charges.
152
153The previous two-slot allocation used about 20.9 GiB VRAM; one 64K slot used
154about 18.4 GiB. [Qwen's model card](https://huggingface.co/Qwen/Qwen3.8-27B) describes
155medium effort as the accuracy/speed balance and cautions that reduced reasoning
156can increase total agent time through retries.
157
158## Ajax investigation
159
160As of 2026-10-05, [PewDiePie's official Ajax page](https://data.pewdiepie.com/)
161says the model will release when ready. It describes an ablated Qwen3.5-9B
162fine-tune for [Odysseus](https://github.com/odysseus-dev/odysseus), but does not
163provide verified model weights or a model license to evaluate yet.
164
165When official weights exist, inspect the model card, license, file formats,
166revision, and hashes before downloading. Use the known llama.cpp runtime with
167GGUF data; don't execute model-repository code or load pickle-based weights.
168Trial inference should have a read-only filesystem and model mount, no network,
169no capabilities, no host credentials or private data, process/memory limits,
170and only disposable scratch storage. Tool-use trials belong in a disposable VM
171with synthetic files and credentials, separate from the personal coding profile.
172Evaluate coding, tool arguments, prompt injection, unauthorized writes, and
173attempted outbound requests against the base model. Passing those tests is
174evidence about observed behavior, not proof that an unrestricted agent is safe.