| 1 | # Snow Globe local AI |
| 2 | |
| 3 | The first model is Qwen3.8-27B, using Unsloth's UD-Q4_K_M GGUF and F16 vision |
| 4 | projector. Both files and the llama.cpp CUDA image are pinned. The model runs |
| 5 | entirely on the RTX 3090, with one inference slot and a 131,072-token context. |
| 6 | Codex compacts at 96,000 tokens to leave room for tool results and output. |
| 7 | The KV cache uses Q8; the single slot leaves more VRAM for vision and runtime |
| 8 | working memory. Additional requests queue so simultaneous generation cannot |
| 9 | reduce the active request's decode rate. An 8 GiB host-memory prompt cache |
| 10 | preserves evicted contexts between coding, reviewer, and helper requests. |
| 11 | The container has a 16 GiB RAM limit for the runtime and checkpoints. |
| 12 | Medium reasoning is enabled for coding, using Qwen's recommended thinking-mode |
| 13 | sampling parameters. This adds |
| 14 | reasoning tokens before answers; the earlier latency baseline disabled thinking. |
| 15 | |
| 16 | Models live in `/srv/clover/Media/AI/LLM/Qwen3.8-27B`, also visible on the Mac as |
| 17 | `/Volumes/clover/Media/AI/LLM/Qwen3.8-27B`. Download once on Zenith: |
| 18 | |
| 19 | ```sh |
| 20 | python3 service/local-ai/download.py |
| 21 | ``` |
| 22 | |
| 23 | The downloader resumes partial transfers, verifies size and SHA-256, and never |
| 24 | replaces an existing file that fails verification. No model download or GPU |
| 25 | service runs in rehearsal VMs. `STUDIO_LOCAL_AI=true` enables the physical host. |
| 26 | |
| 27 | ## API and deployment |
| 28 | |
| 29 | The authenticated API exposes `/v1/responses`, `/v1/messages`, |
| 30 | `/v1/chat/completions`, and `/v1/models`. `/health` is public for deployment |
| 31 | checks. Use `Authorization: Bearer <key>` or Anthropic's `x-api-key` header. |
| 32 | `api_key` is a generated Nomad service secret; don't put it in Git or shell |
| 33 | arguments. |
| 34 | |
| 35 | The production endpoint is `https://ai.paperclover.net`. The |
| 36 | `STUDIO_LOCAL_AI=true` host environment in `nixos/zenith.nix` enables it. |
| 37 | The 3090 has room for one model instance. Stop any GPU preview before deploying |
| 38 | production; download the wrappers again when their endpoint changes. |
| 39 | |
| 40 | `gateway.py` uses Python's standard library and starts llama.cpp on a container |
| 41 | loopback port. Anthropic system-text blocks are combined into one system message; |
| 42 | Chat Completions pass through unchanged. For |
| 43 | Responses, it collects developer/system instructions at the start of the |
| 44 | prompt, maps tool namespaces to flat names, restores returned namespace |
| 45 | identities, and represents raw custom tools as JSON functions. The raw tool |
| 46 | grammar remains in the tool description. Built-in hosted tools such as OpenAI |
| 47 | web search aren't implemented; unsupported tool types return an explicit error. |
| 48 | Codex's hosted web search is disabled in the wrapper. |
| 49 | |
| 50 | ## Coding clients |
| 51 | |
| 52 | Download `snow-codex` or `snow-claude` from the dashboard's **AI / MCP → AI** |
| 53 | tab. They are standalone Bash wrappers around the standard installed Codex and |
| 54 | Claude Code clients; client updates require no wrapper rebuild. Make each |
| 55 | download executable and place it on your PATH. The first run asks for the API |
| 56 | key from the same tab and saves it with mode 0600 under |
| 57 | `${XDG_CONFIG_HOME:-~/.config}/snowglobe-ai/api-key`. No checkout or patched |
| 58 | client is required. The downloads contain endpoint and model settings, never |
| 59 | the API key. The `ai` group opens AI / MCP and grants API-key access and downloads; |
| 60 | infrastructure administrators retain access. Role previews can view the page |
| 61 | and downloads, but cannot access the API key. |
| 62 | |
| 63 | `dashboard/ai/launch.sh` owns the wrapper template. The download endpoint |
| 64 | incorporates the live service hostname and canonical model catalog. |
| 65 | Codex starts with `--approve-for-me`, retaining its workspace sandbox and |
| 66 | automatic approval reviewer; Claude starts with `--permission-mode auto`, |
| 67 | running classification against the same local endpoint. Both use medium |
| 68 | reasoning. Claude's credential store is isolated from ordinary Claude login, |
| 69 | so using Snow Globe does not log the user's subscription out. |
| 70 | |
| 71 | API and event-stream waits allow eight hours. Claude's byte-stream watchdog |
| 72 | has a 30-minute maximum, kept alive by ten-second server SSE pings. Stock |
| 73 | Codex 0.160.0 retains a hardcoded 90-second approval-review deadline that |
| 74 | provider settings cannot extend. The earlier patched build survived a |
| 75 | 95-second forced review delay, but the shipped wrappers intentionally use |
| 76 | standard clients as requested. Queueing can therefore still cause a stock |
| 77 | Codex approval review to time out. |
| 78 | |
| 79 | Override the endpoint or key with `SNOWGLOBE_AI_URL` or `SNOWGLOBE_AI_KEY`. |
| 80 | Explicit client arguments can override wrapper defaults. Existing MCP servers, |
| 81 | hooks, and ordinary commands keep their normal network access. Model inference |
| 82 | and approval classification use Snow Globe; cloud plugins, Apps, automatic |
| 83 | memories, and nonessential client telemetry are disabled for this profile. |
| 84 | |
| 85 | ## Verification |
| 86 | |
| 87 | ```sh |
| 88 | python3 -m unittest discover -s service/local-ai -p 'test_*.py' -v |
| 89 | python3 tools/verify-local-ai.py --url https://YOUR-ENDPOINT |
| 90 | ``` |
| 91 | |
| 92 | The downloadable wrappers were verified against production on 2026-10-05 |
| 93 | with stock Codex 0.160.0 and Claude Code 2.1.289. Both clients fixed the median |
| 94 | fixture and passed all eight unchanged tests, taking 189.3 and 209.4 seconds |
| 95 | while sharing the single slot. Stock Codex's actual automatic approval review |
| 96 | authorized an outside-workspace fixture write in 37.5 seconds. The temporary |
| 97 | file was removed. Download tests also exercised first-run key entry through a |
| 98 | terminal, mode 0600 storage, argument preservation, auth isolation, and catalog |
| 99 | cleanup. Unsigned, cross-origin, and view-as requests were denied; the returned |
| 100 | key authenticated successfully with the live model. |
| 101 | |
| 102 | The second command launches both actual clients concurrently in disposable |
| 103 | fixtures. Each must fix two median bugs, preserve input, pass four tests, and |
| 104 | leave the tests unchanged. It also exercises an actual Codex approval review and removes its temporary |
| 105 | file. JSONL evidence is saved in the printed temporary directory. An optional |
| 106 | `--approval-delay 95` probe demonstrates the stock reviewer deadline; it is |
| 107 | expected to fail with stock Codex. |
| 108 | |
| 109 | The deployed 128K configuration was checked on 2026-10-05 with a synthetic |
| 110 | 122,265-token request. It recovered facts from the beginning, middle, and end |
| 111 | through a namespaced tool call. A 122,384-token continuation consumed the tool |
| 112 | result correctly and reused 122,332 cached tokens. A subsequent 122,409-token |
| 113 | request reused 122,359 tokens and generated 384 tokens at 24.21 tokens/second, |
| 114 | with a maximum streamed token gap of 45 ms. Thinking was disabled only for |
| 115 | these retrieval and throughput probes. The server measured 144.4 seconds of |
| 116 | cold prefill at 846.8 tokens/second. The cached stream's first token took 26.1 |
| 117 | seconds while coding clients shared the slot; that elapsed time includes |
| 118 | queueing. The allocation used 20.87 GiB VRAM and peaked at 9.79 GiB host RAM, |
| 119 | with zero OOM events, cgroup limit hits, or restarts throughout verification. |
| 120 | |
| 121 | Both actual clients then passed the eight unchanged coding fixture tests with |
| 122 | medium reasoning enabled. Codex took 285.4 seconds and Claude 343.5 seconds |
| 123 | while competing with the context probe. This checks real harness compatibility; |
| 124 | the simple retrieval probe does not establish coding quality at 122K tokens. |
| 125 | The actual approval-review check passed in 131.0 seconds, including a deliberate |
| 126 | 95-second hold beyond stock Codex's deadline. Its authorized write and read-back |
| 127 | succeeded, and the temporary file was removed. |
| 128 | |
| 129 | On 2026-10-05, a synthetic 60,756-token Responses request recovered distinct |
| 130 | facts from the beginning, middle, and end through a namespaced tool call. A |
| 131 | 60,875-token continuation consumed the tool result correctly, with 60,823 |
| 132 | cached tokens. A subsequent 60,900-token request generated 384 tokens at |
| 133 | 30.75 tokens/second; the largest streamed token gap was 44 ms. These checks |
| 134 | used the earlier non-thinking preset. They validate context handling and one |
| 135 | simple retrieval task, not general coding quality at that length. |
| 136 | |
| 137 | Cold processing of roughly 60K input tokens took about a minute. End-to-end |
| 138 | latency also included waits behind other active requests. The live Codex |
| 139 | session separately reached about 48K input tokens at 33–34 generated |
| 140 | tokens/second, with recent-prefix prefill below one second. Its compaction |
| 141 | requests completed successfully. The earlier 64K limit and 48K compaction threshold |
| 142 | kept space for tool results, reasoning, and output without enlarging the cold |
| 143 | prefill workload further. |
| 144 | |
| 145 | The cache is bounded RAM storage, with no timed expiry or persistence across |
| 146 | server restarts. In an earlier cross-context probe, 15,912 of 15,916 tokens were |
| 147 | restored after two competing contexts, returning in 1.05 seconds versus 23.99 |
| 148 | seconds without the host cache. Distinct large contexts can still evict one |
| 149 | another. The Anthropic endpoint reports `cache_read_input_tokens` and remaining |
| 150 | uncached `input_tokens`; Responses reports `input_tokens_details.cached_tokens`. |
| 151 | There are no API cache charges. |
| 152 | |
| 153 | The previous two-slot allocation used about 20.9 GiB VRAM; one 64K slot used |
| 154 | about 18.4 GiB. [Qwen's model card](https://huggingface.co/Qwen/Qwen3.8-27B) describes |
| 155 | medium effort as the accuracy/speed balance and cautions that reduced reasoning |
| 156 | can increase total agent time through retries. |
| 157 | |
| 158 | ## Ajax investigation |
| 159 | |
| 160 | As of 2026-10-05, [PewDiePie's official Ajax page](https://data.pewdiepie.com/) |
| 161 | says the model will release when ready. It describes an ablated Qwen3.5-9B |
| 162 | fine-tune for [Odysseus](https://github.com/odysseus-dev/odysseus), but does not |
| 163 | provide verified model weights or a model license to evaluate yet. |
| 164 | |
| 165 | When official weights exist, inspect the model card, license, file formats, |
| 166 | revision, and hashes before downloading. Use the known llama.cpp runtime with |
| 167 | GGUF data; don't execute model-repository code or load pickle-based weights. |
| 168 | Trial inference should have a read-only filesystem and model mount, no network, |
| 169 | no capabilities, no host credentials or private data, process/memory limits, |
| 170 | and only disposable scratch storage. Tool-use trials belong in a disposable VM |
| 171 | with synthetic files and credentials, separate from the personal coding profile. |
| 172 | Evaluate coding, tool arguments, prompt injection, unauthorized writes, and |
| 173 | attempted outbound requests against the base model. Passing those tests is |
| 174 | evidence about observed behavior, not proof that an unrestricted agent is safe. |