Syzygy 1.1.0 shipped on August 31, 2026. In the 70 commits since, one thread has been about pulling inference into the clipboard app's own process: an embedded mistral.rs 0.9.3 loading GGUF, Metal enabled on macOS builds, and an OpenAI-compatible endpoint bound to 127.0.0.1:0. The AI panel, hybrid search, and usage page in 1.2.0 all stand on that base.
What supports it is a set of four design decisions: the model catalogs are generated from upstream snapshots, model capabilities are declared with three states, file verification runs at both download and load time, and search merges lexical and vector rankings with reciprocal rank fusion. Each decision answers a specific way this can fail. Here they are, one by one.
A hand-maintained model list always rots
The app ships with a set of downloadable models: 4 chat, 20 ASR, 9 embedding, and 3 OCR entries, 36 in total. An entry is not a filename plus a link. It is an artifact contract: for every file, the path, format, byte size, SHA-256, and download sources, plus license, language coverage, and capability declarations. The Qwen2.5-0.5B-Instruct (Q4_K_M) entry in the chat catalog, for example, is a single 491.4 MB GGUF with two sources, Hugging Face and hf-mirror, and the upstream revision pinned inside the URL.
A list like that, written by hand, rots in predictable ways: upstream publishes a new revision while the URL still points at the old one; the file quietly grows by 30 MB while size_bytes stays stale; and the sneakiest failure is a checksum copied with one wrong character, so a user downloads 400 MB, verification fails, and nobody can say whether the file or the list is wrong.
So the catalog is generated. A maintainer explicitly runs --sync; the script fetches metadata from fixed upstreams (Hugging Face Hub, the sherpa-onnx release feed, PaddleX), locks the revision and SDK evidence, validates sources and JSON Schema against an in-memory candidate set, and only then writes the snapshot, the catalog, and the lock file, in that order. Everyday generation and --check are fully offline: recompute the output, compare it byte-for-byte against what was committed, and exit non-zero on drift. The same lock input always produces the same bytes, and every upstream or SDK change becomes a reviewable diff.
The tier field (candidate, validated, recommended, featured) expresses curation intent, meaning "what we plan to recommend," not "what we have tested." The current acceptance report records inference_smoke_tested: false in plain text. The catalog guarantees that the artifact contract is complete and the engine binding is implemented; whether inference actually works has to be verified on real hardware, separately. We would rather write that sentence in the docs than let a recommendation label underwrite performance.
If you upgraded from 1.1.0, none of this is visible: same download entry, same progress and resume, just a list that can no longer quietly go wrong.
The four catalogs map onto four closed sets of engine bindings: chat through mistral.rs GGUF, embedding through fastembed ONNX, ASR through whisper.cpp and sherpa-onnx Paraformer, OCR through PaddleOCR on MNN. A binding must reference artifacts that exist with consistent roles and formats; engine selection never reads capability tags or file extensions. The rule also settles a question that is easy to ask wrong: DashScope's qwen-audio-3.0-asr-flash is an API-only model with no open weights, so "download and install" does not apply to it. It is usable only through DashScope's audio endpoint on a pay-per-call basis, and the catalog will never fabricate a downloadable artifact contract for such a model.
"Never tested" and "confirmed unsupported" are two different states
The easiest lie in capability declarations is treating an empty field as a verdict about the model. Upstream metadata is often incomplete: a model card that says nothing about input modalities might mean "no image input" or "nobody filled it in." Confuse those two, and pre-flight logic becomes unwritable.
Syzygy's capability contract gives every axis (capabilities, input_modalities, output_modalities, plus a separate tasks list) three states: a missing or null field is Unknown, meaning "we don't know"; an empty array is Declared([]), meaning "upstream explicitly declared there is nothing here"; a non-empty array is Declared(values). Three states instead of two, and the extra branch preserves the difference between missing information and explicit refusal.
Request behavior follows the distinction. Explicitly unsupported inputs are rejected before the request is sent. Unknown inputs may be attempted, but the system never claims support it does not have, and upstream errors come back as they are. A few boundaries complete the picture: accepting image input does not auto-declare OCR; the 20 ASR entries each carry an implemented engine binding; paths without a binding (some vision-language models, say) are not declared runnable because of a file extension or a capability tag.
If you bring your own OpenAI-compatible account, the same contract applies to your provider profile: the pre-flight capability check does not care whether the model is local or remote.
SHA-256 verification is about timing, not filenames
Per-file checksums are not exotic. What matters is when they run. Syzygy's rule: a download lands in a staging file and is checked for length and SHA-256 before being published into data_dir/llm-models/{catalog_id}/; querying "is this installed" verifies once; and every actual load independently verifies again, building the model identity for that run from the measured digest.
The third check is the point. Between "query says installed" and "the engine reads the file," real time passes, and the file can be replaced, corrupted, or pointed elsewhere by a symlink. So the load-time admission check re-validates regular file, path, length, and digest: same-length tampering is caught by the measured digest, a symlinked file or parent directory is rejected outright, and a failed check never triggers a re-download to repair the file. The verification boundary does not fix things.
The limits are written down too: a passing check proves that the bytes match the catalog contract at that moment. It is not a file lock and cannot promise the disk stays unchanged afterwards. A passing check does not prove the GGUF parses; the native loader's result is handled separately. To be blunt: integrity checking defends against transfer corruption and silent loading of tampered files. It does not defend against a process that already has write access to your disk. That threat model belongs to the operating system.
If your worry is "will an interrupted download leave half a model behind," the behavior is: interruptions keep valid progress, and only files that pass full verification ever appear as installed.
Why the endpoint only listens on 127.0.0.1
mistral.rs runs inside the application process. There is no second inference service. The app then exposes an HTTP endpoint implementing POST /v1/chat/completions (buffered and SSE streaming responses, [DONE] on normal completion) and GET /v1/models. Why bother? Because the OpenAI-compatible protocol is already the calling boundary every AI feature in the app speaks, translation, tagging, editing, and chat alike. Putting the local model behind the same endpoint shape separates protocol adaptation from native inference completely: the upper layers do not need to know whether a given completion runs on a remote account or on local Metal.
The bind address is 127.0.0.1:0: loopback plus an ephemeral port. The exposure surface of this choice has two layers.
Layer one is the network. Loopback packets never leave the machine. A sniffer on the café WiFi, a scanner on the LAN, any host behind the same NAT: none of them can reach the port. Compare with binding 0.0.0.0, which would expose an unauthenticated inference endpoint to the entire local network and turn "local-first" into a slogan.
Layer two is the machine itself, and loopback does nothing there. Any local process can connect to a port on 127.0.0.1, and port-scanning localhost is elementary malware behavior. To be honest: the endpoint is not defended against a malicious local process. Syzygy's defenses sit elsewhere. Load-time integrity admission ensures the bytes come from the catalog contract. The ephemeral port means there is no fixed, predictable port number. One model is served at a time, the old service closes before the new one is published, and the managed source cannot be taken over by ordinary account operations.
If you run this on a work laptop, what else is installed on the machine matters more than which network you joined. The app cannot change that; saying it clearly beats leaving it vague.
BM25 scores and cosine scores must not be added
The first consumer of local capability is clipboard search. Lexical search (Tantivy's BM25) is great at exact matches: filenames, code, identifiers. Vector search is great at nearby meaning: "that link about reimbursement from last week" finds an entry whose body is a travel expense form. You want both, and the scores cannot be combined: BM25 is unbounded, cosine similarity lives in [-1, 1], and a weighted sum amounts to declaring the two scales interchangeable, with any weight being a guess.
Syzygy uses reciprocal rank fusion. Each ranking converts positions into contributions of 1/(k+rank) with k fixed at 60; contributions are summed across the two lists and the totals are sorted. The recipe endures because it needs ranks, not scores: a document sitting at position 200 on the lexical list still wins if it sits at position 3 on the vector list. k=60 flattens the weight of a single list's top ranks, letting documents that do well on both sides rise. Ties break by doc_id lexicographic order, so identical input always yields identical output.
The other half of the engineering is vector identity. Vectors are a derived cache, determined by the body projection and the model identity: when the body changes, privacy eligibility changes, or the embedding model changes, old vectors must be invalidated, and attaching a new model deletes every other model's vectors. Derived data can be rebuilt, so delete-then-backfill beats keeping hidden caches around. Backfill is bounded: while the indexing queue is idle, at most 64 items per batch with at least 15 seconds between batches. Slow beats competing with foreground work. When the model is not ready or the query embedding fails, the lexical order is kept as-is, and a missing model does not spam warnings.
Local embeddings run on fastembed 7.1.0's ONNX pipeline, with files verified exactly like chat models (bge-small-zh-v1.5 is 5 files totaling 95.3 MB including tokenizer and config, each with its own digest). Remote embeddings go through a configured account's /embeddings endpoint, with the response's indices, count, dimensions, and value sanity checked. Both paths implement the same search port.

If you have thousands of clipboard entries, expect no full semantic search on night one. Backfill catches up over time, and semantic hits blend into results as it does.
Usage metering: NULL means "not reported," 0 means "reported zero"
Local inference poses a metering question: where do token counts come from? If the app estimates and writes them into the usage table, the table mixes "what providers said" with "what the app guessed," and no number on the page is quotable anymore.
Syzygy's rule is one line: observed values come only from provider reports or engine counts; local estimates serve input budgeting and are never written to the observation table. Fields a provider did not report are stored as NULL, never filled with 0. The difference is real in SQL aggregation: a day with calls but no reported tokens renders as "—", a day with no calls renders as 0. Displaying the former as 0 would be inventing a day of idle burning on the provider's behalf.
The two ledgers stay separate. The observation table stores one row per call, local and remote in the same table, distinguished by judge_kind (local_llm vs remote_api), and local rows never count toward the remote daily budget. Budget-side estimation, including chars/4-style heuristics, gates only inputs; its error never touches observed truth. Cost estimation reads from an OpenRouter catalog snapshot: a sync script pulls openrouter.ai/api/v1/models into a local snapshot, pricing shares the catalog's lifecycle, and unknown or unpriced models return "no estimate" rather than a fabricated 0. The recorder itself is a side-channel: its own failures are swallowed and can never affect the call being measured.

If you run sensitive content through local models, the ledger's promise is: which call, which engine, how many tokens, all stored on your machine. The numbers on the page are not the provider's bill, and the page says so itself.

Where to start
In 1.2.0 everything lives in settings: download models in the AI sources area (4 chat models from the 491 MB Qwen2.5-0.5B to the 2.5 GB Qwen3-4B, with declarations generated from upstream snapshots), click to use (load, publish the managed profile, and rebind the endpoint in one pass), and pick an embedding model in the search area. The install and first-run guide covers setup and macOS permissions; the indexing and search piece goes deeper into search; the prompts and skills workbench covers the asset layer that feeds these models. The full decision records live in ADRs 0031, 0034, and 0029 in the repository.
Limits, stated plainly: the catalog's context_tokens is a declared value, not a measured runtime parameter; inference_smoke_tested is still false in the acceptance report, and per-model inference acceptance is recorded by platform and artifact revision rather than granted by release. Recommendation tiers are our curation judgment. Measured speed and quality data will have to come from acceptance records on real hardware, and that gap should not be papered over with marketing language.
Frequently asked questions
How does Syzygy run local LLMs?
Syzygy 1.2.0 embeds mistral.rs in-process, enables Metal acceleration on macOS, loads GGUF files that pass SHA-256 verification, and serves an OpenAI-compatible chat endpoint on 127.0.0.1 (POST /v1/chat/completions and GET /v1/models). One local model is served at a time; switching closes the old service first.
Why does local search combine keywords and vectors?
Keyword search excels at exact matches and vector search at nearby meaning, and their scores live on incompatible scales, so adding them is meaningless. Syzygy merges the two rankings with reciprocal rank fusion at a fixed k=60, looking only at ranks. When the embedding model is unavailable, results fall back to the lexical order.
Is downloading local models safe?
The model catalogs are generated from upstream snapshots, and every file carries its own size and SHA-256. Downloads are verified in a staging file before publication; install-state queries and every actual load re-verify the file. Same-length tampering, symlinks, and path escapes are rejected, and failed verification never loads.