Back to the changelog

Development

Syzygy 1.2.0 development cycle: voice dictation, local LLMs, and versioned prompts

Seventy commits in the 1.2.0 development cycle: a from-scratch dictation pipeline (Silero VAD segmentation, multi-engine fallback, retry budgets), in-process mistral.rs local models, gitoxide-backed prompt versioning with a skills registry, and release engineering with macOS notarization and signed Linux AppImages.

Syzygy teamPublished 12 min readUpdated
  • #Syzygy 1.2.0
  • #voice dictation
  • #offline speech to text
  • #local LLM
  • #prompt management
  • #skills registry
On this page

Development-cycle notes, updated October 4, 2026. This article summarizes implementation work, not a confirmed 1.2.0 release date. Available packages and released versions are listed on the download page. The commit counts below describe the period through September 26.

1.1.0 shipped on August 31, 2026. Through September 26, 70 commits landed on main (32 feat, 17 refactor, 9 add, 5 fix), and all three capability blocks of this development cycle grew from zero inside that window: the first dictation code appeared on September 5, the local-LLM crate was born on September 14, and the skills registry on September 23. Add the release engineering that ran through the whole month (macOS notarization and updater signing, signed Linux AppImages) and 1.2.0 has a clear theme: data enters the workbench by speaking as well as copying, AI compute moves from the cloud onto your machine, and the prompts and skills you accumulate get versioned like code.

Voice dictation stage: recording, segmentation, and transcription in one surface
Voice dictation stage: recording, segmentation, and transcription in one surface

Voice dictation: from first commit to full pipeline

One decision made before any code shaped the whole feature (ADR 0001): transcription results materialize as clipboard items and enter Syzygy's existing preview, search, organization, and paste flows. The alternative (refocusing the app you were in and auto-pasting) was rejected. Auto-injection requires cross-platform global hotkeys, focus restoration, and accessibility permissions, and mis-delivered text becomes the user's problem. The workbench paste path already existed, so voice results join a model the user already controls.

The full pipeline:

flowchart LR
  A[Microphone capture] --> B[VAD streaming segmentation]
  B --> C[Per-segment transcription with engine fallback]
  C --> D[Whole-session retranscription]
  D --> E[Punctuation polish and hotword rewrites]
  E --> F[Materialize as clipboard item]

In prose: audio is cut into segments for recognition; failed recognition switches engines under a budget; after you stop, the whole session gets one more pass with full context; deterministic text cleanup runs last, and the result lands in the clipboard.

Segmentation. Long recordings must decide where to cut before you stop. This release makes Silero VAD the default segmenter. The model is locked in the catalog: source, version, and SHA-256 are immutable, and dropping an arbitrary onnx file into a path never activates anything. Inference runs on the sherpa-onnx blocking thread with window and sample-rate semantics aligned to Silero's training setup; each cut trims a tail before the next segment starts. Energy-based detection was not deleted. It remains the fallback branch when the model is not installed or fails to load, and both segmenters share the same window limits and minimum segment length. When you press stop, the detector is flushed so pending speech is not dropped. An honest boundary: the strategy and interfaces have regression coverage, but validation with real microphones across Windows and Linux devices is still on the plan.

Fallback chain and the retry budget. Recognition engines come in four kinds, ordered by your configuration: system recognition (macOS Speech), local Whisper, local Paraformer on sherpa-onnx, and remote engines on AI accounts. Remote engines speak three audio protocols: OpenAI-compatible chat completions with audio content parts, multipart transcription endpoints, and DashScope's native multimodal endpoint (qwen-audio-3.0-asr-flash uses the native protocol; pricing and regional availability follow the provider’s current terms). Tencent Cloud ASR is a separate module.

The headline change is that recognition retry moved from "one failure, fall down the chain" to a triple budget: 12 total attempts per segment, 3 per engine before demotion, and a 60-second wall-clock budget. The budget only gates the start of a new attempt: a request already in flight runs to its own conclusion, with no hard timeout amputating in-flight work. Failure classification decides where to go next: transient failures (timeouts, transport errors, 5xx) retry the same engine with backoff starting at 500 ms, stepping 500 ms, capping at 1.5 s; an unavailable engine and a deterministic "no speech detected" both move straight to the next engine, because re-asking a deterministic verdict is waste. The old code let max_attempts silently truncate your configured engine list during planning. That semantic is gone. Remote request and response shapes are pinned by wiremock contract tests, credentials live in the system keychain, and speech profiles keep only vendor non-sensitive fields.

Whole-session retranscription. Segmentation buys immediacy at the cost of cross-segment context. So after you stop, before post-processing, the runtime concatenates the session's own segment recordings (at least two segments, 300 seconds total or less, which also lines up with the 10 MB request limit on remote engines) and hands them back to the session's winning engine for one full-context batch pass. The result replaces the concatenated segment text. Retranscription is an upgrade, not a gate: if the engine refuses, concatenation fails, or you cancel, the run falls back to the concatenated text and post-processing proceeds, so words are never lost. Session audio stays in memory only and is cleared when the run reaches a terminal state.

Finishing never blocks the next recording. Previously the record button stayed grey while AI post-processing ran. Now the microphone slot is released the moment Stop is accepted; finishing steps (post-processing, retranscription, materialization) complete in the background, attributed to their run, and you can start the next dictation immediately.

Post-processing. Text cleanup has two layers. The first is deterministic punctuation polish. CJK punctuation completion and repeated-punctuation collapse are plain rules, not model calls: linear cost, reproducible results. The second is the hotword dictionary. Recognizers do not return terms the way you typed them; a Chinese engine tokenizes OpenAI into open ai and GPT into G P T. Before this development cycle the dictionary compensated with 12 hand-written phrase rewrites, one per spotted term. The matching rule now strips inline whitespace before matching and maps the hit back onto the original text, which covers all 1,153 enabled entries at once, so those 12 rewrites were deleted. The whitespace span is bounded (no wider than twice the term it interrupts) so distant characters cannot fuse into a term, and pinyin candidates bypass this path entirely.

The free tier's 60 seconds. The free tier caps one continuous recording at 60 seconds, and the implementation was designed rather than defaulted: hitting the limit pauses the take (microphone off, segment untranscribed), resume reopens the microphone on the same take, and stopping transcribes everything as one segment. Time is measured in audio samples, not wall clock: the runtime holds no timer, and samples delivered between the limit firing and the microphone actually closing are kept. Pro has no such cap.

Smaller, real changes round out the release: copying path text (say, an editor's Copy Path) upgrades qualifying plain-text paths to file entries at capture time, with the original text kept as provenance so pasting stays faithful to what was copied; untagging an item now keeps the tag itself, with orphaned tag links cleaned up by migration; translation profiles replaced their priority field with an explicit fallback order; OCR model handling and resource synchronization were refactored, and the model registry gained translation support; AI chat gained usage accounting, and the chat chain introduced deterministic nonce encryption.

Smaller, real fixes round out the voice line: system recognition no longer reports "no speech detected" when your spoken language differs from the system locale; voice run history is bounded so long-term use cannot grow it without limit; the Quick Panel gained the same voice entry as the main window; and the dictation stage was rebalanced so the speaking flow owns the column.

Local LLMs: in-process inference on a loopback endpoint

The goal for local models is a single sentence: clipboard understanding, translation, and chat must work fully offline with data never leaving the machine.

The engine is mistral.rs 0.9.3, in-process, loading GGUF weights directly, with Metal enabled on macOS builds. The model serves an OpenAI-compatible endpoint bound to 127.0.0.1 only: POST /v1/chat/completions supports buffered and SSE responses, and GET /v1/models lists the model. Protocol adaptation is therefore decoupled from the inference engine, and every AI consumer in the app calls through one boundary. Streaming output passes through an incremental filter that strips <think> reasoning blocks as they arrive, so reasoning traces never reach the UI. One local model serves at a time, and the app restores the last intent on restart.

Download admission is stricter than "the file exists": downloads verify length and SHA-256 in a staging file before publishing into the model directory; every actual load independently re-verifies; a file or parent directory that is a symlink is refused outright. A tampered file of identical length does not pass, and a previous successful check is not a license for this load.

Where do the models come from? Four catalogs (chat, ASR, embedding, OCR) are generated reproducibly from upstream snapshots: the sync script locks upstream revisions and SDK versions, every artifact declares role, path, format, size, and SHA-256, a policy decides which entries enter the product catalog, and --check recomputes offline and byte-compares to catch drift. The runtime bindings are a closed set: mistral.rs GGUF, fastembed ONNX, whisper.cpp, sherpa-onnx Paraformer, and PaddleOCR MNN. Catalog entries without a corresponding engine binding are never claimed installable, and API-only models (such as DashScope's ASR) are never dressed up as downloadable weights.

Prompt library: templates, tags, and version history in one workbench
Prompt library: templates, tags, and version history in one workbench

Retrieval is the other main thread. Semantic vector search fused with lexical search via reciprocal rank fusion (k=60): each ranking is computed independently, merged deterministically, and ties break by document ID, so identical input always yields identical output. Vectors are a derived cache: a text fingerprint plus model identity decide validity, and switching models clears other models' vectors rather than mixing vector spaces. When the indexing queue is idle, bounded backfill runs at most 64 items per batch, at least 15 seconds apart. Local and remote embeddings share one port: remote calls go through a configured account's /embeddings with index, count, dimension, and finite-value checks on the response. The remote catalog gained an OpenRouter sync script, and usage accounting lines up with it, including per-model pricing.

Search qualifiers: with semantic and lexical search fused, you still filter by type, tag, and device
Search qualifiers: with semantic and lexical search fused, you still filter by type, tag, and device

One more gate sits before any request goes out. Model capability declarations are tri-state: undeclared, explicitly empty, and explicitly listed mean different things. When a declaration clearly lacks a required input modality, the request is refused before it reaches the provider, with the reason reported; attachments are never silently OCR'd and never quietly dropped. When the declaration is unknown, the attempt is allowed, and a successful attempt never rewrites the declaration.

Prompts and skills: text assets managed like code

Prompts gained real version control this development cycle. The app embeds gitoxide (gix, a pure-Rust git implementation); saving updates the working draft, while explicit publication creates a version and git commit. History and diff commands inspect publications; restore writes the selected revision back as working content for review before another publication. The variable system resolves through five tiers, falling back by scope, with security tests asserting templates cannot leak secrets into rendered output. Prompt templates got a tag system that matches how clipboard entries are organized.

Attachments let prompts reference images or files: attachments have their own identity, storage, and reference syntax; request assembly follows the body's render order exactly; and a missing referenced attachment fails the request rather than being skipped. Before sending, the references run through the tri-state modality precheck: a model declared with vision gets images, and no one else does.

The skills registry is a new domain. A skill source is an ordered, user-managed list (add, toggle, remove, reorder); scanning deduplicates by the SHA-256 of the raw SKILL.md bytes, keeping the first valid instance and marking duplicates. Removing a source never deletes files on disk, and an external directory does not become app-reclaimable storage just because it was indexed. The editor's slash menu filters skills by source qualifier, and prompts, skills, and memory keep separate storage ownership while sharing only the file substrate and editor components.

Release engineering: making the update channel trustworthy

Shipping itself consumed 14 commits this cycle. On macOS we built a repeatable notarization pipeline: a notarization script handles submission and polling of Apple's state machine, a signing script signs the Tauri updater archive without exposing private key material, and a separate script projects the updater's public key into the desktop build config: public key in the build, private key never on the build machine. Along the way we fixed a concrete bug: codesign needs --xml to emit plist output.

On Linux, AppImages build in containers with a dry-run mode for the scripts, signing happens after repacking, and the private key path is validated before use. DMG creation moved to zlib compression to control size. iOS store assets were filled in during the same cycle: App Store screenshots for login, home, and the main feature surfaces. The bundled voice dictionary regeneration also joined the release pipeline: it reproduces the dictionary offline from reviewed CSpell data, so dictionary content is traceable and rebuildable.

The test surface grew in step: 4 of the 70 commits are pure end-to-end test commits covering the skills workbench, slash skill panel, privacy app picker, voice dictation UI, layout regressions, workspace routing, AI model declarations, image previews, and prompt attachments. The voice line alone accumulated 193 passing unit and integration tests by mid-September (83 in the domain crate, 110 in the runtime), and the skills-registry week carried 1,066 passing tests across its 10 touched packages. Localization key checks stayed at 100% sync.

Known boundaries and next steps

Stated plainly, what this development cycle did not finish: cross-device validation of VAD segmentation and the recognition chain with real microphones is pending; streaming partial results for retranscription are not on screen yet (the whole-session replacement lands in one step); error classification has not descended to 429 and Retry-After granularity; the current runtime now wires Qwen Realtime and DashScope Message, StepFun, iFlytek, Volcano Engine, and Tencent Cloud. Runtime integration is not a claim that every provider account or real microphone has passed acceptance testing.

Where this goes next is covered in the roadmap: app-store distribution on mobile, exposing the data space and memory as an MCP server, and deploying encrypted relays. The previous release is in the 1.1.0 notes; voice usage is documented in the voice input guide, and the local models and prompts and skills articles cover their respective workbenches.

Frequently asked questions

Does Syzygy voice input require an internet connection?

No. The local catalog carries downloadable Paraformer and quantized Whisper models; recording, segmentation, and transcription all run on your device. Audio only leaves the machine when you explicitly pick a remote engine and consent to it.

What are the free-tier voice limits?

Continuous recording pauses automatically at 60 seconds on the free tier. Recorded audio is kept, you can resume and keep speaking, and everything transcribes together when you stop. It is a pacing limit and the counter restarts with each new recording.

How was the 1.2.0 version number chosen?

Continuous builds have rolled since 1.1.0 shipped on August 31; the repository version field updates with the release process. This entry organizes the cycle's changes under 1.2.0 — the published package number is whatever the download page shows.