Skip to main content
CHANGELOG

What shipped

Released versions only, generated at build time from the product Keep a Changelog file. Unreleased work is not listed here: if it is not in a version heading, it has not shipped.

Latest released v3.23.4 · · 61 releases

Currently downloadable v3.25.2-preview.1 · preview channel

Those are two different questions. This page lists versions that have been released, from the product changelog. The download page serves whatever the live preview channel points at, which moves ahead of the last released version with every build. Download the current build from /download. Channel and exact version always come from the live release manifest.

Releases 6–10 of 61

Page 2 of 13

v3.22.0

Fixed

  • Daemon could silently lose custom provider profiles on startup. JsonOrchestrationStore.open() was not idempotent: a second call (as NalaOrchestrationService.start() makes, after the constructor's own lazy open) could race the debounced disk persist (F5.8) and reset already-seeded defaults back to an empty store whenever it landed inside the 50ms write window — effectively always, for synchronous construct+start. open() is now a no-op once loaded; custom-provider CRUD (add/update/remove) flushes synchronously since it's a rare admin action, not one of the hot per-tick writes the debounce targets.

Added

  • A2A Reliability Phase 1 (Grok Build / Onyx). Shared lifecycle state machine, file-backed task envelopes under .nala/tasks/<task_id>/ (atomic writes + checksum + events.jsonl), harness adapter contract, and a full Grok Build adapter: splash-menu detection (Grok Build Beta + New worktree/Resume/Changelog/Quit), safe bootstrap (no blind Enter on unknown screens), single-line file-pointer delivery, acceptance + start verification (token/tool activity), and idempotent retries by task_id. agent_prompt routes Grok-like peers through this path instead of silent inbox / multi-line paste. Audit: docs/agents/AUDIT_FINDINGS.md.
  • Sidebar agent marker pipeline. Pure resolveWorkspaceAgentIndicator aggregates all panes/surfaces: green = working (running), yellow = needs you (waiting / awaiting_input), blue = starting (banner only). New AgentStatus: 'starting' for gate/splash so cold boots no longer false-green.

Fixed

  • Left-sidebar false-green agent dots. Selected workspaces no longer paint agent-green without an agent; orphan agentStatus without agentName is demoted; background-pane yellow attention bubbles up to the workspace row; AgentDetector gate match emits starting instead of running (real work still promoted by ActivityMonitor).

Changed

  • A2A prefers fewest round-trips (fast path). Agents no longer need to chain surface_list → pane_list → a2a_whoami → a2a_discover before every peer message. MCP instructions default to a single agent_prompt({ target, message }) call: the server resolves callsign/pane/pty (with an 8s discover cache) and delivers full-body paste in one hop. a2a_discover is documented as "only when target unknown". Shared policy module a2a-fast-path.ts + local zero-MCP fleet-snapshot.ts (reads ~/.wmux/principals.json + sessions.json) for host tooling. Avoids Slate's polish/sidebar-notifications files.
  • decideDelivery Grok branch. When harnessId: 'grok-build', splash / gate-match running / session-only states select file_pointer instead of treating the agent as deliverable via inbox-only success.

Fixed

  • False running / silent A2A handoff to Grok splash (Onyx incident). Discovery-style agentStatus: "running" after a Grok startup banner no longer implies READY or TASK_STARTED. Regression tests prove splash → BOOTSTRAP_BLOCKED and force file-backed recovery.
  • Double "NALA" brand in the window chrome. The native title-bar caption next to the app icon no longer shows "NALA" (logo stays); the sidebar header remains the in-app brand so it isn’t stacked twice.
  • PTT Ctrl+Shift missed the first word(s). Capture waited until both modifiers were down; pressing Control a beat early and talking dropped the start of the phrase. Speculative mic now starts on the first chord modifier (partial hold) so audio from Ctrl-before-Shift is kept; full chord still arms PTT. Ctrl+letter while partial cancels arm so shortcuts still work.
  • Duplicate agent-completion notifications (hook + detector + OSC). One turn could fire 3–6 badge/sound events: recordDetector ran after sendNotification, so hook/detector races and multi-line idle footers each announced. Ledger is checked before fire on PTYBridge + DaemonNotificationRouter; OSC 9/777/99 shares the same ledger and stamps idle-fallback suppression. Panel still collapses near-duplicates as a backstop. (Finishes Slate’s polish/sidebar-notifications ledger work on this branch.)
  • Sidebar workspace cards had inconsistent heights and a 🔔 line. Context line is fixed-height; latest unread is a tooltip on the badge pill (Slate sidebar polish).
  • Notification dimmed the active pane. .pane-ring-glow set opacity: 0.6 on the whole pane, so a notification made the workspace look inactive/darker. Glow is border/shadow only; active focused panes never take flash/glow/unread ring chrome. Workspace-level notifies that resolve to the focused leaf now count as “active view” and skip toast/ring fan-out (previously only pty-matched surfaces did).
  • Windows toast spam / false agent notifications. Main process no longer pops native OS toasts on every agent idle/waiting event (that path ignored mute and fleet scale). Native toasts are requested only after renderer notification policy (mute, active surface, toast settings, window focus). Activity-idle “Task may have finished” / “output stopped” signals stay badge/ring-only. ToastManager also rate-limits native toasts (4s global, 20s per pty+title). Panel collapses near-duplicate entries (1.5s).
  • Agent toolbar tooltips clipped under the sidebar. Left chips (paperclip / folder / star / …) used centered tooltips that overflowed the overflow:hidden main pane. Tooltips now support align: start|center|end; left chips use start, right chips use end.

Added

  • System-wide voice dictation. While the NALA desktop process is running (window or system tray — not daemon-only), turn on System-wide dictation in Settings → Dictation or the tray menu. Ctrl+Shift+Space works globally when NALA is unfocused and pastes into the focused app (browser, Notepad, …) via clipboard + synthetic paste. Same Whisper + mic capture as in-app PTT (no second STT stack). Tray tooltip/menu show listening state; tray icon uses a red indicator while recording. In-app hold Ctrl+Shift still owns dictation when NALA is focused.
  • Read Aloud speech bubble. Paste-to-speak card centers a theme-aware cloud speech orb (from the blue-bubble design): idle sits muted/static; paste/processing lights accent colors with slow wispy motion; speaking dances to TTS amplitude. Colors follow --accent-blue / surface tokens so every skin stays cohesive. WebGL field sleeps when idle to keep cost low.
  • Workspace notes tray. New notepad chip on the agent toolbar opens a per-workspace pad that rises from the bottom bar. In-tray Dictate button (focuses the pad, then records) plus bottom-bar mic / PTT — all land in the pad while open. Auto-save with the session, Stamp, copy/clear, word/char stats. Settings: show chip, stamp timestamps, font size, stats, autofocus. Scoped by workspace id; Ctrl+Shift+N / ⌘⇧N; dropped when the workspace closes.

Fixed

  • Active pane could look "grayed out" after alt-tabbing back into the app. On a pure OS-level focus change, the terminal's own DOM focus never actually cleared, so the existing focus self-heal (which only reclaims focus that was truly abandoned) correctly stood down — but nothing re-invoked xterm's own .focus() method, letting its internal focused/blurred visual state go stale even though the pane's border/tab correctly showed it as active. Added a complementary resync on window refocus that only re-affirms focus already on the right element (never steals it from a palette input, toolbar field, or another pane).
  • False-positive "Task may have finished" notifications for panes you never looked at. A background pane (dev server, tail -f, build watcher) could trip the idle-fallback notification purely from cosmetic terminal redraws, even if you'd never made it your active surface. The activity monitor now ignores non-substantive output (pure ANSI/cursor redraws), and the fallback notification itself is gated on the pane having been visited at least once — its status still updates for accuracy, but nothing pings you about work you never asked to watch. Also fixed the daemon-mode fallback showing the generic "Task may have finished" text even for a known agent, unlike the local-mode path, which correctly preserves the agent's identity.
  • agent_prompt missing to (A2A direct inject broken). The tool passed a free-form target field into a2a.task.send, which requires to, so every call failed with a2a.task.send: missing "to" regardless of callsign or pane id. It now resolves callsign / pane / surface / pty via a2a.discover and sends to + pane anchors with full-body PTY paste (silent: false).
  • Devin CLI idle detection stuck on running. Gate-only detection left sessions permanently "busy," so durable A2A tools refused injection into an idle empty prompt. Added Thinking/working and ready-for-input patterns for Devin CLI.
  • Terminal scrollbar thumb unusable on some panes. Stopped forcing Monaco slider height via CSS; pin scrollable element to the terminal surface; set overviewRuler.width: 6 so track metrics match the thin bar (left-split long feeds can drag-scroll again).
  • Provider logos. NALA Agent toolbar mark is now a lowercase mathematical/algorithmic n (not a block capital). Pi quick-launch uses the community Pi mark from pi.dev (block P + square i dot), not Greek π.
  • Read Aloud clip UI clutter. Removed the multi-clip list under the paste box and the duplicate left/right replay controls. One header action button handles stop / retry / replay last; session clips open from a compact ⋮ menu beside it (click to open, Escape / outside click to close).
  • Read Aloud pause + tighter gaps. One primary button now cycles Stop (while loading) → Pause (while speaking) → Resume (from the exact sample via AudioContext suspend) → Replay when idle. Gap trim default 40→90ms with RMS-window silence detection; pre-buffer 150→220ms; double newlines/spaces collapse for continuous speech; clause-splitting extended past the first sentence so mid-stream underruns are rarer.
  • Read Aloud code/shell paste hygiene. Paste normalization (already on the TTS worker hot path — no added latency) replaces fenced code, git/shell/package-manager lines, stack traces, diffs, dense code lines, bare URLs, long paths, and commit hashes with short speakable cues ("Code block.", "Shell command.", "technical detail.", etc.). Contiguous runs collapse to one cue so a long stack trace becomes a single phrase instead of symbol soup.

Added

  • /bark spoken summary (NALA Agent + Pi-wrapped). Put /bark as a leading command or anywhere in a prompt. When the model calls nala_submit_spoken_summary (or the host falls back to final assistant text), the summary is pre-synthesized via tts.bark.prepare so Desktop Kokoro holds audio before the turn ends. Press Enter to play. Pi-wrapped agents use @nala/pi-bridge; NALA Agent TUI arms bark + event settle hooks. /read-aloud is unchanged. Requires Desktop for real Kokoro playback.
  • Daemon + main-pipe tts.bark.* RPC + Desktop bridge. prepare / play / ready / get are registered on the daemon, forwarded from the main pipe (MCP/Pi clients), broadcast for Desktop pre-synth, and play held PCM on Enter.
  • nala_agent_send force flag. Always durable-stores; with force:true also attempts same-workspace PTY inject (skips busy/idle gates). decideDelivery honors force / skipBusyCheck.
  • nala_search_panes / nala_events_poll tool names (legacy wmux_* kept as aliases). Agent-facing MCP instructions, A2A nudges, and channel wake lines say NALA / nala instead of "wmux".

Changed

  • Terminal tab agent-status mark: three distinct states instead of two. Previously the tab mark only ever animated (busy) or froze in place (everything else — waiting, complete, or awaiting input all looked identical). It now shows a spinning figure-8 while the agent works, a slowly rotating ring of dots when it's idle, and the same ring pulsing gently with light when the agent is blocked on a question. The question-pulse is scoped to Claude Code for now — it's the only agent with a deterministic signal for "blocked on a question" vs. "turn just ended"; other agents fall back to the idle ring rather than risk a false pulse. The sidebar status dot also now gives "awaiting input" its own color, distinct from plain "waiting."

  • STT: single bottom-bar indicator + Ctrl+Shift PTT. Removed the floating listening/transcribing hover pill (VoiceStatusChip); the agent toolbar Dictate control is the only status surface and expands with label + mic level while active. Default push-to-talk is hold Ctrl+Shift together (no Space / no third key); sessions still on Space or Control+Shift+Space migrate automatically.

  • NALA branding: acronym slogan + naming. TUI/CLI launch banners use the acronym expansion Naturally Autonomous LLM Agents (replacing the retired SIT. STAY. SHIP. tagline). Documented naming matrix in docs/nala/NAMING.md (NALA / Nala / nala / wmux-legacy). Purged user-facing wmux product-name copy from English (and Korean) locale strings; empty-workspace copy already Nala-neutral. Locale branding guard test on en.ts.

Added

  • M09-M15: Continuation repair passes (pi-missions branch). Six repair passes closing all spec compliance gaps from the fourth-set missions:

    • Repair pass 1: M09 settings controls (show Pi/NALA buttons, permission profile, A2A mode, button order, per-provider diagnostics, launch direct/Pi, repair integration). M10/M11 CAPABILITY_REGISTRY.md, expanded PROTOCOL.md, RPC performance targets, capability version/schema fields. M13 install flags (--version, --channel, --quiet, --repair, --uninstall), nala doctor --install. M14 real /read-aloud with spokenSummary flag, /read-aloud on/off/status, full /play-back forms, /playback alias, persistent audio library, voice settings. 222 tests pass.
    • Repair pass 2: M09 settings migration infrastructure (versioning, backup, restore), 10 dogfood scenario test stubs (27 tests). M10/M11 nala_environment_get MCP tool, A2A message envelope (16 types), transport selection settings (8 settings), nala tool call CLI command. M13 deb/rpm manifest entries, TTS runtime version, workspace resolver with git-root detection, settings schema versioning. M14 spoken summary capability registration, audio library TUI overlay, /clips alias. 263 tests pass.
    • Repair pass 3: M15 privacy/security audit tests (110 tests: secret redaction, prompt injection, path traversal, envelope security, capability security, settings migration security). Performance benchmark tests (20 tests across 6 categories). Retention audit tests (28 tests: retention, quota, atomicity, crash recovery). Release certification document (23 sections). Documentation/UX consistency audit. Clean-machine test matrix. Dogfood scenario documentation. 290 tests pass.
    • Repair pass 4: /settings slash command (list/get/set/reset + /set alias), clip pinning (pin/unpin/isPinned/pinnedClipIds), TTS daemon interface stub, workspace resolver edge case tests (UNC/junction/long paths/Unicode), install lifecycle audit tests (update/repair/rollback/uninstall), final gate certification (31 gates). 192 tests pass.
    • Repair pass 5: M09 Phase 7 routing/fallback tests (33 tests), Phase 8 context menu + double-click dedup (31 tests), Phase 9 lifecycle state machine (46 tests), Phase 10 automated tests (33 tests), Phase 11 final evidence report. M12 continuation audit report (22 sections, 22 gates). M10/M11 capability parity matrix + provider transport matrix. Code quality audit (14 files). 143 tests pass.
    • Repair pass 6 (code quality fixes): HIGH — RPC spec drift fixed (14 missing method descriptors added, 123→137 methods). MEDIUM — 7 swallowed catches in mcp/index.ts made explicit, /fork stub marked deprecated, runtime-manifest version made dynamic. LOW — temp file permissions set to 0o600. 201 tests pass.
    • Repair pass 7: M10 MCP tool coverage (13 missing tools added, 31/31 registered, 32 coverage tests). M12 gate 7 daemon tool handlers (registerToolRpc.ts, 6 handlers, 23 CLI tests). M12 gate 3 RPC spec expansion (73 new methods, 137→210 total). M13 workspace launch integration tests (37 tests). M09 accessibility tests (42 tests). 165 tests pass.
    • Repair pass 8: M11 §19 documentation (10 files: ARCHITECTURE, PROTOCOL, TRANSPORT_MATRIX, PROVIDER_CERTIFICATION, SECURITY, PERFORMANCE, FAILURE_RECOVERY, PACKAGING, OPERATOR_RUNBOOK, USER_GUIDE). M11/M10 final evidence reports. M14 TTS daemon interface expansion (9 methods, protocol v2), device selection module, spoken summary storage module. M09 Phase 11 audit tests (35 tests). M10 Phase 12 conformance matrix (64 tests). 171 tests pass.
    • Repair pass 9: M10 Phase 10 failure/recovery tests (72 tests, 27 scenarios + degraded states). M10 Phase 11 settings tests (70 tests). M10 Phase 8 permission/security tests (124 tests, 13 scopes + defenses). M14 slash command edge cases (39 tests). M13 packaged build verification (32 tests). 336 tests pass.
    • Repair pass 10: M09 Phase 9 process lifecycle and recovery tests (54 tests covering process startup, command rejection, executable missing, provider self-update, auth flow, quota exhaustion, daemon restart, renderer reload, desktop restart, orphan cleanup, Ctrl+C/terminal close, app shutdown, crash restore, multi-workspace, multi-instance, process-tree cleanup). M10 Phase 9 observability and self-awareness tests (35 tests: fleet view, canonical event stream, trace causality, A2A envelope observability, lifecycle observability, route observability). M14 TTS synthesis pipeline tests (35 tests: synthesis request validation, result structure, playback queue, device selection, service status, end-to-end pipeline, protocol version). M13 install script verification tests (34 tests: install.sh flags/HTTPS/SHA-256/CLI shim/PATH, install.ps1 flags/TLS/Get-FileHash/Squirrel, cross-platform consistency, runtime manifest integration). Final certification documents created (FINAL_RELEASE_CERTIFICATION.md, updated CONTINUATION_AUDIT.md, m15-final-gate-certification.md, FINAL_SUMMARY.md). 158 tests pass.
    • Repair pass 11: Fast-follow tasks closing 8 previously documented limitations. Dedicated nala rollback command (M15 gate 29 → READY, 17 tests with --yes/--dry-run/--version flags). UNC path handling in workspace resolver (M15 gate 8 enhanced, 10 Windows-guarded tests covering \\server\share\project, .git walk-up, trailing backslash, forward slash, long UNC paths). Spoken summary → persistent library pipeline (spoken-summary-pipeline.ts, 11 tests). Audio device selection completion (tts-audio-output.ts AudioOutputManager, 14 tests). RPC spec expansion 210→341 methods covering all 331 daemon-registered methods (M12 gate 3 → READY). Spec validation against live registrations (rpc-spec-validation.test.ts, 13 tests). Capability key system reconciliation (capability-reconciliation.test.ts, 21 tests). Verified all 54 m09-process-lifecycle tests pass (previously 2 failing). 86 tests pass. Total project: 3,335 tests pass across 134 test files, 0 failing.
    • Repair pass 12: Dedicated nala update command (M15 gate 27 enhanced, 21 tests with --check/--dry-run/--yes/--channel flags). Chaos testing infrastructure (chaos-test-harness.ts with 4-phase execution, 12 harness tests + 12 scenario tests covering daemon crash, provider crash, socket break, disk full, network partition, concurrent modification, memory pressure, slow provider). M12/M13/M14 final evidence reports (22 + 14 + 15 sections per spec). Certification updates across 5 files: M12 gate 3 → READY (341 methods), M15 gate 29 → READY (rollback command), M15 gate 8 enhanced (UNC tests). Combined gate totals: 42 READY, 7 READY_WITH_LIMITATIONS, 3 PENDING, 0 BLOCKED. 62 tests pass.
    • Repair pass 13: TTS daemon service implementation (tts-daemon-service.ts with createTtsDaemonService factory, 9 methods, FIFO queue with concurrency cap and timeout, AudioOutputManager integration, 30 tests). TTS daemon extraction design document (TTS_DAEMON_EXTRACTION_DESIGN.md, 13 sections covering current/target architecture, 4-phase migration path, protocol design, security, performance, backward compatibility). Auto-updater E2E simulation (31 tests covering update detection, installation, rollback after failed update, version compatibility, channel management, update lifecycle with graceful failure, checksum verification, download retry). Dogfood scenario automation (24 tests covering 5 E2E scenarios: fresh install, update existing install, multi-workspace session, read-aloud workflow, provider crash and recovery). Cross-mission conformance report (CROSS_MISSION_CONFORMANCE_REPORT.md, 486 lines covering M09–M15 conformance, mission-by-mission analysis, cross-mission dependencies, test coverage matrix, documentation coverage, gate status summary, remaining gaps). Fixed 2 flaky performance benchmark tests in m15-performance-benchmarks.test.ts by separating library construction from the measured operation, making the 1000-clip add/delete benchmarks deterministic. 85 new tests pass across 3 test files, no regressions. Total project: 3,436 tests pass across 138 test files, 0 failing.
    • Repair pass 14: M15 final evidence report (19 sections per spec, 31-gate table). Release readiness checklist (8 sections: pre-release, release process, post-release, rollback, emergency, sign-off). Operator runbook updated with 5 new sections (§9 nala update, §10 nala rollback, §11 channel management, §12 TTS daemon, §13 chaos testing). Test coverage report (151 test files, 3,418 test cases, 60.4% file-level coverage, source-to-test mapping, untested files, recommendations). Fixed 2 flaky performance benchmark tests. Certification updates across 5 files.
    • Repair pass 15: Closed test coverage gaps. 8 high-priority untested source files now have dedicated tests (cli-performance-targets 38, profile-types 44, provider-registry 25, provider-certification 31, provider-manifests 97, llm-attribution 33, ids 26, errors 33 = 327 tests). M12-prefixed tests created (m12-benchmark-gate 28, m12-release-gate 15 = 43 tests). A2A E2E delivery latency simulation (23 tests covering envelope creation/validation/enqueue/delivery/e2e/priority/idempotency/rate-limiting/size). 5 type-definition module tests (delegation-types 22, settings-types 18, trace-types 37, artifact-types 27, message-types 27 = 131 tests). 5 smallest test files expanded (chaos-scenarios 4→22, task-types 5→40, turn-ownership 8→20, secretRedaction 8→40, dev-prod-isolation 9→29 = +77 tests). 675 new/expanded tests pass across 21 files. M12 gate 11 (A2A_PERFORMANCE_READY) upgraded to READY. Total project: 4,069 tests across 154 test files, 0 failing. File-level coverage improved from 60.4% to 73.3% (74 of 101 source files with tests).
  • NALA CLI/TUI Agent — Second Set Missions M01-M04 (pi-missions branch). Complete implementation of the NALA CLI/TUI as a first-class client of the shared NALA daemon, covering protocol, TUI, agent OS intelligence, and distribution.

    • M01: Shared daemon CLI runtime and command surface — Client protocol, daemon discovery, multi-client sessions, event streaming, command registry, public commands. 35+ source modules in src/cli/ with barrel exports via src/shared/nala/index.ts. Path traversal protection in resolveWorkspaceConfigPath(). All tests pass.
    • M02: Flagship TUI, NALA agent, and provider runtime — Agent identity and profiles, TUI architecture and terminal capabilities, interactive user flows, slash command system, provider-native backend integration, single-engine invariant (model-call firewall), NALA tools injection, skills/extensions/MCP UX, tasks/artifacts/A2A/delegation in TUI, context/memory/diff UX, voice and notifications, settings/onboarding parity, edge-case validation. 1146 tests across 46 files.
    • M03: Agent OS intelligence, A2A, observability, and security — Workspace intelligence service, context broker, typed scoped reversible memory, durable A2A + task pipeline, one-writer worktree ownership, canonical local observability ledger, security hardening program, provider privacy/conformance lab, eval-gated quality + self-improvement, 2030-ready extension seams, chaos/adversarial/scale testing, desktop/TUI UX integration. 790 tests.
    • M04: Distribution, installers, updates, release, and ecosystem — Product identity + executable strategy, one-line installers (Windows/macOS/Linux/offline enterprise), PATH + shell behavior + completions (14 shells), version/protocol compatibility, trustworthy atomic update system with journaling/rollback, packaging (Pi/skills/extensions/MCP/adapters), provider prerequisites + first-run experience, entitlements/accounts/local-first, documentation generation, release engineering with SBOM/SLSA, clean-machine cert matrix, performance budgets, failure and recovery UX, future ecosystem seams. 533 tests across 27 files.
    • M05a-d: Continuation audits — Validated M01-M04 with fixes: barrel exports, path traversal protection, import path fixes, Zod .strict() consistency (13 schemas fixed). All audit reports in docs/nala-cli/audits/.
    • M05 Final: Production gate — Resolved pre-existing TTS merge conflicts (4 files: audioPlayback.ts, useReadAloud.ts, ReadAloudInput.tsx, audioPlayback.test.ts), restoring pre-buffering, flushPending(), and doReplay refactor. Full test suite: 13049 tests pass (0 failures). Comprehensive architecture, security, and quality audit completed.
  • M01-M03: Single-engine Pi-wrapped provider runtime, product integration, and certification. Three final-integration missions implementing the complete provider-wrapping architecture:

    • M01: Execution taxonomy + launch grants + recovery state machine — New typed schemas in src/shared/nala/: execution-taxonomy.ts (ReasoningOwner, Mission01ExecutionMode, CredentialOwner, BillingClass, AuthRoute, ToolLoopOwner), launch-grant.ts (WrappedLaunchGrant with HMAC signing, LaunchGrantStore with single-use redemption), provider-launch-intent.ts (ProviderLaunchPreference, ProviderLaunchIntent), recovery-states.ts (19-state RecoveryStateMachine with transition enforcement), llm-attribution.ts (LlmAttribution with verified/estimated/unavailable confidence, 30+ extended event types). Single-engine invariant test with throwing model factory.
    • M02: Session header + launch intent dialog + event renderer — New renderer components: SessionHeaderStrip.tsx (compact accessible header with callsign, provider, execution mode, auth/billing, health), SessionInspector.tsx (expandable details drawer), ProviderLaunchIntentDialog.tsx (mode/profile/worktree/permission selection), NormalizedEventRenderer.tsx (renders prompts, responses, tool calls, approvals, diffs, A2A messages, errors).
    • M03: Provider certification + conformance harness + fake ACP fixtures — provider-certification.ts (ProviderCertificationSchema with 16 fields, CertificationStatusSchema), ProviderCertificationStore.ts (CRUD with atomic writes, expiry, revocation), ConformanceHarness.ts (15 standard tests: installation, auth, read-only prompt, controlled edit, MCP, A2A, artifact, permissions, cancellation, restart, quota, cleanup, fallback, single-engine, redaction), FakeAcpFixtures.ts (15+ fixture scenarios: normal, malformed, oversized, crash, canary, forged identity).
    • S01: SessionBindingStore tests — 65 tests covering stampIdentity, get by runId/piSessionId/agentId, status updates, rebind, write outcome tracking, branch ancestry, atomic write/corruption recovery, workspace listing.
    • S04: DiagnosticsService scaffold — DiagnosticsService.ts with configurable provider registry, content-hash caching, test impact suggestions, bounded results, graceful degradation. 21 tests. 6 new RPC methods.
    • S08: License inventory script — scripts/generateLicenseInventory.mjs reading package-lock.json, extracting license/homepage/author for all deps. 9 tests.
  • S06: Skills manager + extension system scaffold. Two new daemon services with full CRUD, file-backed persistence, and RPC wiring: (1) SkillsManagerService (src/daemon/nala/skills/SkillsManagerService.ts) — durable CRUD for named, versioned agent skills (built-in + custom); 5 built-in skills seeded (code-review, test-runner, doc-generator, bug-triage, refactor); built-in re-merge preserves user enable/disable toggles across reopens; custom skill id collision with built-in ids rejected; 23 tests. (2) ExtensionSystem (src/daemon/nala/extensions/ExtensionSystem.ts) — isolated lifecycle management (install → enable → disable → uninstall); activation errors recorded as status='error' without crashing sibling extensions; manifest validation via Zod; upgrade-in-place preserves installedAt; 20 tests. Shared types: skill-types.ts (SkillDefinition, SkillCreateInput, SkillUpdateInput, SkillQuery) and extension-types.ts (ExtensionManifest, ExtensionRecord, ExtensionInstallInput). 13 new RPC methods (nala.skills.*, nala.extensions.*) registered via registerSkillsExtensionsRpc.ts, wired into NalaOrchestrationService, capabilities, and METHOD_CAPABILITY map. Bug fix: deriveId now uses hyphens (not dots) to match the ExtensionRecordSchema id regex; mutate() now persists draft on throw so error states are durable.

  • S07: Fake provider suite for CI + conformance tests. 5 new conformance modules in src/daemon/nala/conformance/: (1) FakeProviderSuite — deterministic fake provider adapters for CI (no real provider processes); 117 tests covering all fake provider behaviors. (2) CertificationRunner — runs certification suites against providers; 8 tests. (3) ChaosSuite — chaos engineering probes for failure injection; 9 tests. (4) PrivacyProbe — privacy compliance verification probes; 13 tests. (5) RoutePolicyService — route policy enforcement and validation; 15 tests. Total: 162 conformance tests, all passing.

  • S03: A2A end-to-end pipeline tests + worktree lease orchestrator. (1) a2aPipeline.e2e.test.ts — 11 end-to-end tests exercising the full NALA orchestration pipeline: task creation → agent assignment → A2A message exchange → delegation with cycle detection → task failure → worktree lease expiry → subtree cancellation → DAG building. Uses real JsonOrchestrationStore and NalaOrchestrationService with FakeNalaAdapter. (2) WorktreeLeaseOrchestrator (src/daemon/nala/worktrees/WorktreeLeaseOrchestrator.ts) — coordinates worktree lease lifecycle with task state; 9 tests. (3) WorktreeLeaseService + WorktreePrepService tests — 19 tests covering lease acquisition, renewal, expiry, and worktree preparation. Total: 48 worktree/orchestration tests, all passing.

  • S08: SBOM generation tests. 7 tests for scripts/generateSbom.mjs covering SPDX 2.3 document validity, app package version correctness, document namespace, relationships array, packages array, empty dependencies handling, and undefined field cleanup. All passing.

  • S02: EcosystemPanel — package catalog, review, quarantine, and safe mode UI. New renderer component (src/renderer/nala/components/EcosystemPanel.tsx) with 4 tabs: (1) Catalog — all packages (built-in + user-installed) with status badges, enable/disable/remove actions, and capability tags; (2) Review — pending intake reports for reviewer approval; (3) Quarantine — quarantined packages with release action; (4) Safe Mode — safe mode toggle with enter/exit and active status display. 8 tests covering tab rendering, RPC loading, empty states, close behavior, and tab switching. i18n keys added to en.ts and ko.ts.

  • Authenticode verification in install.ps1. Windows installer now verifies Authenticode signatures on downloaded binaries before installation. Verification activates automatically once signing certificates are available; gracefully skips (with warning) for unsigned dev builds.

  • NALA static terminal wordmark and launch banner (Mission 07). First-party checked-in banner assets (src/cli/brand/) with 4 variants (wide, standard, compact, plain) for responsive terminal widths. banner.ts module provides deterministic variant selection by terminal width, ANSI color support with NO_COLOR respect, and renderStartupBanner() composition with workspace/provider/runtime/A2A status lines. 18 tests. Phase 1 audit confirmed no superseded mascot/art code to remove.

  • Scrollback restore Fix B — cap-aware reconcile with promoteSession. Suspended sessions that were cap-skipped during recovery (MAX_RECOVER_SESSIONS=40) can now be promoted on demand when the renderer reconcile encounters their ptyId. New daemon.promoteSession RPC (idempotent on active, not-found for missing, resource-exhausted at MAX_SESSIONS=200) reuses the extracted promoteOneSession helper. daemon.listSessions extended with { includeSuspended?: boolean } param to merge disk-only suspended sessions into the response. AppLayout reconcile routes suspended ptyIds through promote → reconnect instead of fallback-create. 20 new tests across daemon + renderer.

  • Pane split max depth guard. splitPane now enforces a MAX_SPLIT_DEPTH=6 limit in addition to the existing 20-pane count cap, preventing infinitely nested splits. Depth-blocked splits set a blockedAtDepth flag and show a warning toast. 3 new tests.

Fixed

  • TTS corrupt model recovery. On Protobuf parsing failed / corrupt Kokoro ONNX cache load, the TTS worker clears transformers.js model caches and retries once; if q4 still fails, falls back to q8. User-facing errors are marked recoverable (retry / change quality) instead of a fatal crash.

  • nala:rpc cold-start daemon wait. Internal nala:rpc no longer immediately throws DAEMON_DISCONNECTED before first daemon ping — waits up to ~18s with backoff for cold start. Mid-session disconnects after a successful connect still fail immediately.

  • STT push-to-talk default + mic sensitivity. Default PTT hotkey is Control+Shift+Space (avoids bare Space false starts). Settings → Dictation adds a mic sensitivity slider (0–100) that maps to the VAD energy floor in audioCapture (replacing the hard-coded 0.008 threshold). Preference persists in session data.

  • Agent idle detection + tab fig-8 freeze. Claude Code end-of-turn verbs (Cogitated for 28s, Cooked for …, etc.) now map to waiting; esc to interrupt remains non-waiting (in-flight). When ActivityMonitor goes quiet on a known agent, PTYBridge keeps waiting + agentName instead of clearing to idle. Renderer isSubstantivePtyOutput ignores pure ANSI/TUI chrome so the tab figure-8 stops spinning after a turn ends.

  • Workspace auto-trust for A2A / plugin-exec. User-opened and session-restored workspaces call workspaceTrust.setTrust; a one-shot ensureWorkspacesTrustedOnce on AppLayout mount also trusts the slice-init first workspace so boot agents are not stuck behind an untrusted gate (does not re-run, so a later user untrust is not clobbered).

  • Single-pane active border toned down. Solo leaves no longer draw a bright active accent border (shouldShowActivePaneAccent); multi-pane focus border still uses a dimmed cursor accent.

  • Toast duplicate suppression (stale-daemon spam guard). pushToast now suppresses exact message+level duplicates while an identical toast is still in the list, returning the existing toast's id. Prevents stale-daemon error spam from stacking to the MAX_TOASTS=10 cap when multiple components poll the same dead RPC method. Different message or level still admits; re-admits after auto-dismiss. 4 new tests.

  • Git-oracle test flakiness hardened. The g() helper in diff.handler.test.ts and all git calls in diffParse.test.ts now use retry (2x) with 100ms backoff and 15s per-call timeout, matching the applyPatchInRepo pattern. Eliminates bare execFileSync flakiness under heavy parallel I/O.

  • Mission 05 production hardening gaps closed. Three missing Mission 05 deliverables implemented: (1) ReconciliationService (src/daemon/nala/orchestration/ReconciliationService.ts) — detects ambiguous write outcomes (timeout, crash, partial response) by checking actual file/git state against expected hashes; returns APPLIED/NOT_APPLIED/AMBIGUOUS/CONFLICT verdict; bounded retry with reconciliation never auto-retries AMBIGUOUS or CONFLICT; 22 tests. (2) RemotePolicyService (src/daemon/nala/policy/RemotePolicyService.ts) — fetches signed policy manifests from NALA_POLICY_URL, verifies SHA-256 hash pin, caches with TTL (default 24h), falls back to BUNDLED_SAFE_POLICY on any failure; applies remote kill-switches to local route policy; 12 tests. (3) ProviderDriftDetector (src/daemon/nala/policy/ProviderDriftDetector.ts) — compares installed provider versions against compatibility matrix, returns drift severity (none/warning/critical), batch detection for all providers; 9 tests. (4) UninstallCleanupService (src/daemon/nala/maintenance/UninstallCleanupService.ts) — removes Pi runtime artifacts and NALA session state with dryRun default true, bounded by orchestration root and Pi home, never touches user data; 15 tests. (5) NALA model provider manifest — added nala provider to provider-manifests.ts as a first-party orchestration decision provider (not code-writing), internal transport, local_runtime auth, enabled by default false. 8 new RPC methods registered (nala.reconciliation.check/retry, nala.policy.remote.fetch/status, nala.policy.drift.detect/report, nala.maintenance.cleanup.preview/execute). 2 new capability keys (nala.policy, nala.maintenance). API reference regenerated.

  • i18n for NALA renderer components. All 20 new NALA renderer components (SettingsPanel, OperationsDashboard, LaunchCard, UnifiedApprovalDialog, FirstRunWizard, RepositoryMapPanel, SymbolExplorerPanel, MemoryManagerPanel, ContextPreviewPanel, OrchestrationSettingsPanel, ContextMemorySettingsPanel, ObservabilitySettingsPanel, ProfileEditor, ProfileCompareView, RoleCard, WorkflowCard, AlertList, FleetPanel, SpanTreeViewer, NalaPanelsHost) now use the useT() hook and t() translator instead of hardcoded English strings. 60+ new translation keys added to src/renderer/i18n/locales/en.ts under the nala.* namespace. Other locales fall back to English automatically for missing keys (by design).

  • NALA panel wiring and renderer component tests. All 12 top-level NALA panels (Settings, FirstRunWizard, OperationsDashboard, LaunchCard, UnifiedApprovalDialog, OrchestrationSettings, ContextMemory, Observability, RepositoryMap, SymbolExplorer, MemoryManager, ContextPreview) are now mounted in the app shell and reachable from three entry points: (1) Command palette — 12 new Ctrl+Shift+P entries including a launchCard sub-mode that lists each provider; (2) NalaOrchestrationBar "NALA" dropdown — 10 drawer panels; (3) Settings → NALA tab "NALA Panels" section — buttons opening all 10 drawers. Central open/close state lives in uiSlice (NalaPanelId union + openNalaPanel/closeNalaPanel actions, plus nalaLaunchCardProvider and nalaApprovalRequest slots for the modal/dialog). NalaPanelsHost is mounted in AppLayout and renders every panel conditionally. FirstRunWizard auto-opens once per profile via a localStorage marker (nala.firstRunCompleted). Added 7 new renderer test files (148 tests, all passing): SettingsPanel, OperationsDashboard, LaunchCard, UnifiedApprovalDialog, FirstRunWizard, RepositoryMapPanel, MemoryManagerPanel — covering tab rendering, RPC flows, CRUD operations, error states, and closed-state behavior using the established window.electronAPI.nala.rpc mock bridge pattern. Regenerated docs/api/reference.md to include the new Pi bridge alias RPC methods.

  • Workspace intelligence UI (S04 WS9). 4 new renderer components + 1 shared helper: (1) RepositoryMapPanel — file inventory viewer with per-file identity (path, language, size, classification, sensitivity), summary stats, ignore-policy block, indexing controls (scan/pause/resume), and filters by path/language/classification/sensitivity; (2) SymbolExplorerPanel — two-column symbol viewer with searchable list, kind-colored badges, detail pane (signature/docs/source/confidence), and outgoing/incoming dependency edges; (3) MemoryManagerPanel — full CRUD memory manager with scope/category/state filters, inline editor, retention policy editor, and admit/reject/pin/expire actions; (4) ContextPreviewPanel — context assembly viewer with strategy selector, token budget bar, per-item cards with "why this context" selection reasons, and compact action; (5) s04Rpc.ts — shared RPC helper + useActiveWorkspaceId hook.

  • Pi bridge tools — task, worktree, terminal, search, trace (13 new tools). The Pi bridge extension (integrations/pi/package/src/tools/) now exposes 13 new tools beyond the existing A2A and artifact tools: (1) Task tools — nala_task_get, nala_task_update, nala_task_complete, nala_task_list; (2) Worktree tools — nala_worktree_request, nala_write_lease_status, nala_worktree_release; (3) Terminal tools — nala_terminal_list, nala_terminal_send (returns descriptive error — PTY lifecycle is owned by main process); (4) Search tools — nala_workspace_search, nala_symbol_search; (5) Trace tools — nala_trace_note, nala_progress_update. All tools use env-var identity (not model-supplied IDs), Zod input validation, and graceful error handling. 10 new daemon RPC alias handlers in registerPiBridgeAliasRpc.ts bridge tool method names to existing services.

  • Operations dashboard (S05 WS6). 4 new renderer components for NALA observability: (1) OperationsDashboard — 4-tab drawer (Fleet/Timeline/Alerts/Diagnostics) with live fleet view, per-run timeline with span tree, alert management with ack/resolve, and diagnostics with provider status, resource usage, and support bundle export; (2) SpanTreeViewer — expandable/collapsible span tree with color-coded status dots and detail expansion; (3) AlertList — alert list with severity badges, filters by severity/state/source, and acknowledge/resolve actions; (4) FleetPanel — agent grid with status indicators, callsign, provider, model, route, task, worktree, pending questions, and cancel/steer/open actions, auto-refreshing every 10s.

  • Launch card, unified approval UX, and settings panels (S06 WS4/WS6/WS7). 5 new renderer components: (1) LaunchCard — detailed provider launch card with route selector, profile summary, capability badges, billing class indicator, expandable details, compact/full modes, and changed-state indicator; (2) UnifiedApprovalDialog — comprehensive approval dialog for 9 permission types (filesystem, shell, network, git, browser, MCP, spawn, secret, external path) with risk-level color coding, 5 decision choices (deny once, allow once, allow for run, allow for workspace, allow for profile), and capability summaries; (3) OrchestrationSettingsPanel — orchestration mode, delegation limits, default roles, review rounds, supervisor questions, notification preferences; (4) ContextMemorySettingsPanel — context strategy, indexing controls, memory categories with per-category retention, secret scanning toggle, memory search interface; (5) ObservabilitySettingsPanel — trace content mode, retention by event class, exporter consent toggles, support bundle export, "what is stored" inspector.

  • S06 renderer UI — settings, profiles, roles, workflows, first-run wizard. 6 new renderer components in src/renderer/nala/components/: (1) SettingsPanel — tabbed drawer (Settings/Profiles/Roles/Workflows) with scope selector, effective settings list with inline edit and reset, export, migration list; profile CRUD with clone/compare/export/health; role and workflow card grids. (2) ProfileEditor — modal for creating/editing provider profiles with write-policy checkboxes. (3) ProfileCompareView — side-by-side diff modal showing field-level differences. (4) RoleCard — role display with capabilities, boundary policy, delegation rights, and signing status. (5) WorkflowCard — workflow display with ordered stages, role assignments, and parallel fan-out. (6) FirstRunWizard — 5-step wizard (welcome → provider → profile → review → done) that creates a profile and writes initial settings. All components use the established window.electronAPI.nala.rpc bridge and Tailwind CSS variable theming.

Fixed

  • MCP pinned terminal route invalidation when workspace dies (P2). When an external MCP caller claimed a dedicated workspace and the user manually closed that workspace mid-session, the process-lifetime pin in paneResolver kept pointing at the dead PTY, causing the caller's terminal tool to permanently fail until MCP restart. invalidateWorkspaceId() (the stale-identity self-heal path in src/mcp/index.ts) only cleared the workspaceResolved verified-cache flag, not the pin. Added clearPin() export to paneResolver.ts and wired it into invalidateWorkspaceId() so both call sites in callRpc (success-path and error-path stale-result checks) now clear the pin alongside the verified cache, enabling self-healing re-claim on the next call. 4 new tests added.

  • Terminal DSR-CPR response leak after channel mention paste (P2). When a channel mention was pasted into a pane, the cursor position response (ESC[<row>;<col>R, e.g. ESC[40;3R) was not consumed by the TUI during rapid paste/repaint transitions and instead leaked back through the PTY output stream as literal text (e.g. ;3R40 appearing in bulk in the scrollback). This is the same class of bug as the 2026-07-04 CPR feedback storm (replay-time, fixed by replayQuerySanitizer.ts), but on LIVE output. Added a stateful CprResponseFilter (src/renderer/terminal/cprResponseFilter.ts) that strips CPR response sequences from PTY output before they reach xterm.js. The filter is per-terminal (instantiated in useTerminal's main effect) and handles chunk-boundary splits by buffering incomplete CSI prefixes. DSR queries (ESC[6n) are preserved — xterm must see those to generate the responses the TUI expects. CPR responses are terminal reports, never display content, so stripping is lossless for the user. 19 unit tests added.

  • METHOD_CAPABILITY map completed for all 273 nala. RPC methods.* The daemon registers 273 nala.* RPC methods via .onRpc(), but METHOD_CAPABILITY in src/main/mcp/methodCapabilityMap.ts only knew about 72 of them — leaving 201 methods ungated and absent from the RpcMethod union type and ALL_RPC_METHODS array in src/shared/rpc.ts. This caused tsc errors (Record incompleteness), 2 test failures in methodCapabilityMap.test.ts (totality + capability validity), and API reference drift. Added all 201 missing methods to rpc.ts, methodCapabilityMap.ts (all using wmux.internal capability, matching the existing substrate-internal pattern), and regenerated docs/api/reference.md. Root cause: DaemonPipeServer.onRpc() accepts an untyped string, so methods could be registered without being added to the type system.

  • Provider conformance, privacy, chaos, and certification lab (Mission S07). NALA now has a provider conformance lab for certifying that each provider route meets the universal agent contract before being advertised. (1) CertificationRunner (src/daemon/nala/conformance/CertificationRunner.ts) — runs 13 standard scenarios (read-only, plan, write, tool, approval, A2A, cancel, resume, model switch, error, quota, offline, cleanup) against a UniversalAgentAdapter, produces a tamper-evident CertificationReport with SHA-256 resultsHash and per-scenario pass/fail observations. Error and quota categories expect graceful failures. (2) RoutePolicyService (src/daemon/nala/conformance/RoutePolicyService.ts) — bounded JSON persistence (max 512 policies) for per-provider, per-route metadata with 5 statuses (certified, experimental, personal_only, blocked, unknown), kill-switch (disableRoute), re-certification requirement (enableRoute), version-window checking (isRouteSupported), and 30-day offline cache (cacheForOffline). (3) ChaosSuite (src/daemon/nala/conformance/ChaosSuite.ts) — 11 standard chaos scenarios (process crash, network partition, disk full, OOM, slow response, corrupted state, permission revoked, token expiry, concurrent writes, dirty git, missing files) with filesystem reconciliation that never auto-retries ambiguous writes. (4) PrivacyProbe (src/daemon/nala/conformance/PrivacyProbe.ts) — creates a synthetic repo with allowed/ignored/excluded files, nested .git, symlinks, and canary secrets in .env/secrets/node_modules/dist/nested-repo; verdict fails on canary exposure, excluded/ignored file access, or blocked network calls. (5) Shared schemas (src/shared/nala/conformance-types.ts) — Zod schemas for certification scenarios/results/reports, route policy metadata, chaos scenarios/results, and privacy probe results. (6) RPC (registerS07Rpc.ts) — 14 endpoints for scenario listing/validation, route policy CRUD, chaos scenario listing/validation, and privacy probe helpers. (7) Tests — 79 new tests (17 schema, 8 certification, 15 route policy, 9 chaos, 13 privacy, plus existing conformance kit).

  • Settings, profiles, and roles/workflows system (Mission S06). NALA now has a layered settings resolver, provider profile system, and signed roles/workflows for governing agent behavior across workspaces. (1) SettingsService (src/daemon/nala/settings/SettingsService.ts) — effective-settings resolver with 4-scope precedence (defaults → user → workspace → session), setting definitions with securitySensitive flag, CRUD, import preview (quarantines unknown keys), migrations with version stamps, health test, and export (never includes secrets, only secretRefs). Corruption recovery via atomic read with fallback. (2) ProfileService (src/daemon/nala/profiles/ProfileService.ts) — provider profile CRUD, clone, compare (diff of providerId, authRef, billingClass, model, mode, permissions), import preview (detects missing providers, unknown capabilities, executable resources), import (forces autoActivate: false), export (strips secrets), health test. 3 built-in profiles (full-access, supervised, read-only). Bounded to 1,000 profiles. (3) RoleService (src/daemon/nala/roles/RoleService.ts) — 7 signed built-in roles (scout, researcher, planner, implementer, reviewer, verifier, oracle) with capability sets and boundary policies; 7 built-in workflows (research-then-implement, plan-review-implement, parallel-scouts, sequential-review, verify-after-implement, oracle-consult, full-cycle); custom role/workflow creation with signing. (4) Shared schemas (src/shared/nala/settings-types.ts, src/shared/nala/profile-types.ts) — Zod schemas for settings, profiles, import previews, comparisons, migrations, and resets. (5) RPC (registerS06Rpc.ts) — 20+ endpoints for settings (resolve, update, reset, export, importPreview, importSnapshot, definitions, migrations), profiles (upsert, get, list, delete, clone, compare, importPreview, importProfile, export, healthTest), and roles/workflows (list, get, create custom). (6) Tests — 47 new tests (16 settings, 18 profiles, 13 roles).

  • Trace bridge wiring (Mission S05 fix). JsonOrchestrationStore now forwards every appendActivity call to TraceIngestionService via setActivityListener(). Added 30+ dotted-form event kind mappings (task.created → task_created, etc.) to TraceIngestionService. Added eventsByTrace index to TraceLedgerService for efficient trace-level queries. Relaxed AlertService and ExporterService to accept partial upserts with schema-parsed defaults.

  • Production packaging, updates, and release certification documentation (Mission S08). 11 new operational documentation files: (1) pi-release-dependency-graph.md — frozen dependency graph with SBOM, provenance, hashes, licenses, and scans; (2) pi-reproducible-build.md — reproducible build for all components with integrity verification; (3) pi-installer-platform-matrix.md — Windows/macOS/Linux installer integration matrix; (4) pi-data-migration-rollback.md — transactional migrations, backup, compatibility windows, rollback; (5) pi-secure-updates-remote-policy.md — TUF-inspired update metadata, channels, rollout, remote policy; (6) pi-saas-entitlements-usage.md — feature classification, entitlements, usage metering, account lifecycle; (7) pi-clean-machine-upgrade-certification.md — clean VM test matrices for all platforms; (8) pi-security-resilience-certification.md — security/chaos/soak suites with 24-hour soak plan; (9) pi-support-playbooks.md — 10 support playbooks (auth, provider drift, quarantine, etc.); (10) pi-final-release-gate-rollout.md — final gate evaluation with gradual rollout; (11) RELEASE_CERTIFICATION_REPORT.md — final certification report with S01-S08 gate verdicts (46/52 PASS, 6 EXTERNAL).

  • Continuation audit report (Mission S09). docs/desktop-saas/MISSION-S09/CONTINUATION_AUDIT_REPORT.md — full audit of S06-S08 with 10 findings, 5 fixes applied, gate verdicts, and final release gate verdict of READY_WITH_EXPLICIT_LIMITATIONS. All 1,014 NALA tests pass across 71 test files with 0 TypeScript errors.

  • Production hardening (Mission 05). NALA's orchestration layer is now hardened for production with 8 batches of security, governance, and reliability improvements. (1) Electron/IPC/process hardening (B1) — CSP now enforced in development mode with Vite HMR allowances; frame-ancestors 'none' and base-uri 'self' added to prevent clickjacking and base tag injection; shared Zod schemas (src/main/ipc/schemas.ts) for high-risk IPC handlers (pty.create, pty.resize, pty.dispose, shell.openExternal, shell.openPath, clipboard.write, diff.request, fanout.request, worktask.ref, tts.speak) with size limits, format validation, and enum constraints; pty.create handlers validate payload shape before process-spawning code runs. (2) Resource governance and emergency stop (B2) — centralized ResourceLimitsSchema with safe defaults for task runtime, idle timeout, artifact size, storage quota, tool result size, delegation depth, concurrency, message size, provider health-check interval, process count, and retry limits; all limits configurable via NALA_* environment variables; global emergency stop via NALA_EMERGENCY_STOP=1 disables all non-core features in the capability manifest; validation helpers for artifact/tool result/message/delegation depth/runtime; exponential backoff calculator. (3) Provider compatibility matrix and policy manifest (B3) — signed ProviderPolicyManifestSchema with per-provider, per-transport route status (approved, experimental, personal-only, blocked); BUNDLED_SAFE_POLICY as last-known-good fallback; ProviderCompatibilityMatrixSchema with version ranges, known limitations, fallback transports, and per-architecture certification; validateProviderVersion() returns warnings for version drift; policy evaluation helpers (isProviderBlocked, isWrappedSessionDisabled, meetsMinimumVersion). (4) Fake provider conformance kit (B4) — cross-platform fake CLI scripts simulating ACP, Codex, Claude SDK, JSONL, and hostile protocols; 9 predefined conformance scenarios (normal, malformed, oversized, protocol pollution, unauthenticated, slow response); writeConformanceKit() generates all scenario scripts for CI testing without consuming plans or secrets. (5) Diagnostics bundle and privacy controls (B5) — exportable redacted DiagnosticsBundleSchema with versions, providers, launch profiles, compatibility status, recent events, sessions, resource limits, and optional logs; secret scanner detects OpenAI/Anthropic keys, Bearer tokens, passwords, private keys; export refuses if secrets are found; PrivacySettingsSchema with telemetry disabled by default; 7 minimized telemetry event types (no prompt content, code, or terminal output). (6) Versioned state migrations (B6) — MigrationRegistry with idempotent, resumable, transactional migrations; automatic backup before destructive migrations; 3 NALA state migrations (initialize profiles, add schemaVersion, add resource governance defaults). (7) Future NALA provider seam (B7) — NALA_MANIFEST reserves the 'nala' provider ID through the normal manifest system; FakeNalaAdapter proves the seam works without UI special-casing; 'nala' added to BUILTIN_PROVIDER_IDS; nala-default environment policy added. (8) Documentation (B8) — architecture, threat model (8 threats with mitigations), runbooks (6 procedures), and release checklist. 169 new tests across 7 test files.

  • Unified Pi-wrapped launch UX and provider routing (Mission 04). NALA now has a versioned launch-profile and route service that governs how each provider button behaves when clicked. (1) LaunchProfileService — a daemon-side service (src/daemon/nala/providers/LaunchProfileService.ts) with profile CRUD, default route assignment, capability validation, and route resolution with fallback options. Each provider has an AgentLaunchProfile with a route (direct, pi-wrapped, ask-each-time), wrapped execution style, auth/model/mode/permission/cwd/resume policies. Default routes: Pi button → pi-wrapped, all others → direct (no silent migration). (2) Route-aware launcher — the useLaunchRoute hook resolves the route via nala.launch.resolve RPC, then dispatches to direct or pi-wrapped launch. Falls back to direct if the daemon doesn't support the route RPC. (3) LaunchRoutePicker — modal dropdown for ask-each-time route selection. (4) ProviderButtonContextMenu — right-click menu on provider buttons for quick route switching with health badges. (5) SessionRouteBadge — π badge for pi-wrapped sessions in tab titles and Fleet cards. (6) CapabilitiesDrawer — slide-out panel showing session identity, provider health, 15 capability flags, available models, and notices. (7) PiWrappedOnboarding — first-run dialog explaining NALA tools, A2A messaging, permission bridge, and session resume. (8) ProviderDoctorPanel — diagnostic panel integrating the Provider Doctor with per-provider policy status, installation, auth, transport, and model discovery. (9) SessionResumePanel — recovery panel for pi-wrapped sessions with Resume through Pi / Resume directly / Start fresh options. (10) LaunchProfileSettings — settings section in the NALA tab with per-provider route selector, profile list, and Provider Doctor inline view. (11) Feature flag — nala.launchProfiles capability with NALA_DISABLE_LAUNCH_PROFILES=1 emergency kill switch. (12) Pi-wrapped resume — ResumeBinding extended with launchRoute field; piWrappedResumeEnvPatch() re-injects NALA bridge env vars on resume. 6 new RPC methods, 31 new tests.

  • Universal provider adapter contracts (Mission 02). NALA now has a versioned, schema-validated contract system for universal provider adapters. (1) universal-provider-types.ts — defines ProviderManifest, UniversalAgentAdapter interface, NalaAgentEvent normalized event envelope, AuthMode (7 modes: official_cli_subscription, official_cli_api, pi_oauth, direct_api, coding_plan_key, local_runtime, unavailable), ProviderTransport (8 transports), ProviderCapabilityDeclaration (26 granular capabilities), and all associated Zod schemas with validation helpers. (2) provider-manifests.ts — manifests for all 10 supported providers: Codex, Claude, Grok, Cursor, Devin, Z.AI, OpenAI, Anthropic, Ollama, LM Studio. Each manifest declares transports, auth modes, capabilities, documentation URLs, terms URLs, minimum versions, and policy revisions. Z.AI is not enabled by default and does not claim to be keyless. (3) ProviderRegistry.ts — central registry with per-provider environment policies that prevent silent billing-mode switches. Claude subscription mode removes ANTHROPIC_API_KEY; Grok subscription removes XAI_API_KEY; Devin stored-login removes WINDSURF_API_KEY; Codex ChatGPT-plan removes OPENAI_API_KEY. The buildChildEnvironment method constructs sanitized child environments without deleting user environment variables globally. (4) SecureCredentialService.ts — stores provider API keys using Electron safeStorage (Windows DPAPI). Secrets are never returned to the renderer; only redacted metadata (last 4 characters) is exposed. Decrypted values exist only in the trusted main process for minimum time. (5) AcpAgentAdapter.ts — hardened ACP (Agent Client Protocol) client adapter with JSON-RPC 2.0 over stdio, protocol version negotiation, session management, streamed events, permission handling, cancellation, bounded line/message sizes, stderr separation, request timeouts, and process-tree cleanup. (6) StructuredCliAdapter.ts — fallback adapter for providers that don't support ACP or a dedicated SDK transport. Runs the provider CLI in one-shot or streaming-JSON mode with shell:false, prompts via stdin, bounded output, and separate stdout/stderr. (7) ProviderAdapterFactory.ts — factory that creates the appropriate adapter for each provider based on the preferred transport and fallback chain. Codex uses codex-app-server → structured-cli; Claude uses claude-agent-sdk → structured-cli; Grok/Cursor/Devin use acp-stdio → structured-cli; Z.AI uses anthropic-compatible/openai-compatible; local runtimes use local-runtime. (8) ProviderDoctor.ts — diagnostic service that reports on all providers' installation, auth, capabilities, policy status, and billing-mode conflicts. Includes Z.AI plan eligibility caution, Claude policy date notice, and Grok/Cursor collision detection. (9) Fake ACP agent test harness — a minimal ACP-compatible agent (fake-acp-agent.mjs) for contract testing with modes for normal, slow, malformed, crash, permission, and stderr-noise scenarios. 19 new tests (9 ACP contract + 10 ProviderDoctor) validate all paths. 285 total tests pass.

  • Pi project trust bridge (Mission 01 audit). A new PiProjectTrustBridge (src/main/pi/PiProjectTrustBridge.ts) maps NALA workspace trust to Pi project trust with explicit user approval. Unknown workspaces default to untrusted (fail-closed). Trust decisions are persisted to trust/trust-decisions.json with approver, timestamp, and workspace identity hash. Identity changes (path modification) trigger re-approval. IPC handlers pi:trust:get, pi:trust:approve, pi:trust:revoke expose trust management to the renderer. The diagnostic report now includes trust decisions.

  • Pi bridge /nala-doctor command and structured handshake (Mission 01 audit). The @nala/pi-bridge extension now provides a /nala-doctor slash command with full diagnostic output (protocol version, runtime version, handshake status, capabilities, PI_CODING_AGENT_DIR). The handshake protocol is now structured as a versioned NalaPiHello JSON message emitted to stderr, with protocol version negotiation and rejection of unsupported versions. The handshake includes workspace ID, agent ID, task ID, session ID, launch nonce, and capability list.

  • Pi agent harness integration (Mission 01). NALA now bundles and manages a Pi runtime as a sixth built-in provider. Pi (v0.80.7, MIT) is a universal agent harness that runs as an app-managed subprocess in a NALA terminal pane. (1) PiAdapter — a new BaseAdapter subclass (src/daemon/nala/providers/PiAdapter.ts) that resolves the app-managed Pi runtime (never searches PATH for a global pi), builds launch specs with node <cli.js> args, and scopes Pi's config directory to NALA's app-data via PI_CODING_AGENT_DIR. (2) PiRuntimeManager — a main-process singleton (src/main/pi/PiRuntimeManager.ts) that resolves, verifies, and manages the bundled Pi runtime. It checks required files (cli.js, rpc-entry.js, package.json), probes pi --version, and injects NALA_PI_RUNTIME_PATH and NALA_PI_HOME env vars into the daemon spawn. (3) Quick-launch button — a Pi button with a π logo appears in the NALA Orchestration Bar alongside Claude, Codex, Devin, Cursor, and Grok. Clicking it launches Pi in a terminal pane with NALA identity env vars. (4) @nala/pi-bridge extension — a first-party Pi extension (integrations/pi/package/) that injects NALA workspace context, emits lifecycle events, and provides /nala-status and /nala-workspace slash commands. (5) Integrations structure — integrations/pi/ contains the pinned provenance lock (pi-runtime.lock.json), zero-patch ledger, fetch/build/verify scripts, and the vendored runtime. (6) IPC diagnostics — pi:health, pi:diagnostics, and pi:set-dev-override IPC handlers expose runtime health to the renderer for Settings display. (7) Baseline documentation — docs/pi-integration/01-baseline.md documents the full agent launch path, Pi architecture, and ADR-001 (subprocess decision).

Fixed

  • PiAdapter and PiRuntimeManager fail-closed security hardening (Mission 01 audit). The PiAdapter previously had dev-mode fallbacks that searched process.cwd() and __dirname for the Pi runtime, violating the fail-closed invariant (Mission 01 §2.5, §6.1). A malicious cli.js placed in the current directory could have been executed. The adapter now uses ONLY the NALA_PI_RUNTIME_PATH env var injected by PiRuntimeManager. The PiRuntimeManager's dev-mode fallback is now gated on NALA_DEV=1 so production builds never resolve a runtime from the cwd. A regression test verifies that a fake cli.js in the cwd is NOT found. The PiHome fallback was also changed from ~/.nala/pi to os.tmpdir()/nala-pi-fallback to avoid touching the user's global Pi config.

  • Multi-tier agent executable resolution with cross-platform support. NALA's agent launcher now uses a 4-tier fallback chain to find agent CLIs on any platform: (1) Custom path — user-specified path; bare names like claude are treated as callsigns (shell resolves via PATH), absolute paths use the full path. Stale custom paths fall through to lower tiers instead of failing. (2) PATH search — preferred tier; produces a portable callsign (devin ... instead of & 'C:\...\devin.exe' ...). (3) Platform known paths — all 5 built-in adapters (Claude, Codex, Devin, Cursor, Grok) now return per-platform install locations for Windows (AppData), macOS (homebrew, ~/.claude/bin, ~/.local/bin), and Linux (~/.local/bin, /usr/local/bin). Previously only Windows AppData paths were hardcoded — macOS/Linux only worked via PATH search. (4) Package manager dirs — checks npm global, homebrew, cargo, pipx, and volta directories as a last resort. (5) Cross-platform PATH refresh — refreshSystemPath() now spawns a login shell on macOS/Linux (/bin/zsh -l -c "echo $PATH" or $SHELL -l -c "echo $PATH") to catch agents installed via brew install, pipx install, etc. after the app launched. Windows PATH refresh (registry merge) is unchanged. (6) Generalized disambiguation — rejectPathPatterns on BaseAdapter replaces the Cursor-specific resolveExecutable override. Cursor rejects paths containing grok directories; Grok rejects paths containing cursor-agent directories. Rejected candidates are skipped and resolution continues to the next tier. (7) Version validation — each adapter declares an expectedVersionPattern regex; probe() emits a VERSION_MISMATCH notice when the binary's --version output doesn't match, catching shell aliases or wrong tools shadowing the callsign. (8) Probe result enrichment — ProviderProbeResult now includes matchedVia (which PATH entry, known path, or package-manager dir matched) for diagnostics. resolvedVia expanded from ['path','known','custom'] to include 'pkg-manager', 'bundled', 'not-found'. The nala.provider.detectAgents RPC exposes matchedVia. (9) Custom provider platform paths — CustomProviderDefinitionSchema gains optional knownPathsByPlatform ({ win32?, darwin?, linux?, other? }) for platform-specific install locations. The flat knownPaths array remains as a cross-platform fallback. 48 new tests in resolutionTiers.test.ts cover all tiers, platform paths, callsign inference, disambiguation, version patterns, and probe enrichment.

  • Voice dictation master toggle + configurable PTT hotkey. Users can now fully enable/disable voice dictation and customize the push-to-talk activation key from Settings → Dictation. (1) Master dictation toggle — a new "Voice dictation" on/off switch at the top of the Dictation settings tab controls all voice features. When off, the Dictate button is hidden from the toolbar, the PTT hotkey is not intercepted, and toggleDictation() shows a "Voice dictation is disabled in Settings" warning. No microphone audio is captured. (2) Configurable PTT hotkey — the push-to-talk activation key is no longer hardcoded to Space. A new "Push-to-talk hotkey" card in settings lets the user click a capture button and press any key (optionally with Ctrl/Shift/Alt/Meta modifiers) to set the activation combination. A "Reset to Space" button restores the default. The keyboard shortcuts reference card now shows the configured hotkey (e.g., "Hold Ctrl+Space") instead of a hardcoded "Hold Space". The hotkey string format is Modifier+...+KeyCode (e.g., Control+Space, Shift+Alt+KeyV); new utility functions parseHotkey, hotkeyMatches, formatHotkey, and hotkeyLabel handle parsing, matching, and display. The SttConfig type expanded from 12 to 14 fields (dictationEnabled, pttHotkey); all settings persist across sessions via session.json and are pushed to the main process via VOICE_CONFIG_UPDATE IPC. For the default Space hotkey, the existing shouldHandleSpacePtt policy (IME guard, xterm helper check, editable target check) is preserved; for custom hotkeys, a simpler "terminal must be focused" check is used. Short-tap behavior (type a space on quick Space tap) is preserved for Space only — custom hotkeys cancel silently on short tap.

  • Gatsby theme — 1920s Art Deco easter-egg skin. A new "Gatsby" theme in Settings → Appearance transforms the app into a Jazz Age masterpiece. The palette is drawn directly from Fitzgerald's symbolism: deep bay-black night (#0A0E08), champagne-cream text (#F4E9C1), rich gold accent (#D4AF37), and the iconic green light (#7EC8A0) as the secondary/link accent — "a single green light, minute and far away, that might have been the end of a dock." Beyond the standard color tokens, the theme layers Art Deco decorative elements throughout the UI: (1) Gold gradient headings — h1/h2 titles render with a shimmering gold gradient (Art Deco background-clip: text). (2) Gold-bordered Card surfaces — settings cards get a subtle gold border with a gradient surface fill and inset gold top highlight. (3) Green light pulse — running agent status dots (.sidebar-dot-running) pulse with a soft green glow, echoing Gatsby's eternal hope at the end of Daisy's dock. (4) Gold terminal cursor — the xterm cursor is gold (#D4AF37) via the dedicated gatsby xterm palette, with a subtle gold border framing the terminal viewport. (5) Art Deco corner ornaments — modal dialogs get geometric stepped gold corner brackets (Chrysler Building motif) via ::before/::after pseudo-elements. (6) Sunburst status bar — the status bar gets a subtle radiating-line conic gradient background (the most iconic Art Deco motif). (7) Chevron zigzag dividers — <hr> elements render as repeating zigzag patterns in dim gold. (8) Gold-on-black scrollbars — scrollbars get a gold gradient thumb on a black track. (9) Art Deco input fields — text inputs (text, number, search, password, email, textarea, select) get a gold-underlined style with a gold glow on focus. (10) Green-light toggle glow — active toggle switches emit a soft green glow. (11) Gold focus rings — all :focus-visible outlines use gold. (12) Green-light text selection — text selection uses a translucent green-light background. (13) Gold-accented active tabs — the active settings tab gets a gold left border with an inset glow. (14) Gold range sliders — slider inputs use gold accent color. (15) Subtle gold button hover glow — buttons emit a faint gold halo on hover. All animations respect prefers-reduced-motion. The theme also includes a dedicated xterm terminal palette (gatsby) with matching gold cursor, green-light ANSI greens, and champagne foreground. The theme appears as the 11th option in the theme grid — a fun easter egg for users to discover.

  • STT settings overhaul — microphone selection, audio processing, transcript formatting, and live audio level. The Dictation settings tab now has detailed controls organized into sections: (1) Microphone device picker — enumerate available mics via enumerateDevices, select a preferred device, or use auto-selection (best laptop mic scoring). A Refresh button re-enumerates devices after plugging in a USB mic. (2) Audio processing toggles — independently enable/disable noise suppression, echo cancellation, and auto gain control (all default on). These are passed as getUserMedia constraints so the browser's DSP handles them before audio reaches whisper. (3) Transcript formatting — capitalize first letter (default on), auto-add trailing punctuation if the transcript lacks terminal punctuation (default on), and trailing space for multi-phrase dictation flow (default on). These are pure post-processing on the transcript text before injection. (4) Live audio level indicator — the VoiceStatusChip now shows a real-time mic level bar during recording, driven by an AnalyserNode on the mic stream (RMS level at ~60fps, respects prefers-reduced-motion). (5) Keyboard shortcuts reference card — shows "Hold Space = PTT", "Click Dictate = Toggle", "Esc = Cancel" directly in settings. All new settings persist across sessions and are pushed to the main process for diagnostics via VOICE_CONFIG_UPDATE IPC. The SttConfig type expanded from 5 to 12 fields; WhisperService.reconfigure() and getConfig() handle all new fields.

  • Robust agent callsign resolution (PATH-first with live refresh). NALA now reliably launches agents using their bare callsign (e.g., devin --permission-mode dangerous) instead of fragile full paths (e.g., & 'C:\Users\...\devin.exe' --permission-mode dangerous). Key improvements: (1) Live PATH refresh — before each agent launch, the daemon re-reads the system PATH from the Windows registry (HKCU + HKLM), so agents installed after the app launched are detected immediately without restarting. (2) Callsign disambiguation — Cursor and Grok previously both claimed agent as a PATH callsign, causing ambiguity; Cursor now only uses cursor-agent (its dedicated callsign), while Grok uses grok. (3) Detection scan RPC — a new nala.provider.detectAgents RPC returns a summary of which agents are installed, how they were found (PATH callsign vs known path vs not found), and their version. (4) resolvedVia + callsign fields in ProviderProbeResult — the UI can now show whether each agent was found on PATH (portable callsign) or via a known install location (full path fallback). The callsign-first approach makes launches portable across machines, survives binary updates/moves, and produces readable one-liners in the terminal.

  • TTS pause elimination (gap trim + pre-buffer). Read Aloud no longer has mid-sentence pauses. Two new settings (Settings → Read Aloud) fix the root cause: (1) Gap trim (0–200ms, default 40ms) removes the natural silence Kokoro inserts at the start/end of each sentence chunk, tightening the gap between consecutive chunks. (2) Pre-buffer (0–500ms, default 150ms) delays the first sound slightly to build a head-start buffer, so when a long sentence takes more time to synthesize than the previous sentence took to play, there's buffered audio to fill the gap — no silence. Both settings are adjustable via sliders and persist across sessions.

  • STT push-to-talk toggle + smoother state. The spacebar push-to-talk can now be toggled on/off in Settings → Dictation (default: on). When off, only the Dictate button works — the spacebar is never intercepted. The hold threshold slider now ranges from 15–150ms with 5ms steps for finer control. The default hold time is reduced from 45ms to 30ms (33% shorter) so a brief press is enough to trigger dictation. The voice status chip and toolbar button no longer flash "Dictate" before "Listening…" — both show "Listening…" instantly when you hold the spacebar, with no visible state transition (entrance animation and CSS transitions removed for instant feedback).

  • nala update CLI command — always run the latest stable. A new nala update (also wmux update) CLI command checks GitHub Releases for a newer stable version, downloads the installer, SHA-256-verifies it, and launches it — all without needing the app or daemon to be running. On Windows, the Squirrel Setup.exe handles kill-install-restart automatically. On Linux AppImage, the downloaded AppImage replaces the running one and relaunches; deb/rpm are opened with the system package manager. Use --check to only report availability without installing, and --json for machine-readable output. This means every time you run nala update, you end up on the latest stable build.

  • Launch-time update check (no more 30-minute polling). The in-app auto-updater previously checked for updates 15 seconds after launch and then every 30 minutes. The periodic 30-minute check has been removed — now the app checks once on fresh launch and auto-downloads the installer if a newer stable version exists. The user gets a "Restart to update" prompt when the download is verified. This aligns with the "always latest stable on launch" goal without being disruptive mid-session.

  • Dictation settings tab (Settings → Dictation). STT/voice dictation can now be configured from the UI instead of environment variables only. The new Settings → Dictation tab includes: whisper model selection (tiny.en / base.en / tiny / base / small, with size and speed notes), language selection (24 common languages + auto-detect), push-to-talk hold threshold slider (30–150ms), auto-preload toggle, a live engine status indicator (idle / loading / ready / error with model name), and reset-to-defaults. Model changes restart the whisper worker automatically; language changes take effect on the next dictation (no restart needed). Settings persist across sessions via session.json.

  • Instant push-to-talk (speculative capture). The mic now starts the instant you press Space — not after a hold timer. Previously, holding Space to dictate waited 100ms before even starting the microphone, then another 50-200ms for getUserMedia(), so the total press-to-dictate latency was 150-300ms. Now the mic starts on keydown, and the 60ms hold threshold only decides tap-vs-hold (if you tap, the speculative audio is discarded and a normal space is typed). A warm mic stream is also kept open between dictations to eliminate getUserMedia() latency on subsequent presses. Net effect: dictation starts capturing audio in ~0ms on warm presses (vs 150-300ms before).

  • Agent callsign launching (PATH-first). When you click a provider button (Claude, Codex, Devin, Cursor, Grok), NALA now launches the agent using its bare callsign (e.g., devin --permission-mode dangerous) instead of the full file path (& 'C:\Users\...\devin.exe' --permission-mode dangerous). This works because most agent CLIs install themselves on PATH. If the callsign isn't on PATH, NALA falls back to the known install location. The callsign approach is more portable (survives binary updates/moves), more readable in the terminal, and matches how users naturally launch agents. Custom providers also get callsign support via their executableNames field.

  • Dark window frame with white title bar symbols. The main window now uses titleBarOverlay with a dark background (#050C0E) and white min/max/close symbols, so the window frame matches the app theme instead of using the Windows system accent color (which could be orange, blue, or any color).

  • Release notes in Settings → About. The update widget in the About tab now shows a collapsible "Release notes" panel with the CHANGELOG excerpt for the available version, so users can read what changed before clicking "Update & Restart".

  • Full i18n coverage for update UI. All 23 locales now have translations for the StatusBar update keys (checkForUpdates, checking, settings, about, restartToUpdate, updateReady, upToDate) — no more English fallback for any locale.

  • Topbar "Restart to update" button. When an update is downloaded and SHA-256 verified, a pulsing green "↻ Restart to update" button appears in the status bar next to the settings gear — one click installs and restarts. No need to navigate to Settings.

  • Settings dropdown with "Check for updates". The settings gear (⚙) in the topbar now opens a dropdown menu with "Check for updates", "Settings", and "About" — so manual update checks are one click away if auto-detect fails on startup.

  • Agentic tools settings + MCP tools. Agents can now call 6 new MCP tools, each gated by a Settings → Tools toggle. Safe tools (on by default): agent_list_workspace (see agents in workspace), agent_prompt (send a prompt to another agent). Dangerous tools (off by default): agent_spin_up (spin up a new agent in a terminal), agent_open_window (open a new workspace window), agent_open_terminal (open a new terminal), agent_spawn_subagent (spawn a subagent for delegated work). When a tool is disabled, agents get a clear "enable in Settings → Tools" message. Settings persist across sessions with fail-closed guards on dangerous tools.

  • Quiet toasts by default. Most toast notifications are now suppressed — the UI already surfaces those signals inline (badges, panels, status), so the pop-ups were just noise. Only actionable error-level toasts appear by default; routine info/warn toasts ("copied!", "workspace created", "long paste truncated") are dropped at the store boundary. Settings → Notifications adds a "Toast minimum level" selector (Error / Warning / All) to dial verbosity back up, and the existing toast on/off toggle now fully silences them. The setting persists across sessions.

  • Read Aloud: paste text to hear it spoken. A new card at the bottom of the workspace turns a clipboard paste into spoken audio using Kokoro-82M, a fully local, open-source TTS model — no cloud calls, no Python, no GPU required. Click the card, paste (Ctrl+V or any native paste gesture), and the text is read aloud sentence-by-sentence via Web Audio, starting in well under a second on subsequent uses (streaming synthesis — the first sentence plays while later ones are still being generated). The pasted text is never rendered or stored in DOM state. A new Settings → Read Aloud tab offers voice selection (15 curated English voices across US/UK accents, grouped by accent and gender with quality grades and trait descriptions), a voice preview button to audition any voice without pasting, speed (0.75×–1.5×, Kokoro's duration-aware speed so it stays natural, with quick-select preset buttons), a fast/higher-quality toggle, volume, auto-play-on-paste, interrupt-vs-finish policy for pasting while already speaking, and a reset-to-defaults button. First use requires an internet connection to download the ~90 MB voice model; after that, it works fully offline. The model runs in a Node.js worker_thread inside the Electron main process so synthesis never stalls the UI. Pasted text is normalized (markdown fencing, headings, links, bold/italic, list markers, emoji, and smart quotes are stripped) so docs and chat read naturally. The worker auto-terminates after 5 minutes of idle to free memory. When auto-play is off, pasting "primes" the card (green indicator) and a click starts speaking. A retry button appears on error. Download progress is shown as a percentage and inline progress bar during first-run model fetch.

  • Agent callsigns — same-provider instances get distinct names. Launching three Devin panes used to show "Devin CLI" on every tab, sidebar row, and Fleet card with no way to tell them apart. Each agent launch now auto-assigns a short, futuristic callsign (Vega, Cipher, Onyx, …) from a ~50-entry pool, scoped per-workspace so the same name never recurs twice in one workspace at the same time. The callsign seeds Surface.title at launch, so tabs, sidebar, Fleet cards, and nala_agent_list all show it automatically. a2a_discover returns it as agentName (callsign-first) alongside a separate agentType (the detected CLI flavor), so an orchestrating agent can address "Agent Vega" by matching the callsign instead of guessing which "Devin CLI" to send to. FleetCard's display-name priority was reordered to paneLabel || title || agentName so an explicit callsign wins over the generic detected type. Manual rename (double-click a tab) still works and is never overwritten — callsigns only seed new launches. Right-click a tab for a "Reroll callsign" action that picks a fresh random name from the pool (excluding other tabs' current names). The a2a_discover and nala_agent_list MCP tool descriptions now explicitly mention the callsign field so orchestrating agents know to match it when a user says "tell Vega to…".

  • Agent detection for Devin CLI, Cursor CLI, and Grok Build. The agent detector now recognizes three additional CLI tools — Devin CLI (Cognition), Cursor CLI (cursor-agent), and Grok Build (xAI) — so their panes get the correct agent slug, status tracking (waiting-at-prompt detection), and signal routing. The AgentSlug union was extended in all three sync locations (AgentDetector, shared/events, integrations/shared/signal-types) per the existing contract.

  • Create a workspace from a folder. wmux new-workspace --cwd <dir> (also --dir/--path, or a bare directory path) opens a workspace whose terminals start in that folder and is named after the folder by default. In the UI: sidebar + → Open folder…, or the command palette Open Folder as Workspace.

  • Workspace isolation for agent-to-agent messaging. Agents launched inside a workspace can now only discover, message, and broadcast to agents in that same workspace — run separate multi-agent fleets on multiple projects with no bleed-through. Sibling agents stay fully addressable (discover lists every agent pane in your workspace; broadcast pings each sibling pane), replies to human-bridged cross-workspace tasks still work, and your own tools (UI, CLI) keep full cross-workspace reach.

  • Fan-out sub-tasks know their slice. Each spawned agent's prompt now opens with its task name, position ("sub-task 2 of 4"), the sibling task list, an only-your-slice scope rule, and a pointer to the mission channel — instead of N agents all receiving the identical prompt and colliding on the same work.

  • True YOLO. With YOLO on, NALA never interrupts you — agent-to-agent execute requests and spawn proposals approve automatically, even in Connect/Orchestrate modes. The only "agent needs you" moment left is an agent's own CLI asking a question in its terminal, which shows as that tab's figure-8 freezing. Settings → NALA now explains YOLO and each A2A mode in plain language.

  • Per-provider status-bar visibility. Settings → NALA → Providers has show/hide checkboxes for each quick-launch button (all shown by default); the row re-spaces itself automatically. Buttons are now square, logo-only marks — no more Cl/Cx/Dv/Cu/Gk letters.

  • Smart provider launch. Clicking a provider button with an empty, untouched terminal focused opens the agent right there — seamlessly, in the same tab position, no flash — instead of adding a new tab; any activity keeps the new-tab behavior, and clicking several provider buttons in a row spawns one terminal each.

  • Tab figure-8 is the agent status light. A tab hosting an agent always shows the mini figure-8: moving while the agent works, frozen and dimmed when it's idle, finished, or waiting on your input. The bell icon is gone. The motion tracks real activity — an agent parked at its prompt (or a trust dialog) freezes within ~3 seconds even if its coarse status still says running.

  • Close every terminal. Closing the last terminal no longer auto-respawns one — the pane shows a calm empty state (still figure-8, "All quiet here", a New Terminal button) until you start something. Restarts still boot with a fresh terminal.

  • Dictate into any text box. Hold Space in any input, textarea, or editable field — modals, popovers, Idea Capture, Rich Input — and speech streams in at the caret as you talk, same near-live cadence as terminal dictation. A short Space press still types a space. The target is pinned when you start talking, so slow transcription can't land in the wrong place.

  • Bottom-toolbar tools always work. Attach, Snippets, Rich Input, and Dictate are never greyed out by focus quirks — each resolves its terminal via a fallback chain (focused → active pane → any), and creates one if the workspace has none.

  • No console window in the taskbar. Launch via scripts/NALA.vbs (point the desktop shortcut there) — the launcher console runs hidden; output stays in logs/launch.log. The daemon was already windowless.

  • Tasks & Artifacts and Plan are now in the sidebar. The two actions that lived in a dedicated strip under the status bar moved into the sidebar (both the expanded list and the collapsed icon rail), so they're always reachable regardless of sidebar position. The panel and sheet themselves are unchanged.

Fixed

  • TTS pre-buffer sample-rate miscalculation. The pre-buffer threshold was calculated using a hardcoded 24000Hz sample rate. If a TTS chunk arrived at a different sample rate (e.g. 22050Hz), the threshold was miscalculated — causing premature or delayed playback start. Replaced pendingSamples with pendingDurationMs, computed from each chunk's actual sample rate.
  • WhisperService config reported stale PTT hold time. getConfig() hardcoded pttHoldMs: 60 instead of using the VOICE_PTT_HOLD_MS constant (now 45ms). The diagnostics config now reflects the actual default.
  • CLI update: uncaught renameSync on AppImage backup. The renameSync that backs up the existing AppImage before replacing it was not wrapped in try/catch — an ENOENT or EACCES would crash with an opaque error. Now wrapped with a clear error message.
  • CLI update: spawn errors silently swallowed. The Windows installer, AppImage relaunch, and xdg-open spawn calls had no error handlers — if the child process failed to start, the CLI would exit silently. All three now have .on('error') handlers that report the failure in both human and JSON mode.
  • CLI update: write-stream errors in downloadAndVerify. Disk-full, permission, or I/O errors on the download write-stream were not caught — the download would hang indefinitely. The write-stream now has an error handler that fails-closed.
  • CLI update: version resolution in bundled CLI. getFallbackVersion() only read package.json relative to __dirname, which fails in a Squirrel install. Now tries the running app via system.identify RPC first, then falls back to package.json, then 0.0.0 (any published release will be "newer").
  • ChannelsPanel unhandled promise rejections. joinChannelDaemon() and hydrateChannelsCatalog() promise chains had no .catch() — a daemon RPC rejection (not just ok:false) would be an unhandled promise rejection. Both now have .catch().
  • WhisperService non-null assertion crash. workerScriptPath() used candidates[0]! which would crash if all candidates were empty strings. Now uses candidates[0] ?? '' — the caller already checks fs.existsSync() and reports a clear error.
  • TTS config typecheck error. useReadAloud.ts was missing gapTrimMs and preBufferMs fields when constructing a TtsConfig object, causing a TypeScript build failure. The fields are now read from the zustand store and included in the config sync effect's dependency array so changes to either setting propagate to the TTS worker.
  • TTS audio playback test failures. Six audioPlayback.test.ts tests were failing because the default preBufferMs=150 caused small test buffers to be queued instead of scheduled (making isPlaying() return false), and zero-filled buffers were stripped entirely by trimSilence. Tests now use non-silent test buffers and disable pre-buffering via setPreBufferMs(0) to test the scheduling logic directly.
  • TTS config test updated for new fields. ttsConfig.test.ts expected TTS_DEFAULT_CONFIG without gapTrimMs/preBufferMs; updated to include the new default values (40ms / 150ms).
  • Auto-updater 404 on Windows. The updater previously fetched from update.electronjs.org, which returned 404 for private repos and required an external service. Now fetches update-manifest.json directly from GitHub Releases (same self-describing approach as Linux) — no external dependency, no 404.
  • Update check 404 shows "Up to date" instead of "Update failed". When no GitHub Release has been published yet (or the release has no manifest asset), the manifest URL returns 404. This is now treated as "no update available" (the renderer shows "Up to date") instead of surfacing as an error. Only real network failures or server errors (5xx) show as "Update failed".
  • Window frame uses Windows system accent color. The main window title bar and border previously used the Windows system accent color (which could be orange, blue, etc). Now uses titleBarOverlay with a dark background (#050C0E) and white symbols to match the app theme.
  • Old nala-icon fallback removed. The icon resolution fallback chain in createWindow.ts, tray.ts, and index.ts no longer falls back to nala-icon.ico/nala-icon.png — only the standard icon.* files are used, ensuring the white-on-black NALA icon is always shown.
  • White accent color. The teal/cyan accent (#62C8D4) is now white (#FFFFFF) across the default Nala Stream theme — focus rings, figure-8 dots, active tabs, beam glows, and all --accent-blue references now match the white "N" branding.
  • Squirrel installer branding. Added a dark-background loading GIF to the Windows installer so the setup screen matches the app's dark theme instead of showing the default Squirrel appearance.
  • Faster dictation start. Reduced the push-to-talk hold timer from 150ms to 100ms — press-and-hold Space now arms dictation 50ms sooner with minimal false-trigger risk.
  • Release notes in update manifest. Both Windows and Linux update manifests now include releaseNotes (extracted from CHANGELOG by CI), and the UpdateManifest / LinuxUpdateManifest TypeScript interfaces include the field for type safety. The Settings → About update widget captures and displays these notes in a collapsible panel.
  • Agent callsign pool refactored. Removed "gus" and all 3-letter names from the callsign pool. All 56 callsigns are now 4-5 letters (a few 6-letter exceptions like "Signal", "Vector", "Beacon") for better readability in @-mentions and tab labels.
  • Figure-8 tab animation robustness. Added a 5-second startup grace period so the animation doesn't pause prematurely while an agent is booting up (showing a banner, waiting on a trust dialog). Faster 500ms re-sample tick (was 1s) for more responsive state transitions. New first-output tracking in the ptyActivity module ensures the animation reliably starts when an agent begins working and stops when idle, for all agents and terminals.
  • Chord-click paste in Read Aloud. Left+right mouse buttons pressed simultaneously now pastes clipboard content into the TTS textarea (same behavior as the terminal). Right-click context menu suppressed on the textarea.
  • TTS pipeline: double terminal chunk on error. The TTS_SPEAK handler sent a done:true, error chunk from .catch() but didn't set terminalSent = true, so .finally() sent a second redundant done:true chunk. Now the catch path sets the flag so the finally is a no-op.
  • TTS pipeline: stale error clobbering newer request. In TtsService.onWorkerMessage for type === 'error', the error status was set unconditionally — even when the error's requestId belonged to an already-superseded request. A stale error could flip a newer request's status to error. Now only errors for the pending request (or genuine worker-level errors with no pending request) surface to the status handler.
  • Unhandled promise rejection in useNotificationListener. The resolveAgent IPC call (during daemon handler-swap windows) had no .catch(), causing an unhandled promise rejection. Added the same best-effort catch used in AppLayout.tsx for the identical call.
  • Flaky diffParse test under parallel load. The round-trip oracle test creates temp git repos via execFileSync('git', ...) which intermittently fails when 5960+ tests compete for temp directory I/O and process spawn slots. Added a retry wrapper (2 retries with 100-200ms backoff) and a 15s timeout to the git helper so transient git failures don't cause test failures.
  • i18n: last hardcoded string (SnippetsMenu empty state). The snippets menu empty state text ("Nothing saved yet. Add a snippet below — then one click pastes it into the terminal.") was the last remaining hardcoded English string in the renderer. Now uses t('toolbar.snippetsEmpty').
  • SearchBar race condition with split panes. The terminal search history was stored in a module-level mutable array shared across all SearchBar instances. In split panes, multiple SearchBars would race on the same array, corrupting each other's history. Moved to a per-instance useRef so each SearchBar has its own independent history.
  • i18n: remaining hardcoded strings (round 3). 11 user-visible strings were still hardcoded English: DiffPanel header ("Diff") and loading state ("Loading…"), PresetPicker empty option ("Empty" / "Blank single pane"), SettingsPanel keybinding capture ("ESC to cancel"), NalaTaskPanel artifact preview close button ("Close"), and OrchestrationDashboard (7 strings: title, description, buttons, status labels). All now use t() with 11 new i18n keys.
  • Electron Forge packaging config for NALA branding. The forge.config.ts packagerConfig was missing name, executableName, appBundleId, and win32metadata — the packaged executable used the internal "wmux" name instead of "NALA", and Windows metadata (Company, Product, Description) was unset. Added name: 'NALA', executableName: 'NALA', appBundleId: 'com.nala.app', and full win32metadata (CompanyName, ProductName, FileDescription, OriginalFilename) so the installer and taskbar show "NALA" correctly.
  • Removed debug console.log from production main process. src/main/index.ts had a console.log('[DEBUG] registering app.on(ready)') that printed to the console on every startup. Removed.
  • FirstRunWizard layout rollback on dismiss. Dismissing the first-run wizard (Skip / Escape / ×) while the sample task was mid-flight (splitting or awaiting-prompt) left the user with a 2x2 grid layout they didn't ask for — the previous layout was never restored. The wizard now saves the workspace's rootPane, activePaneId, and nextPaneOrdinal before applying the grid template, and restores them on dismiss if the sample task hasn't completed. The saved layout is cleared on success so no rollback occurs after a successful sample task.
  • i18n: SettingsPanel NALA tab hardcoded strings. The YOLO mode label/description, A2A mode label/description/options (Off/Connect/Orchestrate), providers visibility description, section labels (Orchestration, Providers, Theme), and StatusBadge labels (registered/not registered, detected/not detected, OK/Not OK) were all hardcoded English. All now use t() with 18 new i18n keys.
  • i18n: ChannelView and Composer fallback strings. Removed unnecessary || 'Search messages…' and || 'Type a message…' fallbacks — t() always returns a string, so the fallbacks were dead code that bypassed i18n.
  • i18n: Toast and Dialog aria-labels. "Dismiss", "Status messages", and "Close" aria-labels in ToastContainer and DialogHeader were hardcoded English. Both components now import useT and use t('toast.dismiss'), t('toast.statusMessages'), and t('dialog.close').
  • Removed outdated TODO comments. Two "M6 TODOS #4" comments in CompanyPanel.tsx and CompanyView.tsx described a fix that was already implemented (await destroyCompanyWithCleanup()). The stale TODOs were removed to avoid confusion.
  • Accessibility audit: keyboard navigation and screen-reader labels. A full a11y audit found 21 issues across the renderer. Fixed: 6 clickable <div>/<span> elements that lacked role and tabIndex (multiview workspace tiles, department headers, PR badges, mission rows, search results) now have role="button"/role="link", tabIndex={0}, and onKeyDown handlers for Enter/Space; 13 <input> elements without labels (browser URL bar, command palette search, terminal search, pane/surface rename, settings keybinding/color/agent-command inputs, workspace profile env-var inputs) now have aria-label attributes. 9 new i18n keys added for the aria-labels.
  • AudioContext leak in notification sound. useNotificationSound.ts created a module-level AudioContext singleton that was never closed. Added disposeNotificationSound() export so the context can be explicitly closed to release audio resources.
  • i18n sweep: hardcoded English strings → translatable keys. A locale audit found ~70 user-visible strings still hardcoded in English across 6 components — ErrorBoundary crash fallback, EditorPanel (read error, reload, view/edit toggle, loading), FileTreePanel (no-CWD, refresh, empty, close preview, read error, loading), Pane empty state ("All quiet here", "New Terminal", launch-agent hint), SettingsPanel MCP servers section (label, checking, re-register, confirm, cancel, unregister, status unavailable), and NalaTaskPanel (30+ strings: toasts, tab labels, section headers, send/reply buttons, inbox, search, artifact preview). All now use t() with new en/ko locale keys. Two variable-shadowing bugs in NalaTaskPanel (local const t = shadowing the translator t) were also fixed.
  • Unified all loading states on the NALA figure-8 loader. A comprehensive audit found 29 distinct loading state patterns across the app — 3 different spinner types (FigureEightLoader, CSS animate-spin, text-only "Loading..."), plus 8 buttons that just disabled with no visual indicator. All loading states now use the same FigureEightLoader component with appropriate size variants:
    • NalaOrchestrationBar provider buttons: CSS spinner → FigureEightLoader size="inline"
    • FileTreePanel: text-only "Loading..." → FigureEightLoader size="inline" + text
    • EditorPanel: text-only "Loading..." → FigureEightLoader size="pane" + text (centered)
    • DiffPanel: text-only "Loading..." → FigureEightLoader size="pane" + text (centered)
    • SettingsPanel (6 locations): MCP checking/reregister/unregister, LanLink loading/applying, update check, TTS voice preview, LanLink pairing begin/join — all now show FigureEightLoader size="inline" instead of just disabling or showing "…"
    • NalaTaskPanel "+ New" button: "…" text → FigureEightLoader size="inline"
    • FirstRunWizard register/retry buttons: disabled-only → FigureEightLoader size="inline"
    • WorktaskCleanupView rescan/close buttons: text-only → FigureEightLoader size="inline" + text
    • The browser refresh icon's animate-spin is intentionally kept (standard browser convention).
  • Product-name "wmux" → "NALA" rebrand in i18n locales. 14 user-facing strings across English and Korean still said "wmux" in product-name contexts (supervised restart markers, LanLink daemon notice, updates label, taskbar flash description, MCP registration status, first-run wizard errors, Claude integration signal health, channels stale banner, resource exhausted error). All now say "NALA". Technical references to wmux.json (the config file) and wmux mcp register (the CLI command) are intentionally kept as-is.
  • Voice dictation UI timeout was 122s for every clip, even 1-second. voiceTranscribeUiTimeoutMs used Math.max(VOICE_TRANSCRIBE_MAX_MS, ...) instead of just the duration-scaled timeout + slack. A 1s clip should time out in ~22s (20s main + 2s slack), but the UI waited over 2 minutes before reporting a timeout. Now the UI timeout is voiceTranscribeTimeoutMs(duration) + 2s, so short clips fail fast and the user gets feedback in seconds, not minutes.
  • Voice dictation error messages now distinguish engine failure from timeout. "Worker not running" (Python missing/crashed) now says "Voice engine unavailable — check Python + faster-whisper installation" instead of the generic "Dictation timed out — try a shorter phrase". Engine startup timeout also gets a specific message.
  • Read Aloud has a visible text input box in the sidebar. The old paste-only tile (click → paste → hear) was replaced with a visible text input where you can type or paste text and press Enter (or click the speaker button) to hear it spoken. The input shows the speaker icon, a text field, and a speak/stop/retry button — all in one compact row at the bottom of the sidebar, below Tasks & Artifacts and Plan.
  • Read Aloud error state no longer silently reverts to idle. When TTS synthesis failed (watchdog timeout, worker crash, model error), the error status arrived via TTS_STATUS(error) but was immediately clobbered back to idle by the late terminal done=true audio chunk (sent by the handler's finally block). Both useReadAloud and useTtsPreview now clear the requestId in the error handler so the stale done chunk is ignored, leaving the error state visible to the user.
  • Read Aloud cancels in-flight synthesis on unmount. Both useReadAloud and useTtsPreview now call api.stop(requestId) in their cleanup effects, so closing the Settings panel mid-preview or unmounting a pane stops the worker instead of letting it synthesize audio that nobody will hear.
  • Voice preview button shows error state. The Settings → Read Aloud voice preview button now turns red and shows the error message when preview fails, instead of silently resetting to idle.
  • AUMID icon fallback consistent between prestart script and main process. The main process now falls back to assets/nala-icon.ico when assets/icon.ico is missing, matching the prestart script's fallback logic.
  • 12 lint errors fixed (0 errors remaining). Fixed: prefer-const in test files, irregular whitespace in JSDoc, empty interface declarations (→ type aliases), {} type in i18n, control character in regex, and stale react-hooks/exhaustive-deps disable comments referencing an uninstalled plugin rule.
  • SharedArrayBuffer ReferenceError in sandboxed preload. voice.transcribe() failed instantly on every call because instanceof SharedArrayBuffer throws ReferenceError in the sandboxed preload context (the global isn't defined there). Now guarded with typeof SharedArrayBuffer !== 'undefined'.
  • Read Aloud moved from a floating bottom pill into the left nav bar, grouped with Tasks/Plan at the bottom. The paste-to-speech control was a separate pill floating over the terminal area; it's now a square icon tile (32px) matching its neighboring nav icons. Tasks & Artifacts, Plan, and Read Aloud are grouped together in their own section at the bottom of the sidebar (both expanded and collapsed/MiniSidebar rail), just above the workspace-count footer, leaving the top of the nav (header + list) exclusively for workspaces. Same click behavior (paste to speak, click again to stop) and the same underlying useReadAloud() pipeline, just relocated and regrouped.
  • "Resume Claude" pill redesigned and auto-dismisses. The recovered-agent resume pill was a raw inline-styled box floating at the top-left of the pane, overlapping the tab strip, with a clunky "▶ Resume Claude ×" layout that never went away on its own. It's now a polished, bottom-left floating pill with a proper CSS class, fade-in animation, a play icon + "Resume Claude" action button, and a subtle dismiss × button. It auto-dismisses after 60 seconds (the offer is stale by then), on user input (typing into the pane), on explicit dismiss, or when the agent is re-detected live — so it's gone when not needed.
  • No more "Unsupported terminal" warning from Devin CLI and other modern CLI tools. Every terminal pane now exports TERM_PROGRAM=wmux and TERM_PROGRAM_VERSION, the industry-standard variables that CLI tools check to detect a capable terminal emulator. Without them, tools fell back to Windows Console Host (conhost) detection and warned about limited support. Stamped at both spawn sites (local and daemon mode), after all env filtering, so a workspace profile or replayed session can never override the terminal identity.
  • Status-bar A2A / YOLO / provider buttons (Cl, Cx, Dv, Cu, Gk) no longer toast "Unknown method". The daemon build was failing on NALA TypeScript errors (redactionVersion on activity append + a ProviderIdSchema re-export clash), so nala.provider.* never registered and every click failed. Activity append now defaults redactionVersion, first-class provider ids are namespaced separately from agent provider ids, and a rebuilt daemon exposes the RPCs again.
  • Voice dictation writes into the terminal while you speak. Capture flushes speech segments about every 0.7s, transcribes them, and injects text in order (direct PTY write for live partials) instead of waiting until you stop and dumping the full clip at once.
  • useVoiceController requires a parent <VoiceHost> no longer crashes the app. HMR / remount races could leave VoiceButton or VoiceStatusChip without a provider; the hook now falls back to a disabled no-op controller instead of throwing.
  • New terminal inherits the active terminal's working directory. Toolbar "New", the pane tab-strip "+", Ctrl+T, the command palette's New Surface, and provider quick-launch all seed from the focused (or owning-pane) terminal's live cwd when available, so a new tab opens in the same path you are already in.
  • NALA A2A MCP tools work under permission enforcement. The eight agent-facing nala.* RPCs the bundled MCP server calls (task get, agent list, message send/reply/inbox, spawn propose, artifact/claim list) were missing from the method-capability map and the first-party grant, so nala_* tools were denied in enforce mode. They now carry a2a.read/a2a.send capabilities and are first-party granted; the rest of the nala.* surface stays wmux.internal.
  • Hold-Space dictation no longer cuts off at 45 seconds. The hard recording cap was a per-buffer limit that ended the session when hit — even with the key still held and the user still talking. Hitting the cap now seamlessly rolls into a fresh capture segment: the current buffer is flushed to whisper in the background and capture continues with no perceptible gap, for as long as the key stays down. An ordered, serial queue guarantees injected text stays chronological even if an earlier segment's transcription finishes after a later one's. The silent-drop case (releasing the key after a rollover already finalized) now surfaces a subtle toast instead of the status chip vanishing without explanation.
  • Rollover boundary no longer drops audio with the live-partials flag. rolloverSegment now flushes with minSec: 0 so the pending buffer is always cleared at the 45s cap, even when a recent live flush left less than 0.7s buffered. Without this, flushSegment would return null without clearing pending, while the capture loop reset its sample counter — corrupting subsequent segment tracking.
  • A NALA toolbar crash no longer blanks the whole app. The NalaOrchestrationBar is now wrapped in an ErrorBoundary with a compact fallback — if the A2A bar's render throws, a small "NALA ↻" retry button appears in its place instead of a full-screen red crash panel covering the status bar.
  • YOLO and A2A mode no longer "revert by themselves". The Settings → NALA toggles for YOLO and A2A mode were backed by independent local state, so a 15-second status-bar poll could overwrite a just-changed setting before the RPC round-trip completed. Both are now backed by store fields (nalaYoloActive, nalaA2aMode) so the settings panel and the toolbar control always agree instantly.
  • Company destroy no longer races the store flip. destroyCompanyWithCleanup was fire-and-forget — the store flipped to company === null before async pty.dispose handlers finished, which could dereference the just-cleared store. The handler now awaits the cleanup Promise.all before the store mutation.
  • Sandboxed preload no longer blank-screens on boot. getPipeName() in shared/constants.ts had a top-level require('os') that Electron's sandboxed preload require shim cannot resolve — an eager module-scope import broke preload loading entirely (silent blank-screen boot). The require is now confined to the function body, which only executes in main/daemon contexts.
  • Voice transcription no longer hangs on detached ArrayBuffer views. The preload transcribe() IPC call could hang or throw when passed a detached ArrayBuffer or a SharedArrayBuffer view — structured-clone across IPC can't handle either. The preload now clones the audio into a standalone Uint8Array before invoking, so the renderer's audio pipeline never trips the edge case.
  • Read Aloud: taskbar shows the NALA icon instead of the Electron atom in dev mode. The icon-stamping script was only called from launch-nala.cmd, not from npm start / start:fast, so running the app from the command line showed the stock Electron atom in the Windows taskbar. A new cross-platform prestart script (scripts/prestart-stamp-icon.js) stamps electron.exe with the NALA icon via rcedit AND registers the AppUserModelID with the icon in the Windows registry, so the taskbar always resolves the correct icon. The main process also registers the AUMID at boot as belt-and-suspenders.
  • Voice dictation failed instantly on every clip, of any length ("Dictation timed out" on the very first word). The SharedArrayBuffer defensive check added to preload's voice.transcribe() (see the entry above) referenced the bare SharedArrayBuffer global directly — but it isn't defined at all in Electron's sandboxed preload context, so audio.buffer instanceof SharedArrayBuffer threw a ReferenceError synchronously on every single call, before the audio ever reached the whisper worker. Because that throw happened while building the argument to the renderer's withTimeout(...) wrapper, it landed in the same catch block as a real timeout and surfaced as "Dictation timed out — try a shorter phrase" regardless of clip length — making a 100%-reproducible crash look like an intermittent timeout. Fixed by guarding the reference with typeof SharedArrayBuffer !== 'undefined' first. Verified live: both a 3s and a 50s synthetic clip now transcribe successfully end-to-end (whisper worker → main → preload → renderer).
  • Read Aloud: robustness fixes (dtype race, worker exit hang, error correlation, memory leaks, nested markdown). A final audit pass fixed: (1) dtype race in loadModel that captured the wrong dtype on ready; (2) loadModel hanging for 120s if the worker was terminated during load (missing exit listener); (3) setConfig dtype invalidation only firing on ready state, missing downloading/loading; (4) intentional worker termination clobbering the idle status with an error; (5) ReadAloudCard missing the speaking CSS class (dead selector); (6) handlePaste not always clearing the hidden textarea; (7) worker error messages missing requestId so the service couldn't correlate errors to the pending request; (8) handleLoad not clearing the previous model before loading a new one (ONNX model leak); (9) playChunk with a zero-length buffer creating a source that never fires onended (stuck isPlaying); (10) nested bold/italic markdown (**_text_**, ***triple***) not fully stripped; (11) ensureReady not verifying loadedDtype matches config.dtype; (12) "stopped" rejection from intentional cancellation showing as an error on the Read Aloud card (now returns to idle gracefully).
  • Daemon interval timer leak on shutdown. Five periodic setInterval handles (A2A GC, WorkTask GC, resume spool drain, dead session reaper, buffer snapshot) were never cleared in shutdown(). All five used .unref() so they didn't block process exit, but they could fire mid-teardown and race with session disposal and state saves. The handles are now module-level and cleared at the start of shutdown().
  • Deprecated wmic.exe replaced with PowerShell Get-CimInstance. The daemon used wmic.exe in three places (boot ID retrieval, sync boot ID, shell process verification). wmic.exe is deprecated on Windows 10/11 and may be removed in future updates. All three now use PowerShell Get-CimInstance via the shared windowsPowerShell51Path() helper.
  • Unhandled promise rejections in AppLayout. Two fire-and-forget promises (daemon.whenReady() and probeProjectConfig()) had no .catch() handler, causing unhandled rejections in the renderer. Both now have best-effort catch handlers.
  • Silent error swallowing in daemon event switch and clipboard write. The daemon event switch's default case silently dropped unknown event types — now logs a warning. The browser panel's clipboard write silently swallowed errors — now logs to console.warn.
  • Duplicate Windows PowerShell path resolution consolidated. portWatch.ts had its own windowsPowershellPath() function duplicating windowsPowerShell51Path() from shared/shellResolution.ts. Now imports the shared helper instead.

Changed

  • WhisperService refactored for speed and reliability. Defaults are now sourced from shared/voiceConfig (tiny.en model, int8 compute, en language) instead of hardcoded constants, with env overrides for model, language, threads, and timeouts. CPU thread alignment between BLAS env vars (OMP/MKL/OPENBLAS) and ctranslate2's cpu_threads was fixed — previously BLAS was hardcoded to 4 while Python capped at 6, under-using available cores on higher-core machines during the latency-critical transcription path. Hard timeouts with stuck-worker kill prevent UI hangs. Tmp file pruning reduced from 1h to 20min.

  • Provider quick-launch buttons show real brand marks (Claude, OpenAI, Devin, Cursor, Grok — path data from lobe-icons / Simple Icons) next to the short labels.

  • Bottom-toolbar popovers and NALA sheets rebuilt on the shared dialog design — Rich Input, Snippets, File Explorer, Idea Capture, and the Tasks panel now use the ui primitives: 12px-radius surfaces, larger inputs, Ctrl+Enter submits, theme-token backdrops, one primary action per flow.

  • Bottom toolbar slimmed to icon-only chips. The agent toolbar now matches the status bar's compact density — square icon-only buttons (h-6, rounded 6px) with labels demoted to tooltips, instead of full icon+label buttons. The separate NALA chrome strip that sat between the status bar and the terminal content is gone; its Tasks & Artifacts and Plan triggers moved into the sidebar. The Dictate button was slimmed to match — icon-only with the label and Space hint in its tooltip.

v3.20.0

Added

  • Experimental: hidden panes can skip output parsing (Settings → "Skip hidden pane rendering"). Even with the shared output scheduler, hidden agents' output was still parsed eventually — and measurement showed that parsing total is what drags the visible pane once several background agents stream at once (4 hidden flooders pulled the visible pane down to ~10–20fps). With this toggle on (daemon sessions only, default off), hidden panes' output is queued but never parsed: the renderer does no parsing work for panes you aren't looking at. A pane whose backlog outgrows its cap is marked stale and transparently re-synchronized from the daemon's session buffer when revealed — the daemon replays the authoritative bytes onto a reset terminal, so what you see on reveal is the pane's true current state, never a duplicate or a half-parsed frame. Agent-facing buffer reads (wmux_search_panes, terminal_read) hydrate a stale pane before reading so orchestrating agents never see old output. If a re-sync can't complete (dead session, legacy daemon), the pane degrades to its last-known screen instead of sticking or losing its identity.

  • Diff comments now wake the task agent (J4). Commenting on a hunk in a fan-out task's diff surface no longer just records a note — it @-mentions the task's agents on the mission-channel post, so the existing mention→wake loop nudges them to read and act on the feedback. Every non-human member of the mission channel (excluding you, the commenter) is mentioned at the workspace level, so multiple agent panes sharing one workspace all get woken; if every agent has left the channel the comment still posts, just without a mention. The post's body also carries a [diff: <file> @ <hunk>] <comment> prefix so an agent reading the channel over the CLI or MCP (which don't render the structured anchor) still sees which file and hunk the comment is about. The success message reports how many agents were pinged.

  • Fleet cards surface an agent's completion evidence. A fleet card now shows a small ✓ evidence n/m badge when the pane's most recently completed A2A task carries structured completion evidence — n is how many of the m evidence items are actually verified (a passed command, or a verified inspection/artifact). It's the "trust it ran unattended" proof made legible on the card: the check reads green once at least one item is verified and stays muted when nothing is (verified is a grade, not a claim), and the task title plus the evidence summary live in the badge's tooltip so the on-card text stays a single compact token. The badge reads existing task state only (no new store or round-trip), is addressed per-pane (a pane-pinned task shows on exactly that pane; a workspace-level task shows on the workspace's active pane), and simply isn't drawn when there's no such task.

Fixed

  • Multiple workspaces full of busy agents no longer stutter the visible terminal. Every pane used to push its PTY output straight into its own terminal the moment it arrived over IPC — including panes in hidden workspaces — so a fleet of background agents ran that many independent parse/render pipelines on the one renderer thread, and the pane you were actually typing into starved between them. Terminal output now flows through a single shared scheduler: the visible pane keeps the exact direct-write path it always had for ordinary output (zero added latency), while hidden panes' output is batched and drained cooperatively under a hard per-tick time budget, so no amount of background agent chatter can pin the UI. Even the visible pane's own output floods are chunked through that budget rather than parsed in one blocking pass, so watching a chatty agent stays responsive too. Nothing is dropped — a hidden pane's backlog is handed over in full when it becomes visible (before its reveal repaint), when a reconnect replay needs it, or if it ever exceeds the scheduler's memory cap (which simply restores the old behavior for that pane).

  • Diff-panel comments now actually post to the mission channel. The diff comment post omitted the sender identity the daemon requires, so every comment was rejected with a "코멘트 발사 실패" authorization error instead of being recorded. The comment now posts as the diff's owner workspace (its own mission-channel member row), which is also what lets the new @-mention wake the agent.

Security

  • events.poll no longer lets an agent eavesdrop on another workspace's channels (audit B3). The event-poll RPC previously scoped its results by a caller-supplied workspaceId, so a same-user pipe/MCP client could live-subscribe to any workspace's private channel messages, channel lifecycle, and A2A task pointers just by naming that workspace's id — no pane identity required. Those confidentiality-sensitive event types are now scoped to a server-resolved workspace derived from the caller's verified senderPtyId (the same identity anchor the a2a.channel.* mutations already use), and the caller-supplied workspaceId is ignored for them; an unresolvable caller receives none of these events (fail-closed). The bundled MCP wmux_events_poll tool forwards its own PID-walked senderPtyId, so a legitimately-placed agent still sees its own channels and tasks unchanged. The first-party operator surface (the app's own renderer/plugin host) keeps scoping across the local workspaces it names. Ordinary lifecycle events (pane/process/agent/workspace metadata) are unaffected — their all-workspace firehose was already reachable by any events.subscribe subscriber, so their workspace scope was never a confidentiality boundary and external lifecycle subscribers keep working.

v3.19.0

Added

  • Task lifecycle: close, one-click PR, and a cleanup list (J3). A fan-out task's diff surface now carries 닫기 (Close) and PR buttons, so you can finish a harvested task without touching the terminal. Close runs in a deliberate order — it removes the task's git worktree first and only commits the close (and archives the mission channel) once the worktree is gone, so you can never end up with a "closed" task whose output still litters disk. If the worktree is dirty, close is held: the task stays open, the output is preserved, and a toast tells you to review the diff and commit/PR or discard it. If there are committed-but-unpushed commits, close warns instead of silently dropping them. PR is one click (with a single confirm that names the branch and warns a pre-push hook may run): it gates on gh being installed and authenticated, refuses if the worktree is dirty (uncommitted work wouldn't be in the PR), pushes the branch, and opens a PR against the repo's default branch — and it's idempotent, so a second click after a half-finished attempt recovers the existing PR URL instead of erroring. The PR URL is recorded on the task and the PR-status cache is refreshed immediately. A new "태스크 정리 목록" (Task Cleanup List) command in the palette scans the dedicated worktree root against live tasks and surfaces four kinds of leftovers — unmaterialized-open, disk-missing, dirty-preserved, and orphaned directories (reverse-mapped by an on-disk task.json stamp so they're identifiable even after a closed task ages out of memory) — with an inline Close for the ones that are still open tasks. If a fan-out agent pane comes up but its prompt never fired, you now get a "프롬프트 미발사" toast with a 재발사 (re-fire) action that re-sends the task's original startup command (agent launch + prompt together, same sanitization as the normal path) after checking the prompt file still exists — it never pastes the raw prompt into a bare shell. Finally, a task workspace whose pane wanders outside its worktree boundary gets a small ⚠ 이탈 badge in the sidebar (best-effort, warning only — nothing is blocked).

  • Operators can now join private agent-made channels. The channels panel grows a collapsed discovery section listing every channel on the daemon — including private rooms agents created without inviting the human, and archived rooms for audit visibility — with a one-click join. Joining seats the operator as a regular member with full history, and appends a server-published, viewpoint-neutral system marker ("Operator joined this channel") to the channel as an audit row; the marker consumes a sequence number but owes no member an unread, so agents are not nudged by it. The join surface is strictly human-side: the RPC methods are unreachable from agent transports (pipe router unregistered, first-party MCP exclusion), pinned by boundary tests.

  • Fan-out missions are now visible in the sidebar and fleet panel. Workspaces created by a J1 fan-out now show up under a "Missions" group at the top of the sidebar (title, open/closed status, and a link into the mission's channel) — the group only appears when a workspace has fanned out, so ordinary workspaces are unaffected. The fleet panel's cards also grow a mission line when they belong to a fan-out task. The existing worktree badge (⊕) is untouched — it marks the low-level "this is a git worktree" fact, while the new Missions section marks the higher-level "this is a fan-out task" fact, and a workspace can carry both. Mission data is read-only and pulled (mount + workspace-set changes + a 15s background poll for status drift + an immediate refetch right after a fan-out completes), since the daemon doesn't push mission updates.

Changed

  • Fleet view is now always-on chrome instead of a full-screen modal. Ctrl+Shift+A still toggles it, but it now mounts as a fixed-width panel alongside the workspace sidebar and channel dock (mirroring the channel dock's existing flex-sibling layout) rather than a fixed overlay with a backdrop — other panes stay visible and interactive while it's open, and closing it no longer drops keyboard focus into <body>: the element that had focus when it opened is restored. The fleet/approvals/remote tabs, keyboard row-navigation, and approve/deny shortcuts are unchanged; the card grid narrows to fit the panel's width instead of a full-screen layout. Two focus bugs found in review were fixed before this landed: opening the panel now lands real DOM focus on the active card/row (not just the panel container, which used to leave keyboard users unable to reach any card when only one was present), and row shortcuts (Enter=approve, Backspace/Delete=deny) now only fire when the option row itself is focused — previously an auto-approve checkbox could steal focus and cause those keys to mis-fire as an approval/denial.

  • Type scale: apply the wave-1 semantic tokens to the always-visible chrome. The sidebar (WorkspaceItem, MiniSidebar), channel dock (ChannelsPanel, ChannelView, ChannelMembers), and fleet panel (FleetCard) now use .text-caption/.text-body instead of hardcoded text-[11px]/text-[13px] — swapped only where the token's actual size (caption=11px, body=13px) matches the literal exactly, so there is no size change. Elements that already carried an explicit font-*/leading-* utility are unaffected (utilities win over the token's own weight/line-height); a handful of small mono labels that had no explicit weight now pick up the caption token's weight 500 instead of the browser default 400 — a deliberate, disclosed exception, not a bug. 8px/9px/10px/12px literals in these six files are left untouched (no matching token without a size change) for a later pass.

  • Design tokens: promote hardcoded modal shadows, z-index literals, link accent, and typography to named tokens (visual-invariant). Internal design-system cleanup with no visual change: the six-way-duplicated 0 25px 60px rgba(0,0,0,0.75) modal shadow and the rgba(0,0,0,0.6) backdrop are now --shadow-modal/--backdrop-modal; eight ad-hoc z-[…] literals map to a named --z-* stacking scale (values and relative order unchanged); the link accent gains an accentSecondary token wired to the existing accent value across all eight built-in themes (a hook for future differentiation, currently identical); and a four-tier typography scale (--text-display/-title/-body/-caption) is defined with three representative applications. All values are byte-identical to the originals — verified against the pre-change literals by a three-model review — so themes render exactly as before. The sidebar's two bespoke "Copied!" DOM toasts (workspace-info copy and cwd copy), which each hand-built a bottom-center element and bypassed the canonical toast surface, now route through the shared toastSlice/ToastContainer so copy feedback is styled by one token-driven container instead of duplicated inline CSS (they adopt the app-wide bottom-right/5s presentation as a result). Four dark-only hardcoded hex values that broke the light themes are tokenized: the browser title bar and URL-bar resting state (#11111b → var(--bg-mantle)) and the browser-close / palette-item hovers (#3b1e1e/#2a2a3d → var(--bg-overlay)) now read correctly under hinomaru/taegeuk — these four spots intentionally normalize to the sibling components' tokens, so dark themes see a subtle shade shift there (e.g. #11111b → #181825, and the two outlier hover tints join the twenty sibling hovers already on --bg-overlay) rather than staying byte-identical. The custom-theme-editor, contrast-warning, and color-inspect chrome keep their fixed high-contrast hex by design (they must stay legible while the live theme is being edited/broken), and the webview inspector overlay keeps self-contained hex because it is injected into arbitrary guest pages that have no wmux theme variables.

Fixed

  • UI responsiveness: clicks no longer contend with a background re-render storm. Interaction latency ("every button feels sluggish") had two dominant causes, both fixed. (1) Renderer re-render fan-out: seventeen always-mounted components (sidebar, status bar, channels panel, composer, palette, fleet view, …) subscribed to the entire workspaces tree, which is replaced on every agent-output metadata tick — and the renderer had zero React.memo barriers, so agent activity re-rendered large components continuously and clicks landed on an already-busy render thread. Subscriptions are now minimal derived selectors backed by a reference cache (unchanged projections return the same array/element references, so components only commit when a field they actually display changes), workspace list items self-subscribe by id behind React.memo, title/cwd/git-branch metadata writes are coalesced to one store write per frame, and the 1-second status-bar clock is isolated into its own tiny component. A new re-render regression suite (React Profiler commit counting + selector reference-contract tests) pins the fix: unrelated workspace churn now produces zero commits in unrelated components. (2) Main-process stall: the 5-second periodic session autosave performed a synchronous atomic write on the main event loop, delaying whatever IPC a click had just issued. The periodic path is now an async atomic write with a write-epoch guard and post-write recovery — if an in-flight async write races a newer event-driven synchronous save (the reboot-survival path), the newer snapshot is re-committed immediately, so the final on-disk state matches the latest save under any interleaving (crash-loss window unchanged at ≤5s; exit paths still flush synchronously).

Added

  • Diff review & hunk adoption: harvest a fan-out task's output (J2). Fan-out tasks now have a fourth surface type — a diff surface — that reads a task worktree's uncommitted changes against its merge-base and lets you review, comment, and cherry-pick them into the target repo. Fan-out's result toast gains a "diff 열기" action that opens the diff for that task's workspace. The panel shows a file tree (numstat), a unified diff (+/- coloring only — no full IDE editor, by design), per-hunk checkboxes, and an adopt button. Adoption is all-or-nothing: the selected hunks are reassembled into a single patch (file headers and hunk bodies preserved byte-for-byte, only hunk line-counts recomputed) and applied with one git apply — the target is either fully changed or fully untouched, never half-applied. Adoption is gated hard: a target snapshot (HEAD/branch/dirty set) is captured at read time and re-verified at apply time (rejects if the target moved), any selected file that is dirty in the target is refused (conflict avoidance), a combined pre-apply --check is the gate (so hunks that only apply together aren't wrongly blocked), and hunks already applied to the target are surfaced as an explicit failure so you can deselect them. Untracked files are synthesized into proper new-file patches (regular files only — symlinks/FIFOs are labeled unsupported so a symlink can't leak a file from outside the repo); rename/copy/mode/binary changes and files over the 512KB/2MB caps are display-only (adoption refused, double-checked). File names with spaces, non-ASCII, or quotes are handled correctly (-z porcelain, quotepath off). Comments post to the task's mission channel with a diff-comment anchor (file + hunk header) and render inline under the matching hunk on reload; comments whose hunk header no longer matches the current diff drop into a "위치 이동됨" group (v1 anchor precision is hunk-header granularity — line-level anchors are deferred). The whole path is backed by a validation rig that proves adoption atomicity under a mid-apply kill and catches a re-serialization corruption (dropped no-newline marker) as a shipping blocker.

  • Perf harness: N-pane instrumentation + boolean consistency gates (W2, dev/CI-facing). Extends the existing A1 app benchmark (scripts/perf-bench.mjs + scripts/perf-compare.mjs, driven by .github/workflows/perf.yml) rather than adding a new harness, turning the B2 engine-resume decision from an undefined "feels blocked" call into recorded numeric + pass/fail gates. Four scenarios now run by default on a dedicated bench instance (isolated from the coldStart/input/RAM numbers): (1) N-pane concurrent-streaming frame budget — the 8-pane split loop is generalized to spawnPanes(client, page, n), and at N=4/8/16 every pane's PTY is flooded with continuous output while the renderer's rAF cadence is sampled; each N is gated independently (scenarios.frameBudget.N{n}.frameDeltaMs.p95, ratio 2.0 = the strategy doc's "budget 2×"). (2) Korean IME composition — since CDP/playwright-core cannot drive a real IME, the scenario synthesizes the DOM composition contract xterm's CompositionHelper consumes (compositionstart/compositionupdate/compositionend + input + textarea.value diff) on the focused pane's hidden helper-textarea and verifies the PTY echoes the composed string (안녕하세요) back byte-for-byte; self-validating (a non-equivalent synthesis would echo nothing and fail). (3) Long scrollback — reuses the existing --scrollback-lines flag as a run combination (no new logic). (4) WebGL context-loss/restore — forces WEBGL_lose_context.loseContext()/restoreContext() on the focused pane's canvas and measures recovery via the webglcontextrestored event + !isContextLost() (plus a live-canvas re-count), recording recoveryMs. perf-compare gains a BOOL_GATES array (baseline-independent: scenarios.ime.pass / scenarios.webglContextLoss.pass FAIL immediately when present-but-not-true) alongside the three new numeric frame-budget gates; both stay record-only until an owner blesses a CI baseline (existing bench/baseline-ci.json convention). New CLI flags: --frame-budget-panes 4,8,16, --skip-frame-budget, --skip-ime, --skip-webgl-recovery. Pure logic (frame-stat summary, IME echo comparison, gate judgment) is factored into scripts/perf-scenarios.mjs and unit-tested; the CDP-driven scenario bodies are validated on the Windows CI target only (this being a macOS worktree, they cannot run locally — an honest, documented limitation). No product-code (src/) changes.

  • Fan-out: one prompt → N isolated agent tasks (J1). The AgentToolbar gains a fan-out entry that spawns up to 8 WorkTask missions from a single prompt, each with worktree isolation by default: a dedicated git worktree under {wmux home}/worktrees/{repoHash}/{taskSlug} on a fresh wtask/{slug} branch, a dedicated task workspace (agent pane + shell pane, startupCwd pinned to the worktree), an auto-opened private mission channel (task workspace invited as a member), and the prompt delivered via a file-backed initialCommand (prompt body lives outside the worktree so task diffs stay clean; the path is shell-quoted for POSIX and PowerShell). The whole call is idempotency-keyed end to end — double-clicks and IPC retries can never mint duplicate worktrees — and a global preflight validates the repo and every task's slug/branch before any task or channel is created (unfit input rejects the batch with zero side effects). Per-task failures compensate individually (mission closed, channel archived, any created worktree preserved — never deleted) and surface in a per-task result report (materialization / channel-link state). Worktree operations are serialized per repo (no index.lock races), dirty worktrees refuse removal (preserve-and-list; no force-delete API exists), and bare/submodule/LFS repos fail closed. The daemon activates the reserved task.update materialization path (branch/worktreePath/paneGroupId, write-once monotonic, owner-or-CEO gated) and enforces the canonical-worktree-path exclusivity invariant. A separate broadcast-only action (send text to every terminal pane in the current workspace) is deliberately kept apart from fan-out — non-isolated "fan-out" does not exist. Includes a reboot-survival demo script (single task round-trip: daemon restart → projection restored, worktree intact on disk).

  • WorkTask mission channels: durable task canon + minimal mission-channel lifecycle (J0, dev-facing). Introduces WorkTask — the worktree-mission unit (domain:'task' in the append-only event log) that J1 fan-out and J2 diff will build on — as a projection-first daemon service (daemon/worktask/WorkTaskService), kept deliberately distinct from the A2A Task (different lifecycle + transition graph). Two new pipe RPCs plus their thin MCP tools (channel_mission_start / channel_mission_close) create a WorkTask AND a bound private mission channel in one call, and close flips the task to closed while archiving the channel. Ownership is server-constructed and born-owned (owner = createdBy, never caller-supplied); close authz is a task-level gate (owner OR CEO), the first line of defense over the channel gate. Identity rides the same senderPtyId → verifiedWorkspaceId server stamp as a2a.channel.* mutations (fail-closed on unresolvable identity). Crash-safety is enforced end-to-end: mission channels carry a wmux:mission:{taskId} topic anchor, boot runs a fixed replay → bidirectional reconcile → closed-GC order (an orphan channel from a crash between channel-create and task-append is archived; a closed task whose channel is still active is re-archived — both idempotent no-ops when already settled), and an append-failure on start triggers an immediate compensating archive (the empty-channel reaper cannot reap it — the creator remains a member). Start/close are idempotency-keyed so a lost-response retry never creates a duplicate mission + channel, and re-closing an already-closed mission is a no-op success. Closed tasks are GC'd from the projection after 7 days (log untouched — a view bound only), with archive-unconfirmed tasks exempt. J1+ materialization fields (branch/worktreePath/paneGroupId/prUrl) and the §6.M lease / born-pending contract are schema-reserved but not yet active; task.mission.list is pipe-only in J0 (MCP exposure deferred to J1). Renderer unchanged.

  • E0 conformance harness: recorder + corpus + differential runner (§6.A M1/M2, dev-facing). Introduces the terminal-emulator conformance harness under top-level core/harness/, the measurement scaffolding for the future clean-room VT core. M1 (recorder + corpus): a script-driven recorder (recorder.ts) spawns a real PTY via node-pty to exercise initial geometry + resize, then emits a deterministic recording.bin (raw bytes), events.jsonl (init/resize/reflow_mode trail with monotonic byte offsets), and meta.json (seed + workload-script sha256). PTY spawn, resize, and abnormal-exit failures are escalated (thrown) rather than swallowed, so a broken geometry-exercise path fails the gate instead of silently no-op'ing. The committed corpus (corpus/) is six deterministic synthetic workloads only — scroll flood, resize roundtrip (80→79→80, an explicit non-reflow control at 40 chars where no wrap occurs), resize reflow (120 chars that wrap into two rows, so the 80→79→80 roundtrip actually exercises the rewrap path — its golden pins xterm.js's observed deterministic post-roundtrip state, not an idealized restoration), alt-screen enter/exit, CJK/emoji/VS16/ZWJ width cases, and the SGR spectrum (16/256/truecolor + attribute flags) — each carrying ≥3 golden assertions next to its definition. A companion miner (miner.ts) scrubs {stateDir}/buffers/*.buf dumps (multi-layer: api-key/token/secret key=value, AWS uppercase-snake credential envs, URL userinfo, JSON "key": "…" credentials, PEM private-key blocks, known token prefixes sk-/ghp_/gho_/xox…, Bearer headers, OSC 52 payloads, and a base64 high-entropy heuristic) to a local-only, git-ignored output whose write root is pinned to core/harness/corpus-local/ (an isolation guard rejects any in-repo non-ignored path) — .buf preserves only the ring tail (no geometry), so mined output is for mid-stream robustness and fuzzer seeds, never the deterministic corpus. M2 (differential runner): differ.ts feeds a recording into @xterm/headless@6 (with @xterm/addon-unicode11 pinned to Unicode 11 as the baseline width model) behind a Subject interface (our E1 core and a third reference plug in later), extracts a full-cell grid snapshot (char, width, fg/bg + portable color booleans, 9 style flags, cursor, active buffer), and diffs two snapshots cell-by-cell into a report whose classification schema encodes the four-way ledger (our-bug / xterm-bug / spec-ambiguous / intended) — where intended is admitted only via an explicit approval list (intended-diffs.json, loaded onto the diff path via loadIntendedDiffs), never implicitly. The diff compares the active buffer (normal vs alternate) before cell comparison and excludes xterm.js's non-portable raw color-mode integers from cross-subject comparison; before replay, the event stream is validated (first event is init, byte offsets are monotonic non-decreasing in original order and within range) and violations throw rather than being hidden by sorting; reflow_mode events encountered during replay are honestly recorded on the result. The four-part baseline gate ships as tests: determinism (two xterm.js runs identical) — including a chunk-boundary robustness check that feeds each recording one byte at a time and requires an identical layout to whole-buffer feed (a narrow, documented ZWJ-joiner-at-write-boundary char difference is the only tolerated exception; widths/cursor/colors/flags must match) — no-crash full-corpus completion, golden-assertion pass, and record→replay round-trip stability that reads the committed corpus into memory first and regenerates into a separate temp dir (the gate never writes the repo corpus, so the drift check is no longer a self-comparison). Throughput is recorded as the xterm.js baseline (steady-state feed MB/s + full-cell extraction time). Wired as a fourth vitest lane (vitest.harness.config.ts, tsconfig.harness.json, npm run test:harness). Zero product-code changes; existing test lanes and typecheck unaffected.

Added

  • Append-only event log: crash-safe primitives (envelope PR1). Introduces the segmented NDJSON append-only log (daemon/eventlog/AppendOnlyLog) and the shared event-envelope schema (shared/eventlog) — the foundation for rewiring the channels and A2A canonical state to a crash-safe commit log (§6.L). Key properties: fsync coalescing (group-commit batches), single-ftruncate per-batch rollback, boot-time forward-scan recovery (trim at the first corrupt byte, no partial promotion), Lamport/seq high-watermark resume (reuse forbidden, gaps permitted), and fail-stop on truncation failure rather than silently diverging coordinates. Includes machine-id minting and recovery, and a durable option for atomicWrite (fsync sequence). No service is wired to this log yet — that lands in subsequent PRs.

  • Event log migration engine (envelope PR2). Adds the zero-downtime boot gate (daemon/eventlog/migrateToEventLog) that promotes legacy channels.json to log mode, plus the durable-only EventLogManifest (atomic migration-complete marker) and SnapshotStore (latest → .bak → reseed → genesis fallback chain). Detection uses three branches: inexplicable state is quarantined under quarantine/ and retried rather than silently accepted. Conversion failures leave the legacy file intact and are idempotent on retry. Downgrade detection uses a Lamport + state-hash watermark — a record of an older daemon's writes triggers a reseed snapshot. Compaction safety: no truncation before durable confirmation; genesis and reseed snapshots are never truncated. Not wired into daemon boot yet.

  • A2A tasks are now durable in the daemon event log (envelope PR4). Canonical A2A task state moves from the renderer's in-memory store (30-min GC, lost on restart) into A2aTaskService in the daemon, persisted as domain:'a2a' envelopes in the append-only log. Create, transition, and cancel all reach the log under fsync commits; tasks survive restarts via projection replay. VALID_TRANSITIONS is enforced daemon-side — out-of-graph transitions are rejected at the canonical source. Background ClaudeWorker transitions (working / completed / failed) now route through the daemon rather than writing directly to the renderer, carrying completion evidence along. The renderer a2aSlice is demoted to a read cache that applies daemon commits verbatim without re-validation; when the daemon is unavailable the existing renderer validation path is the automatic fallback (no degraded behavior). Workspace close force-fails in-flight tasks in the log so they do not resurrect on restart; completed tasks are periodically pruned. Daemon canonical state wins over a stale cache on reconnect, including immediately after restart.

  • A2A event authContext is now server-stamped; daemon.ping exposes the active log format generation (envelope PR5). The authContext.principalId in every A2A task event (create, transition, cancel) is now derived by the daemon from stored task coordinates rather than accepted from the caller's claim — actor pane for transitions (to.paneId), caller-side pane for cancel/create, workspace fallback for headless workers or unpinned tasks. principalId and trustTier are display/routing/audit fields only; the authorization anchor remains the server-pinned verifiedWorkspaceId invariant. trustTier is always 'semi-trusted', resolved unilaterally by the server (the temporary caller-override field from PR4 is removed — callers cannot claim a trust tier). daemon.ping responses now carry eventLogFormatVersion additively: present when log mode is active (value = the active format version integer), absent in the legacy fallback. Absence signals a pre-envelope daemon to the auto-replacement logic, which treats unknown format generations fail-closed.

  • A2A completion evidence: schema and pure validator (§6.M P1). Introduces the CompletionEvidence schema and a pure, side-effect-free validator (shared/completionEvidence.ts). Gate = structure: non-empty summary, well-formed items, sanitized paths, DoS caps on body lengths and item counts. verifiedItemCount is derived honestly — an all-unverified completion is accepted at grade 0 rather than rejected (grade is observability, not a gate requirement). Path sanitization rejects colons, leading separators, .., and C0 control characters (undecoded literals enforced). Untrusted-wire normalization: plain-object check, hasOwn gating, fresh-object copy to prevent prototype pollution. Not wired to any transition at this point — gate activation is the next PR, after envelope PR4.

  • A2A completion evidence: production and transport wiring (§6.M P1). ClaudeWorker now produces structured completion evidence from its Claude run results. Both success and failure paths emit inspection + unverified self-report — run-success is never promoted to verified (no laundering). MCP a2a_task_update transports evidence via a dedicated evidence parameter; the contract is fixed in the tool description and coexists with the existing artifact channel. The renderer bridge normalizes untrusted wire shapes before they reach the store: a poisoned shape is stored as completion_evidence_malformed (additive-inert — no task state change at this stage), and server-only stamps like recordedBy are stripped on ingestion. No rejection gate yet — that is the next PR.

  • A2A completion-evidence gate activated (§6.M P1). completed/failed A2A task transitions now require structured completion evidence: completed needs a non-empty summary plus at least one well-formed item (command/inspection/artifact), and failed needs a summary (the failure reason). The daemon A2aTaskService.transition is the single enforcement point; the renderer fallback writer applies the same gate for pane-pinned tasks driven by a pane-identity caller or when the daemon is unavailable. Rejections return actionable reason codes (completion_evidence_missing, completion_evidence_no_items, completion_evidence_empty_summary, completion_evidence_invalid_item, failure_reason_missing) and leave task state unchanged with no log append. verifiedItemCount remains an honest grade rather than a gate requirement — an all-unverified completion is still accepted (grade 0). Workspace-teardown force-fail and verbatim application of daemon commits intentionally bypass the gate to prevent split-brain.

  • Completion evidence grade is now observable in A2A task events (§6.M P1). a2a.task events received via wmux_events_poll now carry verifiedItemCount (count of independently-verified evidence items; 0 = unverified completion) on completed and failed transitions. Event pollers can now distinguish an unverified completion (grade 0) from a graded one without querying the task separately. The count is derived from task.status.evidence at terminal transitions only — non-terminal transitions such as working carry no count. The renderer's primary publisher emits it; workspace-teardown force-fails emit a separate grade-0 event. The trust boundary admits only non-negative integers (forged or out-of-range values are dropped silently). created and cancelled pointers carry no grade field.

  • Validation rig: harness core + SIM smoke (§6.G, dev-facing). Introduces the self-verifying harness under top-level rig/. Components: run isolation (isolation.ts — fresh temp home per run, 4-env wipe of HOME/USERPROFILE/APPDATA/LOCALAPPDATA, WMUX_DATA_SUFFIX='-rig-{runId}'), headless daemon wrapper (daemon.ts — dist/daemon-bundle spawn with a detached process group, daemon.ping ready-poll, group tree-kill, respawn, explicit error on missing bundle), daemon pipe client (pipe.ts — persistent-socket JSON-RPC, dual-ok-layer unwrap, G6 honest-main discipline: one workspaceId binding per persona, throws on cross-workspace impersonation or reserved identity claims), state assertion helpers (assert.ts — seq integrity, full-body cross-check, unread counts, canonical coordinate comments), and deterministic seed (seed.ts). SIM scenario S1 (flood ×8 concurrent senders → getMessages full cross-check: all-delivered, seq-continuous, no-duplicate) lands as a third vitest lane (vitest.rig.config.ts, npm run test:rig:sim, requires npm run build:daemon first). Zero product-code changes; existing two test lanes unaffected.

  • Validation rig: simulator scenarios S2–S8 + SIM regression-detection evidence (§6.G, dev-facing). Completes the synthetic multi-agent simulator on top of the R1 harness. The persona framework (rig/harness/persona.ts) handles identity assignment, channel preamble, seed wiring, and member lifetime; behavioral scripts are owned by each scenario. Deterministic scenarios S2–S8 each run against an isolated daemon: S2 channel integrity under ping-pong load; S3 dead-member expiry — unread, membership, and message-ledger remnants asserted against the client-side cursor only (avoids cursor-circular derivation from lastReadSeq); S4 hung-member: post commits immediately with no infinite hold, unread stays accurate; S5 deliveryStatus receipt contract pinned at current behavior (ack-only pending→delivered); S6 cap-boundary ±1 at the wire level (body 8192 B, mention cap 64, evidence item count 64 / item string 4096 B — string overflow is too_large at the gate, item-count overflow is malformed at wire normalization); S7 SIGKILL mid-flood → respawn → one-way subset assertion {ok-commits} ⊆ replay (at-least-once tail promotion: "no uncommitted resurrection" is intentionally NOT asserted); S8 full A2A lifecycle (send→working→completed, gate-rejection→retry, idempotent resend) plus detection of the #354 idempotency-authz ordering bug (non-participant key-replay is blocked after authz, not before). EPERM chaos: chmod 000 on the Unix socket → client isolation, daemon survival, and recovery confirmed; skipped under root (DAC bypass). CL7 early gate opened via stage-1 detection evidence (rig/EVIDENCE.md): #354 fix reverted on a scratch branch → S8 red confirmed → main green restored. Dogfood script catalog (rig/CATALOG.md): 29 scripts triaged — absorb 4, keep 24, retire 1 (zero physical deletions). Zero product-code changes.

v3.17.0

Added

  • wmux now updates its own background daemon — no manual restart. When an upgraded app reconnects to a daemon left running by an older version, it replaces it automatically: the old daemon suspends every session durably (scrollback, running commands, agent conversations), a current-version daemon starts, and your panes restore themselves — scrollback replayed, supervised commands relaunched, agents resumed. Same session preservation as a full quit-and-restart, without the quit. A brief "Updating the background daemon" toast explains the pause. The 3.16.0 stale-daemon banner remains as the fallback for the cases the replacement deliberately refuses (a NEWER daemon is never downgraded; a daemon that won't shut down cleanly is left running rather than force-killed pre-save).
  • Every agent in a channel now has one honest name — owned by the server, not typed by the agent. Channel display names are derived by the daemon from its pane registry (the same auto-names you see on panes, like w26-1(claude)), so an agent can no longer post under an arbitrary label and two Claude panes can never collapse into one indistinguishable "Claude Code". Names even follow agent swaps: replace claude with codex in a pane and its next message posts under the new name automatically.
  • Recovered agents show up as invite and @-mention candidates right after launch. Previously a workspace you hadn't visited yet contributed nothing to the "Add an agent pane" picker until you clicked into it once; the app now asks the daemon which panes are running agents at startup.

Changed

  • Quitting the app during a daemon replacement now does the right thing for both quit flavors: a normal Quit leaves the fresh daemon running with your restored sessions (tmux-style persistence), while "Shut down wmux completely" guarantees no daemon survives — including one spawned mid-replacement.
  • While the daemon is shutting down for a replacement (or full shutdown), new pane creation is rejected with a clear error instead of silently creating a pane that would be lost in the handover.

Fixed

  • Agents no longer get re-nudged about their own messages. A CLI/MCP agent posting under a stale member id matched no roster seat, so its own post counted as its own unread and the wake worker kept poking it. Posts are now mapped onto the workspace's actual seat (when unambiguous) — and when a workspace has several seats and none match, the sender gets an explicit warning instead of a silent identity fork, including on idempotent retries.
  • The same pane can no longer hold two channel seats. Joining once via the GUI and once via the CLI (or joining before and after agent detection) used to create duplicate roster rows — double nudges, double delivery entries. Joins now converge onto the pane's canonical seat and name the existing seat when they collide.
  • CLI agents stopped colliding on the shared "agent" identity. Panes are spawned with a unique $WMUX_MEMBER_ID, wmux channel join requires an identity instead of silently defaulting, and the join reply reports the seat you actually got.
  • Channel mention nudges are no longer typed into a plain shell terminal. When a member's agent pane was busy (its real Claude pane owned by the on-screen window), the wake worker could auto-submit its wmux channel read … hint into an agent-less shell, where it ran as a stray command; it now stays silent there and leaves delivery to polling.

v3.16.0

Added

  • You are ONE person in channels now — everywhere. Your channel identity is a single app-wide seat instead of one seat per workspace: the roster shows just "Me" (no more "Me · Workspace 2"), your channel list / memberships / unread badges are identical no matter which workspace is open, and joining or creating a channel no longer stamps whichever workspace happened to be active. The daemon merges your previously scattered per-workspace rows into the one seat at boot (deterministic, crash-safe, keeps your earliest join date and furthest read position).
  • Upgrades can't silently wipe your channels anymore. wmux keeps the background daemon alive across app restarts by design, so an upgraded app could attach to an old daemon and channels would look missing (posts failed with no explanation). The channels panel now detects the stale daemon and shows a "quit wmux fully and start it again" banner; it clears itself after the restart.

Changed

  • The unread badge is honest now. Agent posts from the workspace you're looking at used to be silently muted (workspace-level self-mute); with the unified seat, only YOUR OWN posts stay quiet — an agent posting from any workspace counts as unread, because it's news to you.
  • Adding a whole workspace as a channel member is retired — you are already in your channels as one seat, and agents join as individual panes.

Fixed

  • Private agent-only channels no longer leak into your dock. A private channel between agents whose workspace happened to be active could bump your unread badge for a channel you can't even open (phantom badge). Display is now scoped to channels you are actually in.
  • The channel wake worker no longer sweeps the virtual human seat every tick (it owns no terminal, so the sweep was pure CPU drift that grew with history).

Security

  • The reserved human seat cannot be invited, claimed, or targeted from the agent pipe — an agent could previously seed a phantom "human" member row that force-injected its channel into your always-on view. Rejected at both the pipe router and the daemon, so a direct-socket caller cannot bypass it either.