macOS diagnostics watchdog

The repository includes an opt-in, local watchdog for capturing evidence when the Celorga app becomes unresponsive, consumes sustained CPU, or crosses a resident-memory threshold. The helper runs outside the app, so it can still sample all app threads when the main thread is blocked.

The watchdog is a development tool. It never restarts, quits, or mutates the app, and it does not promote diagnostic output into canonical notes.

Run the watchdog

Build and start the helper from the repository root:

npm run diagnostics:macos

By default it follows the daily app with bundle identifier org.org.workspace and reads that app's remembered corpus from machine-local preferences. The corpus must have a valid portable identity in org2.json. Pass an explicit corpus or monitor the isolated Codex app when needed:

npm run diagnostics:macos -- --corpus /path/to/corpus
npm run diagnostics:macos -- \
  --bundle-id org.org.workspace.codex \
  --corpus /path/to/disposable-corpus

Press Control-C to stop. Run npm run diagnostics:macos -- --help for every threshold and override.

The installed app must include the matching heartbeat responder for direct main-thread hang detection. When monitoring an older build, the helper reports that the responder is unavailable and continues with CPU and memory monitoring instead of treating the missing responder as a hang.

Runtime cost and triggers

While idle, the helper sends one distributed heartbeat and reads one proc_pidinfo resource snapshot per second. It does not run ps, top, Instruments, or continuous log collection. The heavier /usr/bin/sample profiler runs only after a trigger.

Default triggers are intentionally conservative:

  • a main-thread heartbeat gap of 5 seconds,

  • process CPU at or above 95 percent for 15 seconds,

  • resident memory at or above 1,536 MB,

  • a 10-minute cooldown after any captured incident.

The stack capture lasts 5 seconds at a 10-millisecond sampling interval. Thresholds and sampling duration are configurable from the command line. Setting --hang-seconds, --cpu-percent, or --memory-mb to zero disables that trigger.

Interaction latency

The app separately instruments ordinary UI responsiveness at frame scale. It records pointer-event-to-window-update, source-editor-key-to-draw, and AI composer-key-to-draw intervals in the org.org.workspace/InteractionLatency unified-log category and as Instruments points of interest. A sample over the 60 Hz frame budget (16.7 ms) is logged immediately; every 120 samples the app also reports rolling p95. This is deliberately distinct from the watchdog: the watchdog captures multi-second hangs, while these intervals expose visible jank that never becomes a hang.

Run the optimized regression boundary with:

npm run test:macos-performance

The GitHub Actions workflow OpenOrg macOS correctness and performance runs the correctness suite before deterministic release-mode performance regressions and records its Xcode, Swift, and SDK versions. Once the shared runtime builds, the performance step still runs after a correctness failure so performance does not go unmeasured. A failure in either step fails the workflow; cancellation stops further work. Inspect the failed step before treating the workflow result as a measured slowdown. Existing regression thresholds remain enforced in the release-mode step. The full interactive scale benchmark runs separately through npm run test:macos-performance with its generated corpus and measurement report.

On Apple Silicon the test launcher runs Swift natively even when npm itself is running through Rosetta, so the optimized test bundle matches the app's target architecture.

CLI latency

Every CLI and bundled renderer process launched by the Celorga app emits one structured completion record in the org.org.workspace/CLILatency unified-log category. Records include a sanitized command identity, elapsed milliseconds, outcome, exit status, and standard-input, standard-output, and standard-error byte counts. They never include command data arguments, corpus paths, search queries, environment values, or output contents.

Recent records can be inspected during development with:

log show --last 10m \
  --predicate 'subsystem == "org.org.workspace" && category == "CLILatency"'

The record is emitted once, after the process and its output collectors settle, so it does not add polling or per-byte work to a CLI invocation. The Swift CLI wrapper also accepts an injected structured metric sink for focused tests and in-process diagnostics consumers.

Incident storage

Captured evidence is primary input, so it belongs under raw/ rather than compiled/:

raw/diagnostics/org2-workspace/YYYY/MM/DD/
  TIMESTAMP-TRIGGER-PID-ID/
    incident.json
    stacks.sample.txt

incident.json uses the versioned org2:workspace-diagnostic:v1 envelope defined by spec/v0/workspace-diagnostic.schema.json and records the trigger, app/build identity, recent resource snapshots, sampling status, and a privacy-safe hot-path fingerprint. stacks.sample.txt is the macOS thread-sampling report.

These files use .json and .sample.txt extensions, so the Celorga app's corpus event classifier does not parse them or invalidate workspace projections. Generated grouping, diagnosis, and fix reports should go under views/diagnostics/ and cite the raw incident directory.

The watchdog validates the portable corpus identity before writing and remains pinned to that corpus for its lifetime. It does not record note contents, chat bodies, credentials, the absolute corpus path, or an absolute app path. Treat stack reports as local development evidence: symbol and framework names may still describe what the app was doing.

Current recovery boundary

The first version is capture-only. It deliberately does not kill or relaunch a frozen app because unsaved editor state may exist. Safe cache shedding, cancellation of obsolete rendering work, and an explicit capture-and-relaunch action can be layered on after incident capture is proven reliable.

Full workspace refresh timings

The org.org.workspace/WorkspaceRefresh unified-log category records each awaited full-refresh stage and total wall time as stage and elapsed_ms. Stage names contain no corpus paths or document text. Independent stages overlap, so their durations do not sum to total wall time. Background health and indexing tasks may finish later. The existing CLILatency category separates command execution costs; workspace agent-state also returns timings for its individual list sections.

/usr/bin/log show --last 10m --info --style compact \
  --predicate 'subsystem == "org.org.workspace" && (category == "WorkspaceRefresh" || category == "CLILatency")'