Skip to content

Observability & Telemetry ​

Local logs for debugging plus a switchable, content-free telemetry stream for product metrics, launch diagnostics, and sanitized errors.

Logging ​

Use createLogger(scope) from electron-log everywhere in the desktop app:

ts
import { createLogger } from '@/lib/logger'
const log = createLogger('Sync')
log.info('pull complete', { count, durationMs })
log.error('pull failed', err)
  • Never use console.*. A pre-commit hook flags it.
  • Logs land in the OS-standard log directory and rotate automatically.
  • Renderer and main process logs are separate files.
  • Dev runs log at debug; packaged installs are lowered to info (file) / warn (console) at startup based on app.isPackaged, since NODE_ENV is undefined at runtime in packaged builds.
  • Important launch, renderer, and main-process errors are mirrored as telemetry events when product telemetry is enabled.

Choosing a Level ​

The redacted diagnostic log stream is floored at warn (see Error Logs in Grafana (Loki)), so warn and error are the levels an operator actually triages. Reserve them for conditions someone can act on, and log expected steady states at debug:

  • Certificate pinning falling back to standard TLS when no pins are configured (main/sync/certificate-pinning.ts) — a deliberate fallback, not a failure.
  • sodium_mlock / sodium_munlock being absent in the WASM libsodium build (main/crypto/memory-lock.ts) — the expected state in Electron. An mlock/munlock call that is present but fails still logs at warn.
  • Progress notes the embedding worker writes to stderr, such as transformers.js reporting an unknown content-length during a model download (main/lib/embeddings.ts). A stderr chunk is only downgraded when every line in it matches a known-benign pattern, so a real failure interleaved with progress output still reaches error.
  • Vector-clock bumps the increment*ClockOffline helpers make while the sync runtime is down (packages/sync-client/src/offline-clock.ts) — the normal offline-edit path, one call per edited row (per changed field for tasks and projects).
  • An index rebuild triggered by a missing index DB (emitIndexRecovered in main/vault/index.ts) — checkIndexHealth reports missing when index.db is not on disk, which is the expected first open of a vault: fresh install, newly linked device, or a deleted file. The index is a derived cache, so rebuilding it costs nothing but time. It is emitted at info, with the same index_recovered action and errorCode as before. corrupt and migration_failed stay at warn — those are genuine data-corruption recovery.

A log payload is built before the transport decides whether to keep it, so a hot path pays for its arguments at every level. Keep these lines to identifiers: log the row id and the changed field names, not the resulting clock or field-clock objects.

Log Locations ​

PlatformPath
macOS~/Library/Logs/memrynote/
Windows%USERPROFILE%/AppData/Roaming/memrynote/logs/
Linux~/.config/memrynote/logs/

Older installs logged into an @memry/desktop directory (the raw package name). On startup the main process moves those files into the directories above (a name collision gains a legacy- prefix), then removes the emptied legacy directory. Dev profiles (MEMRY_DEVICE) log into a per-device memrynote-<device> directory instead.

The same applies to the app identity as a whole: production launches adopt the memrynote runtime app name and move userData (Application Support/@memry/desktop → …/memrynote, leaving a compatibility symlink for downgraded binaries and stored absolute paths). The macOS Safe Storage keychain item is copied to the new name so existing encrypted secrets keep decrypting; on Linux a populated safeStorage store keeps the install on the legacy identity (the keyring item cannot be carried over). See src/main/app-identity.ts.

Launch Phase Timeline ​

src/main/launch-timeline.ts stamps each startup milestone's offset (ms) from process start and emits them as one structured line, launch timeline, when the main window is revealed. Phases are recorded with recordLaunchPhase(phase), which also forwards each one as the per-phase app_launch_phase_completed telemetry event.

FieldMeaning
appReadyMsElectron app.whenReady() startup work finished
windowCreatedMsmain BrowserWindow constructed
vaultOpenStartMs / vaultOpenReadyMsvault restore started / reached isOpen
rendererLoadedMsrenderer did-finish-load
readyToShowMsfirst ready-to-show (absent when it never fired)
shownMswindow actually revealed
reasonready-to-show, fallback-timeout, or did-fail-load
fallbacktrue when the 10s reveal fallback fired
vaultOpenPendingvault open was still running at reveal — the prime suspect

Vault-open timing stops at isOpen, not at the autoOpenLastVault() promise, which also waits on the first full sync. A launch onto the vault picker records no vault phase at all.

The line is logged at warn when the reveal came from the fallback or took ≥5s — only warn/ error records reach the diagnostic log sink — and at info otherwise, so healthy launches stay local instead of flooding the sink.

Post-Reveal Startup Queue ​

src/main/post-reveal.ts holds startup work the first frame does not depend on, so it cannot sit between window creation and the reveal and inflate shownMs. Register with onceWindowShown(name, task); the reveal calls schedulePostRevealTasks(), which drains the queue one second later. The delay is deliberate: ready-to-show means the renderer can paint, not that it has finished booting, and draining on the next tick would move main-thread contention rather than remove it.

name is the at-most-once key. A macOS dock reopen re-creates and re-reveals the main window, so a caller on that path registering again must not start a second copy. Work registered after the drain runs immediately. A task that throws or rejects is logged under the Startup scope and never stops the others. If the app begins quitting inside the delay the queue is left undrained, so a deferred task cannot re-arm a service the shutdown sequence has already torn down.

The updater is the first tenant: initializeUpdater() reconciles install health, reads its prefs, selects its backend and fires the startup-check update check, none of which is on the way to the first frame.

Login-Shell PATH Gate ​

A packaged app launched from Finder or the Dock inherits only the minimal system PATH, so which claude and which codex fail and Agent Chat greys out providers the user has installed. src/main/agent/cli/login-shell-path.ts recovers the real PATH by running the login shell, which sources the user's whole rc chain and costs anywhere from 180 ms to over a second.

That probe is not on the boot path. startLoginShellPathAugmentation() fires it at module scope and returns immediately; whenLoginShellPathApplied() settles once process.env.PATH has been merged. Anything that resolves an executable by name must await that gate rather than read PATH and hope. runBinaryCommand() in src/main/agent/cli/binary-detection.ts is where the agent CLIs do it, which covers which, the --version reads and, transitively, the per-turn spawns that run after detection. Headless --cli awaits it too, because it shells out before any window exists.

The probe still logs Augmented PATH from login shell for packaged launch when it changed PATH, and still reports login_shell_path_probe_failed with a failure class per breakage: status_<n> for a shell that exited non-zero, spawn_error for one that never started, marker_missing for output the parser could not read. A silent failure here is indistinguishable from an uninstalled CLI, so the breadcrumb is the only way to tell the two apart from a user's logs.

Telemetry ​

Telemetry is enabled by default in production builds and off by default in development builds. Users can turn it off via Settings → General → Privacy.

The build channel comes from MEMRY_ENV when set (dev/staging profiles) and otherwise from app.isPackaged, so packaged installs report production. A telemetry choice the user has already saved always wins over the channel default.

What Ships ​

Only enums and event metadata:

ts
trackTelemetry('page_viewed', { surface: 'notes', action: 'viewed' })

Recognized surfaces (TelemetrySurface in packages/contracts/telemetry-api):

app, home, onboarding, vault, notes, journal, tasks, inbox, calendar, search, graph, settings, sync, ai, voice, updater, canvas, projects, tags.

What Never Ships ​

  • Note content
  • Note titles
  • Identifiers (note IDs, task IDs, project IDs)
  • Search queries
  • Tag names
  • User file paths and note filenames
  • Raw (unredacted) exception messages — a desktop error message can embed a note title or content
  • Scraped page metadata — the <title>, description or address of a URL the user clipped

The contract uses string-typed enums for surfaces and actions; arbitrary strings can't sneak through. The one exception is error diagnostics, which additionally ship a redacted stack trace of code locations and, optionally, a message — but only after the client has run it through redactText (packages/contracts/src/redact.ts), never the raw string — see Error Reporting.

The dimension allowlist ​

dimensions is the one free-form slot on an event, so its key namespace is closed. TELEMETRY_DIMENSION_KEYS in packages/contracts/src/telemetry-api.ts lists every key allowed to leave the device; sanitizeTelemetryDimensions keeps at most one entry, drops any key not on the list, and drops any value that fails the safe-value shape (no @, no ://, no slash, ≤ 64 chars, not UUID-shaped).

The shape check alone is not enough, which is why the allowlist exists: a scraped page title such as Divorce settlement calculator passes every one of those rules. Only an enumerable key namespace can tell a bounded enum from arbitrary user or page content.

Enforcement sits in createTelemetryClient's track() (apps/desktop/src/main/telemetry/client.ts) — the single chokepoint every event crosses before the durable queue and the network, whether it came from trackMainEvent, the renderer's IPC handler, or a direct runtime.track call. A new call site cannot opt out by skipping a helper. The renderer wrapper (src/renderer/src/lib/telemetry.ts) applies the same shared function so the UI drops rejected dimensions before the IPC hop.

Adding a key is the review gate. If the value cannot be enumerated ahead of time, it does not belong in a dimension — send a metric instead (a count, a duration, or a bucket label such as result_bucket).

The allowlist is deliberately not part of TelemetryDimensionsSchema. The sync-server validates /telemetry/batch with that schema and rejects a whole batch on one bad event, so narrowing it would make a newly deployed server 400 batches sent by already-shipped desktop builds. The allowlist is enforced on the client, where the data still is.

Failure detail on a failed request ​

An event that reports a failed HTTP request may also carry a failure object (TelemetryFailureDetailSchema) with three bounded fields:

FieldShapePostHog property
httpStatusinteger 100–599http_status
serverCodethe server's SCREAMING_SNAKE error.codeserver_code
retryableboolean — was the failure classified retryableretryable

This exists because sync_error used to ship one opaque server_error label covering 400, 403, 404, 409 and every 5xx alike (#1584): a permanent client-side contract bug and a transient edge 5xx were the same row, so no chart separated them and no alert threshold could be set. The label is unchanged — error_code still reads server_error — and these fields sit beside it, so error_code='server_error' AND http_status >= 500 now answers "is the backend down?" and retryable=false answers "how many of these will never succeed?".

An absent server_code on a 5xx is itself the signal: the response never reached the Worker's error handler, so it came from the edge and the backend has no record of it.

It is a field of its own rather than a dimension for two reasons. An event may carry at most one dimension and sync_error already spends it on transport; and unlike a dimension value, every field here is bounded by construction (a 3-digit range, an anchored enum-ish token, a boolean), so it can never become a free-text channel. sanitizeTelemetryFailure runs at the same createTelemetryClient.track() chokepoint as the dimension allowlist and drops any field that would fail validation — one bad field must never cost the whole batch.

The schema addition is optional and additive in both directions: an older desktop simply omits it, and a sync-server on older contracts strips the unknown key rather than rejecting the batch, so neither deploy order can 400 events from an already-shipped build.

On the desktop the fields are assembled in one place — apps/desktop/src/main/sync/sync-error-telemetry.ts — from what classifyError already computed, so a new sync_error call site cannot forward the category and forget the rest.

Because warn/error log records ship too (Path A, see What Ships and Error & Diagnostic Logs in PostHog), the same rule applies to log lines and thrown error messages in content-handling code: the inbox scraper (apps/desktop/src/main/inbox/metadata.ts) logs the clipped URL and scraped title at debug — below the ship floor — and keeps a content-free warn so the event stays countable.

Tracking Pattern ​

All telemetry calls are fire-and-forget — never await:

ts
void trackTelemetry('onboarding_completed', {
  surface: 'onboarding',
  action: 'completed',
  result: 'success'
})

The void makes the call non-blocking and unfailable from the UI's point of view.

WebGL Availability ​

Every Sigma graph surface (graph-page, graph-canvas, local-graph-panel) probes hasWebGLSupport() (apps/desktop/src/renderer/src/lib/webgl-support.ts) before mounting. Chromium no longer falls back to SwiftShader for WebGL on its own, so main/index.ts appends --enable-unsafe-swiftshader before ready (enableSoftwareWebglFallback() in gpu-crash-guard.ts): a launch without GPU WebGL — hardware acceleration disabled by the crash guard, a blocklisted driver, Remote Desktop, a VM — renders the graph in software instead of showing the "Graph isn't available on this device" fallback. When even that fails, the probe records one app_log_recorded warn per session (source: WebGLSupport, log_action: webgl_unavailable) so affected installs can be counted in PostHog. The fallback is an expected device condition, never an app_error_seen exception.

Events Before the Runtime Exists ​

trackMainEvent used to no-op while getTelemetryRuntime() was still null, so anything that failed during early startup — the app-identity migration, the Safe Storage keychain carry-over, the GPU crash guard disabling hardware acceleration — was never reported. Early events are now buffered in memory (apps/desktop/src/main/telemetry/track.ts, capped at 100 events) and forwarded by drainEarlyMainEvents(), called from main/index.ts immediately after initializeTelemetryRuntime. Each event keeps its original occurredAt. The buffer is set to null once drained, so a runtime disposed during shutdown cannot quietly re-accumulate events nobody will ever flush.

registerMainDiagnostics() is registered before any other startup work for the same reason — the later call in the ready handler is an idempotent no-op — so a failure in the identity carry-over / pending-install window still produces a report.

Shutdown Drain ​

One flush sends at most TELEMETRY_BATCH_LIMIT events while a session can queue up to TELEMETRY_QUEUE_LIMIT, so the previous single flush('shutdown') silently dropped everything past the first batch. Shutdown now drains in bounded rounds (ceil(TELEMETRY_QUEUE_LIMIT / TELEMETRY_BATCH_LIMIT)), stopping at the first failed round so an offline quit never stalls the exit.

Event Categories ​

CategoryEvents
Surface viewspage_viewed — one per active-tab change; carries the tab type as objectType
App lifecycleapp_started, app_backgrounded, app_active_heartbeat, app_update_installed, deep_link_opened (coarse target only — open, billing_start, billing_complete, billing, oauth, pair, unknown; never the URL)
Onboardingonboarding_started, onboarding_completed
Vaultvault_created, vault_opened
Notesnote_created, note_opened, note_updated (throttled 1/doc/5 min), note_deleted, note_imported, note_exported
Journaljournal_opened, journal_updated (throttled 1/doc/5 min)
Taskstask_created, task_completed, task_reopened, task_updated, task_deleted
Projectsproject_created, project_opened, project_updated, project_archived, project_deleted, project_item_linked
Tagstag_created, tag_renamed, tag_deleted, tag_merged, tag_category_created
Canvascanvas_created, canvas_opened, canvas_deleted, canvas_card_added, plus the rollout counters canvas_sync_conflict_copy, canvas_too_large, canvas_asset_uploaded, canvas_asset_dedup_hit, canvas_asset_gc_reaped
Inboxinbox_captured, inbox_filed, inbox_archived, inbox_snoozed
Calendarcalendar_event_created, calendar_event_updated, calendar_event_deleted, calendar_google_connected, calendar_google_disconnected, calendar_google_sync_completed
Remindersreminder_created (source = preset or custom; preset id in the value dimension), reminder_deleted — emitted by the reminder picker, so surface is the page it was opened from
Importimport_completed
Homehome_board_customized
Searchsearch_performed, search_result_opened (search_opened is defined but unused — it would duplicate command_palette_opened)
Command palettecommand_palette_opened, search_result_opened (palette context)
Graphgraph_opened — on graph page mount
Voicevoice_recording_completed (duration + bytes), transcription_completed (success/failure + processing duration)
Settingssetting_changed — surface only, never the value
Agent chatagent_chat_started, agent_chat_message_sent, ai_action_completed (turn result + duration)
Sync healthsync_enabled, sync_run_completed, sync_error (counts/status only, plus the failure detail)
Authsignin_started, signin_succeeded
Diagnosticsapp_log_recorded, app_error_seen, app_launch_phase_completed, app_crashed (see Crash & Unclean-Shutdown Detection)

Crash & Unclean-Shutdown Detection ​

A hard crash — main-process abort, OOM kill, force quit — used to discard the in-memory telemetry queue, so the crash itself never shipped: the classic "it crashed and there are no logs" report. Two mechanisms fix that: a marker file that notices the crash, and a durable queue that keeps the resulting event alive long enough to send.

A marker file (apps/desktop/src/main/telemetry/crash-marker.ts):

  • session-marker.json is written into userData at startup with the session id, startedAt, lastAliveAt, and app version, then refreshed every 60s while the app is alive.
  • clearCrashMarker() removes it once the shutdown cleanup chain completes.
  • A marker still present at the next launch means the previous session died uncleanly, and that launch emits app_crashed on its behalf. Detection runs before the new session writes its own marker.

The marker's presence is the signal; its contents only enrich the event. An unparseable marker still reports the crash, just without the observed-uptime metric (metrics.durationMs, derived from lastAliveAt − startedAt). The previous session's app version ships as a prior_app_version dimension.

errorCode separates the failure modes:

errorCodeMeaning
UNCLEAN_SHUTDOWNno shutdown was attempted — hard crash, OOM kill, force quit
SHUTDOWN_TIMEOUT_<STEP>shutdown ran but its budget expired while <STEP> was still running
SHUTDOWN_TIMEOUTsame, but the marker carries no step (written by an older build)
SHUTDOWN_CLEANUP_FAILEDthe cleanup chain rejected

The last three are stamped by markShutdownFailure() immediately before the forced exit: the log line for that failure never flushes, but the marker survives to the next launch. Only the process that wrote a marker may remove one — a second instance that loses the single-instance lock shares userData and must not erase the primary's marker on its way out. Marker write failures are logged and swallowed; a read-only disk must never break startup.

The overrunning step rides in the errorCode rather than in a dimension, because an event ships at most one dimension and that slot already carries prior_app_version. The SHUTDOWN_TIMEOUT prefix is preserved so a query written against the old code still matches. Only a bounded kebab-case token is accepted from the marker; anything else degrades to the plain code.

app_crashed also carries an assembled message naming the shutdown failure, the overrunning step, the prior version, the observed uptime and whether the marker parsed — without it the Error Tracking issue was titled UNCLEAN_SHUTDOWN and held nothing else. The marker is a file on disk, so every string field is rejected outright unless it matches an enum-ish token (a character-substituted path still leaks its structure), and the assembled message is capped at 512 characters. That cap is not cosmetic: an over-length message fails TelemetryErrorDetailSchema at the sync-server, which rejects the whole batch with a 400, and the desktop client treats a 4xx as permanent — one corrupt marker field would otherwise discard up to 100 unrelated events on every launch until the marker cleared.

Shutdown Budget ​

before-quit runs its cleanup as an ordered list of named steps under one shared deadline (apps/desktop/src/main/shutdown-sequence.ts):

ConstantValueRole
SHUTDOWN_BUDGET_MS8,000 msthe whole graceful chain
SHUTDOWN_LAST_CHANCE_MS1,500 msdurability flush granted after the budget is gone
SHUTDOWN_HARD_BACKSTOP_MS10,000 mstimer outside the sequence; the process always exits by then

The budget is derived from the bounded waits the chain contains, not guessed: 2,000 ms for the renderer flush handshake (windows in parallel) plus 3,000 ms for the voice, image-processing and embeddings utility stops — which run concurrently, so 3,000 ms together rather than 9,000 ms in a row — leaves 3,000 ms of headroom for the unbounded steps. Every step is also handed a cap() that clamps its own bounded wait to what is left of the shared deadline, so no set of waits can collectively overrun it.

Two rules keep a slow quit from becoming a lossy one:

  • Order by durability. The window flush and flushPendingWritebacks() run first, so a wedged teardown step behind them degrades a quit to slow rather than to lost edits.
  • Never force-exit with pending writes. When the budget expires, the step that overran is stamped into the marker, then flushPendingWritebacks() and closeAllDatabases() run inside the last-chance window before app.exit(1). closeAllDatabases() matters because both SQLite files run synchronous = NORMAL, which defers durability to the checkpoint that close() performs. The cleanup-error path does the same.

A quit where nothing is wedged still completes in milliseconds; these ceilings are only reached when a teardown step is genuinely stuck.

Durable Queues ​

The marker only detects the crash — app_crashed still has to survive long enough to be sent, and both telemetry queues flush on a 30s interval. A second hard crash inside that window would otherwise take the crash report with it, which is why both queues mirror to userData (telemetry/queue-store.ts):

QueueMirrorCarries
Event queue (telemetry/client.ts)telemetry-event-queue.jsonapp_crashed, app_error_seen, all events
Log-ship queue (telemetry/ship-queue.ts)telemetry-log-queue.jsonPath A redacted warn/error log lines

Every enqueue reaches disk before it returns, and the mirror is rewritten after every flush, so what is on disk is what has not yet been accepted by the server. The next launch restores it and drains it; a drained batch is removed from the mirror, so nothing is sent twice.

Rules the mirror follows:

  • Format is a journal: a {"version":2} header line followed by one JSON item per line. An enqueue appends its own line rather than re-serialising the queue, which otherwise made each event cost a full rewrite of up to 500 objects — worst exactly during the error bursts and offline sessions that keep the queue pegged at its limit. The file is rewritten (compacted) on drains, on trims, and once the journal outgrows its bound, so it stays bounded.
  • Format changes cannot wedge startup. The previous {"version":1,"items":[…]} format is still read, so an upgrading install keeps whatever its last session queued. A file whose version this build does not recognise is discarded rather than parsed, builds that predate the mirror never read it at all, and a build that predates the journal sees it as unparseable and discards it — so a downgrade costs one session's queue, never the launch.
  • Corruption is expected. The mirror is most likely to be truncated by exactly the crash it was written to survive, so an unparseable file is logged, deleted, and treated as empty.
  • Write failures are non-fatal. A read-only or full disk costs durability, never logging or shutdown; the failure is logged once per streak rather than once per line.
  • The limit applies to the restored set too (TELEMETRY_QUEUE_LIMIT / SHIP_QUEUE_LIMIT), so a mirror written by a build with a larger limit cannot resurrect an unbounded queue.
  • Opting out deletes the mirror. Turning telemetry off clears the file, not just the in-memory queue, and a launch that starts with telemetry disabled discards the mirror instead of restoring it.

Restored events keep their own occurredAt; the batch envelope is stamped with the session that ships them, so an event resurrected from a dead session is attributed to the launch that sent it.

Native Crash Dumps ​

crashReporter.start({ uploadToServer: false }) runs before app.ready, so main, renderer, and utility processes all write minidumps for native crashes that no JS handler ever observes. The dumps stay in app.getPath('crashDumps') for the Path B diagnostic bundle the user submits deliberately. uploadToServer must stay false: a minidump is raw process memory — PostHog does not ingest minidumps, and there is no way to redact one, so uploading would breach the redaction model every other telemetry path is built around.

PostHog Event Capture ​

PostHog is the sole telemetry store (POSTHOG_HOST, https://us.i.posthog.com by default). The sync server transforms each accepted desktop event into a PostHog event (services/posthog-transform.ts) and posts it to the PostHog capture API's /batch/ endpoint (services/posthog.ts, capturePostHogEvents). Event names are preserved from the existing 50-event contract, with one rename: page_viewed → $pageview, which unlocks PostHog's native path-analysis and web-analytics views. Batch metadata (platform, arch, locale, app version, build channel, auth state, sync state, timezone offset) becomes person properties ($set); the event's own dimensions are flattened onto event properties first, then overwritten by server-derived keys (surface, action, environment, session_id, $session_id, $lib, $lib_version) so a client can never spoof a trusted key by naming a dimension after it. $session_id carries the same per-launch UUID as session_id: PostHog's session-scoped metrics and Error Tracking's sessions count read only $session_id, so without it every desktop event was session-less and all desktop traffic collapsed into a single session. $lib is memry-desktop ($lib_version is the app version), which is what Error Tracking's library filter matches on. Server-side business and error events (services/analytics.ts) post to the same capture API, tagged surface: 'server', so one PostHog project holds every event from both desktop and server.

Errors additionally become $exception events for PostHog Error Tracking (exceptionEvent in services/posthog-transform.ts), fingerprinted on our own errorCode when one is present so grouping follows the app's own error taxonomy rather than PostHog's pattern-hash default. They carry $session_id, $lib and $lib_version too, so an issue reports a real session count and is reachable through the library filter.

Generic constructor names do not group. toErrorCode falls back to the error's constructor name, so every bare new Error('...') reports errorCode: 'Error'. Pinning the fingerprint to that merged roughly thirty unrelated production failures — offline network errors, a Squirrel read-only-volume update failure, data.db failed PRAGMA quick_check, missing agent API keys — into one issue titled after whichever stack the first sample happened to carry (#2134). The transform therefore omits $exception_fingerprint for the built-in constructor names plus the UnknownError / StringError fallbacks (NON_DISCRIMINATING_ERROR_CODES), handing those back to PostHog's pattern hash so they split into real issues. A generic name still rides along as $exception_list[0].type; only the grouping key changes. Give a failure a typed code if you want it grouped as its own issue.

Three server-side guards keep this stream from flooding the PostHog quota:

  • Warn-level log lines never reach Error Tracking. The desktop demotes expected failures to warn-level app_log_recorded lines precisely so they stay out of Error Tracking (#1587); exceptionEvent promotes an app_log_recorded event only when its action is error. Warn-level lines stay fully queryable as events and log records.
  • Legacy drop-tripwire noise is dropped at ingestion. Desktop versions before 2026.821 ship the local_mutation_dropped tripwire once per polled row with no throttle or eligibility gate (#1579) — at peak 45% of the project's entire event volume, triple-billed as product event, $exception and log line. isLegacyMutationDropNoise drops those events before all three sinks; fixed clients still forward their throttled diagnostic trickle.
  • Per-install hourly exception budget. claimExceptionBudget (services/exception-budget.ts) caps $exception forwards at 60 per install per hour, reusing the rate_limits table. Only the $exception stream is trimmed — the product events and log lines for the same failures still forward, so a capped install stays diagnosable. The claim fails open on D1 errors: a flaky database costs extra PostHog events, never a swallowed crash report.

Stack frames ​

Error Tracking renders code locations only from $exception_list[].stacktrace — it never parses the exception's value. The desktop sends its stack as redacted text (that is the shape the client-side frame filter and redaction produce), so parseStackFrames in services/posthog-transform.ts turns each at fn (file:line:col) line back into a raw frame:

json
"stacktrace": { "type": "raw", "frames": [
  { "platform": "custom", "lang": "javascript", "function": "push",
    "filename": "~/app/sync.ts", "lineno": 12, "colno": 5,
    "resolved": true, "in_app": true }
] }

Four rules that are load-bearing:

  • Reversed. PostHog treats the last frame as the throw site; a JS stack string is innermost-first. The cap of 50 frames is applied before reversing, so a deep stack loses its outermost callers rather than the frame that actually failed.
  • platform: 'custom'. Claiming web:javascript enters PostHog's symbolification path, which needs uploaded source maps and a per-frame chunk id. We ship neither, so it would resolve to nothing; custom frames render verbatim.
  • in_app: false for node:*, internal/*, node_modules and electron/js2c frames. The UI hides non-in-app frames by default, falling back to showing all of them when an exception has none — so a fully-vendor stack is never blank.
  • Omitted, not empty. An exception with no parsable frame carries no stacktrace key at all. frames: [] would claim we resolved a stack and found nothing; utility-process crashes and log-derived errors genuinely have none.

value holds the redacted message alone — it is the issue title, and the stack belongs in frames. A React component stack is promoted to frames when there is no JS stack, and always ships intact as $exception_component_stack.

When there is no message the transform falls back to the error code, which makes the issue title identical to the code and tells an engineer nothing new. That fallback sets exception_message_missing: true on the event, so a message-less reporting site is countable on a dashboard instead of looking healthy:

sql
SELECT properties.$exception_fingerprint AS fp, count() c, uniq(distinct_id) u
FROM events
WHERE event = '$exception' AND properties.exception_message_missing
  AND timestamp > now() - INTERVAL 7 DAY
GROUP BY fp ORDER BY u DESC

Errors that carry no JS stack by construction — child-process-gone for a crashed utility worker, where the process that died is not the one reporting — instead carry a synthesized message naming the worker, reason and exit status, so their issue page is not blank.

This transform is entirely server-side: it applies to batches from already-installed desktop versions as soon as the sync-server deploys.

No raw identifiers are stored: the install ID is HMAC-hashed server-side (TELEMETRY_HMAC_KEY, hashTelemetryId) and used as the PostHog distinct_id; server-side user_id/device_id/ vault_id are hashed the same way before they ride along as event properties.

Account identity ​

The desktop attaches its access token to /telemetry/batch and /diagnostics/report as an optional bearer. Neither route runs the auth middleware: resolveTelemetryAccountHash verifies the JWT if one is present and returns undefined for a missing, malformed or expired token, so telemetry is never rejected for auth reasons — that batch simply reports anonymously against its install hash.

The resolved account id is HMAC-hashed before it can become a distinct_id, exactly like the install ID. TransformContext names the field accountHash, and resolveDistinctId shape-checks it against hashTelemetryId's output (64 lowercase hex chars); anything else — most plausibly a raw account id — degrades to the install hash rather than reaching PostHog. This is deliberately strict: a PostHog $identify merge is permanent and cannot be undone or re-keyed, so a raw account id that reached a person profile could not be removed afterwards.

When a batch resolves to an account, a $identify event aliases the anonymous install person onto the account person. It fires once per app session, guarded by the telemetry_identify_sessions D1 table (claimIdentifySession, migration 0003, swept by the cron cleanup after 24h). Without the guard, the desktop's ~30s flush cadence would emit one identified event per batch. The guard fails open: a D1 error emits $identify anyway (idempotent in PostHog) rather than leaving the install unlinked.

Diagnostic reports resolve identity through the same resolveDistinctId path as events and logs, so a report lands on the same person profile as the events around it.

Usage segmentation ​

Two dimensions exist so "how many people use MemryNote" can be split without identifying anyone.

auth_state (anonymous | signed_in | signed_out) comes straight from the batch and is written as both an event property and a person property. The event property is the one to break down on: a person property holds the latest value, which answers "is this install signed in now" rather than "had yesterday's active users ever signed up".

plan and plan_status are person properties only, read from sync_entitlements by resolveTelemetryPlan — never sent by the client, which must not be trusted with a dimension like free vs pro. They resolve once per app session, behind the same claimIdentifySession claim that gates $identify, so the lookup costs one D1 row per session rather than one per ~30s batch. Every other batch omits both keys entirely rather than sending them as null, which would wipe what the session's first batch wrote. The status travels with the plan so a canceled pro cannot be counted as a paying user. resolveTelemetryPlan fails closed to undefined: a token whose account row is gone costs one person property, never the whole batch.

resolveTelemetryAccount returns the raw userId alongside accountHash purely so this lookup can read the account's own rows. It stays inside the worker — accountHash remains the only identity that reaches PostHog.

Note that an install which merely runs in the background still produces telemetry, so counting unique persons over "all events" measures installs that were running, not people who used the app. app_active_heartbeat (emitted only while a window is focused) or a real product event is the honest signal for the latter.

Known limitation: telemetry identity is verified but not revocation-checked. A revoked device's still-unexpired access token (≤15 min) can attribute telemetry until it lapses. Telemetry is not an authorization decision, so a per-batch device lookup is not worth the D1 read.

Environments are separated by an environment property on every event inside one PostHog project, not by separate projects.

Additional events in the same pipeline:

EventSource
app_launch_phase_completedElectron main/renderer startup milestones
app_log_recordedSanitized desktop diagnostic breadcrumbs
app_error_seenRenderer, React boundary, and main errors
server_error_seenSync-server request/background failures
server_log_recordedStructured sync-server diagnostic logs
release_download_count_snapshotDaily GitHub Releases download-count pull (not a user event)

Batch Validation & Retry ​

Each /telemetry/batch payload is schema-validated on the sync server. A malformed batch is rejected with 400 VALIDATION_ERROR. The server logs the failing field paths (Zod path + issue code only — never the field values, which may hold the raw identifiers the schema is designed to strip) so rejections are diagnosable without leaking data.

The desktop client treats a permanent 4xx (any 4xx except 429) as unrecoverable and drops that batch, so one malformed event cannot wedge the queue head and replay the same rejected batch on every flush. Transient failures — 5xx, 429, and network errors — leave the batch queued for a later retry.

Token Lookup & Request Deadlines ​

The batch client asks the token manager for an access token only when the install is signed_in; an anonymous or signed_out install ships every batch without touching the secret store. Looking the token up regardless cost an OS keychain round-trip per flush, and on a machine whose keychain hangs (see Cryptography) that round-trip wedged the whole pipeline: events stopped while the log shipper — which never asks for a token — kept going. That "logs but no events" split is the diagnostic signature; the diagnostic report upload (sendIncidentReport) sat behind the same lookup and showed as Sending… forever.

Every net.fetch on the telemetry path — event batches, log shipping, incident reports — goes through boundedNetFetch (telemetry/bounded-net-fetch.ts) with a 30 s abort deadline. A request still in flight when Chromium's network service dies never settles on its own, and a flush loop awaiting it would otherwise stall for the rest of the run with every later batch queued behind it.

The keychain side of that hang is bounded twice over: OS keychain calls are serialized through one process-wide single-flight queue with a run-wide unavailability latch, so a wedged Secret Service can burn at most one libuv threadpool thread instead of all four, and the main process raises UV_THREADPOOL_SIZE to 16 before anything can use the pool. See Cryptography.

Autosave Event Throttling ​

note_updated and journal_updated events fired by the autosave path are throttled to at most one emission per document per 5-minute window (in-memory, resets on restart). This prevents high-frequency editor flushes from inflating event counts.

Body edits reach the note through the CRDT provider rather than the notes UPDATE IPC, so typing never registered as usage at all. trackNoteBodyEditThrottled (telemetry/diagnostics.ts, called from ipc/crdt-handlers.ts) now emits note_updated with source: 'editor_body' on the samenote_updated:<noteId> throttle key the UPDATE handler uses, so metadata saves and body edits share one 5-minute window per note. Only the throttle key ever sees the note id; the event itself carries no identifier.

The shared throttle map (telemetry/throttle.ts) is bounded to 1000 keys. Because the keys are per-document (note_updated:<noteId>, journal_updated:<date>, and the CRDT writeback keys), exceeding that inside one window is ordinary operation — a vault import or a writeback pass over a large vault does it — so the cap is enforced rather than advisory: keys whose window has elapsed are swept first, and if nothing has expired the oldest-inserted keys are dropped anyway. Dropping a key only forfeits its throttle, never an event.

Each entry records the window it was written under, and the sweep judges expiry per entry rather than by the window of whichever call happened to cross the cap. Callers do not share one window — the Google Calendar sync runner throttles on 60 seconds while the autosave keys use the 5-minute default — so a short-window caller must not be able to expire a still-live 5-minute entry and make note_updated re-emit early.

Release Download Counts ​

Downloads happen on GitHub Releases, where PostHog cannot see them — the landing site's download-click event measures intent, not a download. A daily cron on the sync server (services/release-downloads.ts, run from the scheduled handler at 04:00 UTC) reads GET /repos/memrynote/memry/releases and emits one release_download_count_snapshot event per asset.

It is a metric snapshot, not a user event. One emission per poll per asset, from a fixed service distinct_id — nobody downloaded anything at the moment it fired. Person and user counts on it collapse to the service identity and mean nothing, and it must never be used as a funnel or conversion step next to genuine user events like landing_download_click or app_started. The name says so: it was release_asset_downloaded until 2026-09-10, which read as a user action and was mistaken for one. The old name is marked deprecated and hidden in PostHog data management; events emitted before the rename still carry it, so any query spanning the cutover must ask for both names.

assets[].download_count is cumulative per asset. Emitting it raw would produce a monotonically increasing counter that is useless as an event stream — and it fails silently, producing meaningless numbers rather than an error. The last total seen per asset is therefore stored in D1 (release_download_counts, migration 0003) and only the delta is emitted:

  • The first run for an asset seeds its row and emits nothing; a cumulative counter carries no meaningful delta until it has a baseline.
  • A total that went down — GitHub recounting, or a replaced asset — reseeds the baseline rather than emitting a negative delta.
  • The store is written before the events are captured. A D1 failure then throws, the cron reports it, and the untouched baseline makes the next run emit the full delta. Emitting first would double-count that delta after a failed write.

The pull rides its own cron entry (crons = ["0 */6 * * *", "0 4 * * *"]) so it runs once a day while the cleanup sweep keeps its 6-hourly cadence. The daily entry deliberately avoids the 6-hourly times — colliding entries collapse into one invocation.

Events are not person-scoped: an anonymous downloader has no identity to key on, so distinct_id is a fixed memry_releases_<environment>. Staging and production both poll the same public repo, so — as everywhere else — an insight that does not filter environment blends them.

PropertyMeaning
release_tagRelease the asset belongs to (v2026-08-06)
asset_namePublished filename
platformmacos / windows / linux / unknown, derived from the filename
asset_kindinstaller, update_metadata, or update_package (see below)
downloadsThe delta — sum this, never cumulative_downloads
cumulative_downloadsTotal GitHub reported at pull time, for context only

asset_kind is load-bearing. Some assets are only ever fetched by the auto-updater. The landing site never links to them (apps/landing/api/download.ts), so every count on one is an installed app updating itself, and counting it as a download swamps the number that matters:

asset_kindAssets
update_metadataPolled on every update check: electron-builder latest*.yml and .blockmap, Velopack releases.<channel>.json and legacy RELEASES
update_packageDownloaded when an update applies: Velopack .nupkg, and the macOS .zip Squirrel.Mac installs from
installerEverything else: .dmg, -setup.exe, MemryNote-win-Setup.exe, .AppImage, .deb, the portable -win.zip

Filter to asset_kind = 'installer' for real downloads. Two caveats remain. -setup.exe and .AppImage are also fetched by electron-updater on NSIS and AppImage installs, and the filename cannot tell the two apart, so installer still carries some update traffic. And until 2026-09-28 the Velopack feed, .nupkg, and the macOS .zip were all labelled installer (Velopack assets also platform: unknown for RELEASES and .nupkg). releases.win.json alone put ~7,000 phantom Windows "downloads" into the week of 2026-09-21. Queries spanning that cutover must classify by asset_name, not asset_kind.

Downloads cannot be joined to activation. An anonymous downloader and a desktop install share no key. The funnel only works for people who sign up on the landing site and sign in on the desktop, where the identity merge puts both on one person. This is a limitation to state plainly, not to engineer around.

Landing Site Telemetry ​

The marketing site (apps/landing) runs analytics browser-side via posthog-js (apps/landing/src/lib/analytics.ts), ingesting through PostHog's reverse-proxy subdomain (https://e.memrynote.com) — session replay cannot be routed through the sync server, so landing traffic does not go through /telemetry/batch. The old POST /telemetry/web endpoint and its LandingTelemetryBatchSchema contract are gone.

  • Client: init() lazily configures posthog-js once per page load, keyed on VITE_POSTHOG_KEY with api_host = VITE_POSTHOG_HOST (defaulting to https://e.memrynote.com), person_profiles: 'identified_only', and masked session recording (session_recording: { maskAllInputs: true }). It no-ops with no window (SSR/prerender) or no key configured. trackLandingPageView fires PostHog's native $pageview; trackLandingEvent fires one of a fixed set of landing_* event names.
  • Payload: pages and targets are path-only — query strings and hashes are stripped client-side (stripQueryAndHash) before either ever leaves the browser. UTM params (utm_source/medium/campaign/content/term) are read from the query string, trimmed, and capped at 120 characters.
  • Environment: environment is registered once via posthog.register — Vercel's VITE_VERCEL_ENV when present, otherwise a production/development split on Vite's build MODE — so landing traffic is filterable apart from desktop/server events in the same PostHog project.
  • Scanner noise: a before_send filter drops one $exception fingerprint — Object Not Found Matching Id:N, MethodName:update, ParamCount:4. That is Microsoft Office / Outlook SafeLinks pre-fetching a link out of an email, injecting its own scanner into the page, and then losing its own object handle; MethodName / ParamCount are a COM bridge's idioms and appear nowhere in this repo. It arrives in same-day bursts from a handful of readers a few times a quarter and has no type and no usable stack, so it is pure noise in the error list. The match is deliberately narrow — only that exact COM signature. A broad "drop every non-Error rejection" rule would hide real bugs.

Session replay coverage ​

Landing replay is not sampled. In PostHog project 412311, session_recording_opt_in is on and session_recording_sample_rate, session_recording_minimum_duration_milliseconds, session_recording_linked_flag, and the URL / event trigger configs are all unset — every session is eligible. The client passes only masking options (see #860 for the cookie-consent posture: no banner, no consent gate, so nothing client-side suppresses the recorder either).

$sdk_debug_rrweb_start_attempted is a per-event property, not a per-session one. It reports the recorder's state at the moment that single event was captured. analytics.ts captures the first $pageview immediately after posthog.init(), and the recorder only starts once PostHog's remote-config response comes back — so in almost every session the first event carries start_attempted = false (or no value at all) even when the recording starts a few hundred milliseconds later. Grouping sessions by their first $pageview therefore undercounts coverage badly; it is the wrong shape of query, not a coverage number.

Aggregate over the whole session instead — max(properties.$sdk_debug_rrweb_start_attempted = true) grouped by $session_id — or just count raw_session_replay_events. Measured over the 7 days to 2026-09-10 ($lib = 'web'): 350 of 408 sessions (85.8%) started a recording, and raw_session_replay_events holds stored replays for 365 sessions. The same window read first-$pageview-only says 49%.

Of the 58 sessions that never started a recording, 46 captured exactly one event — a bounce that ended before the remote-config round trip completed. That residual is structural: the posthog-js bundle is imported lazily (it is the largest dependency on the site; see the comment on load in analytics.ts), so init, remote config, and recorder start all happen after first paint. Eagerly importing it would shrink the gap at the cost of first paint, which is a trade the lazy import deliberately makes in the other direction.

Practical consequence for anyone reasoning about the landing funnel: replay covers ~86% of sessions with no sampling bias, but the uncovered slice is skewed towards the shortest sessions. Do not read replay as evidence about immediate bounces.

Error Reporting ​

Desktop error reporting follows the same product telemetry setting. Each captured error ships stable metadata — process area, component/source, action, phase, and the error's code (errorCode, a typed code where one exists — see Vault File Errors) — plus a redacted stack trace and, for React boundaries, the component stack.

errorCode prefers a typed code the error carries over its class name: a richer telemetryCode (NoteError's note error code plus the originating errno, see Vault File Errors), then a plain .code (better-sqlite3's error.code, a Node system code), walking the cause/AggregateError.errors chain to a bounded depth so a fetch failed TypeError still surfaces the underlying ECONNREFUSED. A note write failure therefore reports NOTE_WRITE_FAILED:EBUSY rather than collapsing every note fault to NoteError, and a locked database reports SQLITE_BUSY rather than an un-triageable SqliteError. A code is only trusted when it looks like an enum token (^[A-Za-z][A-Za-z0-9_.:-]{0,63}$); anything else — a path, an email, a URL, free-form prose — is rejected outright and the class name is used instead, because a character-substituted path (_Users_kaan_secret.md) still leaks its structure. The class name itself still passes through the safe-token rules (no @, ://, /, \, ≤64 chars).

An unhandled rejection can carry any value as its reason — a string, a plain object, or a cross-realm Error that fails instanceof Error — and those carry no stack, which previously landed in Loki as an unactionable bare Error with an empty stack. Reasons are normalized before reporting: a real Error passes through, a cross-realm error's own frames are adopted, and anything else gets a stack synthesized at the handler plus a code naming the reason's type (Rejection_string, Rejection_Object, Rejection_undefined). The reason's message rides along redacted, not dropped: it goes through the same redactText pass as any other error message before it leaves the device. Omitting it made every Rejection_* row in Error Tracking an issue titled after its own error code, with nothing inside to triage. A reason that crossed a structured-clone or IPC boundary keeps its .name but loses both its stack and its constructor; that name is preferred over the constructor name, so it reports Rejection_TypeError rather than collapsing to Rejection_Error. When the code is a Rejection_* name the stack is the handler's own frames, not the fault's — the code is the actionable part.

A window error does not always carry an error object: cross-origin scripts and some Chromium failure paths report only a message and a source location, which previously landed as StringError with an empty stack and nothing to triage. The error class is recovered from the message's leading token (Uncaught TypeError: … → TypeError, subject to the same enum-token rule) and the filename/lineno/colno are rebuilt into a stack frame, so the code location survives the same frame filter and redaction as a real stack. The message text ships too, redacted on the device by the same pass — without it a cross-origin failure was a WindowError issue titled WindowError.

Because a rejection reason or event.error can be any value — including a Proxy whose traps throw or an object with throwing getters — every property read in this path (including instanceof, which can trap getPrototypeOf) is individually guarded. A hostile value can no longer throw out of the diagnostics handler and destroy the report being built.

The free-form exception message was historically never sent at all, since on the desktop it can embed a note title, filename, or content. buildErrorDetail now ships it (TelemetryErrorDetailSchema.message, optional, capped at 512) after running it through redactText (packages/contracts/src/redact.ts) — the server re-runs redaction in mask mode as a backstop. This is what makes an issue readable: without a message PostHog titles every issue with the bare error code, which is how a whole family of production issues came to read StringError / StringError. redactText strips known-sensitive shapes (secrets, tokens, emails, ids, home-directory paths, content-file basenames) rather than proving the remaining prose is note-free, so this is narrower than the earlier all-or-nothing "no message field" guarantee. The stack is separately reduced to code-location frames only — the leading Name: message header line is stripped — so a crash's location shows up as, for example, TypeError at pushRecords (…/sync-engine.js:120). Frame file paths are app source/bundle locations (not user files); any home-directory prefix (/Users/<name>, C:\Users\<name>) is rewritten to ~, and emails, UUIDs, JWTs, and bearer tokens are scrubbed from anything that ships.

IPC Error Throttling ​

The action an IPC error reports is the channel it was registered on (notes:create), not the handler's function name. Handlers are registered as ipcMain.handle(Channel, createValidatedHandler(Schema, async (input) => …)), and an arrow passed straight in as an argument has an empty name — so every inline handler in the app used to collapse into one literal action, validated_handler, and a schema rejection could not be attributed to a channel from the wire data alone (the captured stack names only the bundled wrapper, and Zod strips its own frames). installIpcChannelLabels (main/ipc/lib/ipc-channel-labels.ts) records the pairing once at the ipcMain.handle boundary — the only place that knows both halves — and registerAllHandlers calls it before the first registration. A handler registered without it keeps the old generic label.

Every IPC envelope error becomes a telemetry event, so a handler stuck in a failure loop could flood the queue. trackIpcError (main/ipc/validate.ts) therefore emits at most one event per action:errorCode per 60-second window, in-memory and reset on restart. The key includes the action because keying on the error name alone let one handler's benign recurring Error mask a genuine Error from an unrelated handler for the whole window. An expected condition (Ollama not running, an abandoned OAuth flow) is skipped before the key is claimed, for the same reason — otherwise the suppressed error would keep refreshing a key it never reports on.

That key set is bounded to 1000 entries. Once past the cap, keys whose window has elapsed are swept; if a burst of previously unseen codes fills the map inside a single window with nothing to expire, the oldest-inserted keys are dropped instead. Dropping a key only forfeits its throttle — the next error for it is reported rather than lost.

Vault File Errors ​

A class name alone is often too coarse to act on: every failed note save reported NoteError, which cannot tell an antivirus or cloud-sync file lock apart from a full disk. NoteError therefore carries the originating fs error as its cause, and reports a composite errorCode of its note error code plus the errno — for example NOTE_WRITE_FAILED:EBUSY (locked) versus NOTE_WRITE_FAILED:ENOSPC (out of space). The errno is admitted by a strict allowlist (/^E[A-Z0-9]+$/), so the vault file path is never part of the code — paths are user data and stay out of telemetry, as above.

Writes to a locked file are retried a bounded number of times before failing (see withTransientFsRetry in main/vault/file-ops.ts). Each retry is written to the local log with its errno and attempt number — again never the path — so a slow or failed save is explainable from a user's log file even when telemetry is switched off.

Sync-server error reporting is server-side. Because the sync server is end-to-end-blind (it only ever holds ciphertext), its own error strings are operational: the redacted message and stack ship to Cloudflare's own first-party Workers console logs (with the raw userId/deviceId/vaultId attached, since that sink is trusted) and to PostHog Logs (with those same ids HMAC-hashed) — never to the PostHog event itself, which only gets the coded server_error_seen event with no message or stack; this is what makes sync failures debuggable without handing ciphertext-adjacent detail to a third party. The server_error_seen event itself carries no user, device, or vault id either — its distinct_id is a fixed memry_server_<environment> value, so the event alone cannot be traced to an account. Only the PostHog log record carries the HMAC-hashed userId/deviceId/vaultId (when the caller has them), which is what lets a failure be correlated to an account inside Logs without exposing the raw id. Dynamic path segments and query strings are normalized away. Expected handled 4xx responses (e.g. SYNC_PAYMENT_REQUIRED) are still counted as server_error_seen but are logged at warn rather than error level, keeping real failures distinguishable from expected noise.

Server Business Events ​

Server-side product events — user_signed_up, user_logged_in, device_registered, vault_registered, vault_deleted, and the Paddle subscription events — go through captureBusinessEvent in apps/sync-server/src/services/analytics.ts. Unlike server_error_seen, these have a real actor, so their distinct_id is the HMAC hash of the acting user's id (hashTelemetryId, same key and shape as the install hash). The raw id never reaches PostHog, and the hash is also kept in the user_id property so queries written against that property keep working. If a business event ever has no acting user, it falls back to the fixed memry_server_<environment> id rather than an empty one.

Before this, every server business event shared that fixed id, so uniq(person_id) over any of them evaluated to 1 and every unique-user funnel or retention metric touching a server event was wrong. Events emitted before the fix keep the old distinct_id and person history is not backfilled, so person-level metrics over server events are only correct from the fix forward. Event counts were unaffected in either period.

Error & Diagnostic Logs in PostHog ​

Error events also become searchable log lines in PostHog Logs. This replaced a self-hosted Grafana + Loki instance the sync server used to push to; PostHog Logs now carries the diagnostic detail (stacks, operational messages) that a PostHog event deliberately omits.

  • Transport: the sync server posts log lines to PostHog Logs' plain OTLP-JSON receiver ({POSTHOG_HOST}/i/v1/logs, services/posthog-logs.ts, pushPostHogLogs) — no OpenTelemetry SDK is used — authenticated with the PostHog project token (POSTHOG_KEY) as a bearer. Pushes are fire-and-forget in waitUntil: a missing key (local dev) is a silent no-op, and a failed push can never affect request handling. Records are grouped into one resourceLogs entry per app (desktop / server) so service.name and deployment.environment stay resource-level attributes rather than being duplicated onto every line.

  • Desktop errors: /telemetry/batch events carrying an errorCode or error detail are forwarded as app="desktop" lines containing the event name, error code, surface/action/source, app version, platform, a redacted message (TelemetryErrorDetailSchema.message is optional and only accepted after the client has run it through redactText; the server re-runs redaction as a backstop) and redacted stack frames, log_action — the operational breadcrumb that keeps log-type error events (which carry no stack of their own) identifiable — and exit_code, the platform exit status for process-lifecycle events (empty string when absent, since exit code 0 is itself meaningful). A failed request additionally contributes http_status, server_code and retryable (see Failure detail on a failed request) — also empty string when absent, because retryable: false is a real answer and must not read as "not reported".

  • Redacted diagnostic logs (kind=log, Path A, always-on): a main-process electron-log transport (apps/desktop/src/main/telemetry/log-ship.ts, installed once from main/index.ts — never from logger.ts, which must stay electron-free for worker bundling) intercepts every warn/error record and redacts it via redactLogLine (packages/contracts/src/redact.ts) before anything leaves the device, using a per-install salt (diagnosticsSalt, persisted in telemetry.json) and the active vault root. Redacted lines batch (queue 500 / batch 50 / flush 30s, drop-4xx-except-429) to POST /telemetry/logs, gated on the telemetry toggle (getTelemetryRuntime().getSettings().enabled) and disabled outright in dev builds. Repeated identical level|scope|message lines within a 3s window are throttled into one line with a repeatCount field. The Path B ring those lines also feed is a fixed-capacity circular buffer (200 slots, 5-minute window) that stores each line's epoch ms at push time and evicts oldest-first through a head index, so one warn/error costs O(1) instead of re-parsing every retained timestamp — the error path has to stay cheap when something is looping. A record's arguments are flattened by parseRecord: the first string wins the message slot, plain objects merge into the fields, and an Error contributes errorName plus errorMessage. That last part matters — logger.error('updater error', err) used to ship {"errorName":"Error"} and nothing else, because the label had already claimed the message slot and the Error's own message was dropped. Worker processes (embeddings, image processing, voice transcription) forward their own warn/error records to main over process.parentPort (apps/desktop/src/main/lib/log-forward.ts, electron-free) for the same redaction + ship pass, tagged origin: 'worker' and workerName. The server re-runs redactLogLine in mask mode (no salt) as defense-in-depth before writing to PostHog Logs (desktopLogRecord in services/posthog-logs.ts) — the client-side redaction is primary; the server pass is a second net, not the source of truth.

  • Incident reports (kind=report, Path B): on a real error, the app offers a one-time "Send diagnostic report" action (the tab error boundary, IPC-error toasts, and a Settings entry — available independent of the telemetry toggle). The tab error boundary sends the report automatically, with no dialog and no button, when Settings > General > Privacy > "Automatically Send Error Reports" is on. That flag is autoSendDiagnostics in telemetry.json (telemetry:setAutoSendDiagnostics); a missing key means on, so fresh installs and upgrades default to on and only an explicit false opts out. When it is off, or the automatic send fails, the boundary shows the Send button and the consent dialog instead. The diagnostics:previewReport / diagnostics:sendReport IPC calls build a DiagnosticReport via the same pure buildIncidentReport function: a generated incidentId (MEMRY-XXXXXXXX, random base32), the last ≤200 redacted lines from the Path A ring buffer (≤5 min), a redacted device/sync snapshot (app version, platform, locale, uptime, sync/auth state, queue depth — no content), and the triggering error's redacted stack (header line dropped, frames only). Preview and send call the same builder with the same incidentId, so the consent dialog's preview is byte-identical to what ships. sendReport posts to POST /diagnostics/report; the server writes one summary line plus one line per log entry, all tagged incident_id, under kind=report.

  • Redaction guarantees: same non-negotiables as What Never Ships — no note content, titles, attachment filenames, absolute home/vault paths, emails, JWTs/tokens, vault/device keys, or IPs. kind=log/kind=report message text runs through the same redactText as the /telemetry/batch error message above; redactLogLine additionally redacts each structured fields entry: secrets are dropped first, paths collapse to ~/ / <vault>/, note/attachment basenames are salted-hashed to [name:hash8].ext, known id fields (noteId, deviceId, installId, …) are salted-hashed, and emails become [email:hash8]. IPs are masked to <ip>. UUID-shaped ids in free text are salted-hashed on the client (correlatable, like the id fields above); the server's re-redaction pass has no salt, so it masks them to a fixed <id> instead. A fixed field allowlist (level, scope, action, errorCode, appVersion, buildChannel, platform, arch, origin, workerName, reason, phase, mode, status, kind, result, plus numeric metric keys like durationMs/itemCount) ships verbatim; most other field values run through the same redaction as the message.

  • Updater backends: initializeUpdater() picks one backend at init and keeps it for the session. Velopack when process.platform === 'win32', the app is packaged, and new UpdateManager('https://github.com/memrynote/memry') constructs; electron-updater in every other case, which covers macOS, Linux, and Windows installs made by the older NSIS installer, where Velopack's constructor throws This application is not properly installed. Velopack reads the GitHub releases feed itself (releases.win.json plus the .nupkg assets) and has no app-update.yml. Both backends drive the same state machine in apps/desktop/src/main/updater.ts, so the phase values, the severity classification, the install marker and the install-health streak below are backend-neutral. What differs is the library's own log lines, which carry scope ElectronUpdater on one path and scope Velopack on the other.

  • Updater failures: every failure in apps/desktop/src/main/updater.ts logs a describeUpdaterError() field bag alongside the raw error, so a silent auto-update failure is diagnosable from Loki alone: phase (startup-check, scheduled-check, auto-check-enable, auto-download-enable, or — for electron-updater's own error event, which carries no phase — one of check / download / downloaded / install inferred from the status at the time), plus errorName, errorMessage, errorCode, httpStatus, url, errorCause and the top stack frames as errorStack. The field names are chosen against the redaction allowlist above: phase and errorCode ship verbatim, url is path-redacted (query string stripped), the rest are text-redacted and capped.

  • Updater severity classification: the local main.log line and the user-facing error state are unchanged — every updater failure is still logged at error and still flips the UI to the error state. What is classified is the telemetry severity (apps/desktop/src/main/updater-error-severity.ts). A failure during a check (check / startup-check / scheduled-check / auto-check-enable) whose message or cause chain carries only allowlisted Chromium transport codes — net::ERR_NAME_NOT_RESOLVED, ERR_INTERNET_DISCONNECTED, ERR_NETWORK_CHANGED, ERR_TIMED_OUT, ERR_CONNECTION_TIMED_OUT, ERR_CONNECTION_RESET, ERR_CONNECTION_CLOSED, ERR_CONNECTION_REFUSED, ERR_CONNECTION_ABORTED, ERR_ADDRESS_UNREACHABLE, ERR_ADDRESS_INVALID, ERR_NETWORK_ACCESS_DENIED, ERR_NETWORK_IO_SUSPENDED, ERR_HTTP2_PROTOCOL_ERROR, ERR_HTTP2_SERVER_REFUSED_STREAM — ships as an app_log_recorded warn instead of an app_error_seen exception. Being offline is a normal state for an offline-first app, and those events were 33.2 % of every exception in the product. The cause chain matters because electron-updater's GitHubProvider wraps a transport failure in a parse-shaped ERR_UPDATER_INVALID_RELEASE_FEED; a feed that is genuinely malformed has no network cause and stays an exception. The set is an allowlist, never a net::ERR_ prefix test: net::ERR_CERT_* / net::ERR_SSL_* are security signals, and anything unrecognised fails closed to error. ERR_CONNECTION_ABORTED, ERR_ADDRESS_UNREACHABLE, ERR_ADDRESS_INVALID and ERR_NETWORK_ACCESS_DENIED were added by #1994, measured on the only population that can evidence a gap in this set — builds already carrying this classification (2026.822.1 and newer), since an older build reported every code as an error regardless. An upstream 5xx stays an exception deliberately: a 504 on the releases feed is GitHub failing for everyone at once, the one check-phase shape meaning the whole fleet has stopped receiving updates, unlike the per-device transport codes above that scale with the number of flaky networks. Everything else is untouched — HTTP 4xx/5xx (including the HTTP_ERROR_618 jwt:expired on GitHub's pre-signed asset URLs), signature failures, install-phase errnos, ENOENT … app-update.yml, and any failure in the download / downloaded / install phases, where a network drop can leave a half-applied update. Reclassified events are never dropped: same error code, same redacted message and stack, and each one carries retryCount — the consecutive-failed-check streak — so a cross-install signal can separate one laptop on a train from many installs failing in a row. An install that has not completed a single check in 24 hours and has failed at least 6 checks in that time raises one exception (latched until the next successful check), so a genuinely stuck updater is still loud.

  • Post-exit teardown aborts: a child-process-gone report for a worker the owning module released as graceful_stop is demoted to warn. That release is only ever recorded on an observed code === 0, so the pairing is proof the worker already exited cleanly and the report is the native runtime aborting while unwinding afterwards — not a failure the user felt. This was 70% of macOS installs (#1990): the embeddings worker handled its shutdown message with a bare process.exit(0), which skips JS cleanup and runs onnxruntime's static destructors with sessions still live, aborting with SIGABRT. Node reported 0 and Electron reported 6 for the same process, which is why both codes appear. The worker now disposes the pipeline and drops its last message listener — Electron's ParentPort pauses itself on removeListener, releasing the handle that holds the loop open, and there is no close() to call — with an unref'd fallback that must stay below SHUTDOWN_TIMEOUT_MS in embeddings.ts so a wedged disposal is never force-killed into the very teardown death this removes. The demotion keys on the recorded release alone, never the phase: idle_shutdown is also reachable from a force-kill, where no exit code was ever observed. Demoted reports keep their error code, message, stderr tail and metrics and stay queryable as app_log_recorded; they only leave Error Tracking.

  • Repeated install failures: a failed check is loud in telemetry, but a failed install was silent to the user. On macOS, Squirrel.Mac stages a downloaded update into its own ShipIt copy, and when that copy fails (ditto: Could not lstat …, No space left on device) the error lands in a session that then quits normally. Nothing survived the quit, so the next launch re-served the same cached zip and failed the same way — two production installs sat on 2026.817.1 through four releases with no signal of any kind (#1999). apps/desktop/src/main/updater-install-health.ts persists the streak in update-install-health.json under userData, keyed on the pair (running version, target version), and after three consecutive failed attempts sets installFailed so the existing manual-download dialog appears. It counts attempts, not launches: electron-updater re-serves an already-validated cached zip on every auto-check, so a genuinely stuck install surfaces within ~30 minutes. The streak clears the moment the app boots as a different build, which is the only honest evidence an install applied — display versions cannot be ordered, so "newer than the failing one" is not a test that can be written. Every field is re-validated on read and a corrupt file degrades to "no streak"; an install that has never failed has no file and behaves exactly as before. Both backends share this streak and the pending-install marker (update-install-attempt.json, reported as UPDATE_INSTALL_DID_NOT_APPLY). On Velopack the marker is written before the hand-off to Update.exe and read on the next launch, and an install-phase failure advances the streak exactly as above. Velopack does no download-time staging, so it never produces the downloaded-phase attempts Squirrel.Mac does. The marker also records which installer was handed off to (installer: electron-updater, velopack, or velopack-handoff), and the silent install-on-quit path writes it too, so a Windows update that applies on a normal quit and never comes back is no longer invisible.

  • NSIS to Velopack hand-off: a Windows install made by the older NSIS installer keeps electron-updater for checking and downloading. Once the NSIS setup.exe is on disk, apps/desktop/src/main/installer-handoff.ts looks for MemryNote-win-Setup.exe on the same release tag (HEAD, then a streamed download into userData/installer-handoff/, folded into the downloading state) and verifies it with Get-AuthenticodeSignature: status Valid and a signer subject containing CN=Open Source Developer Kaan Karaca, or the file is deleted and never run. The install step then spawns a detached batch script (rendered by installer-handoff-script.ts, so its exact command lines are pinned by tests) that waits for the app's PID to exit, runs a copy of Uninstall MemryNote.exe /S _?=INSTALLDIR (no --updated, so the tolerant removal path runs, and the copy blocks until the old install, its shortcuts and its uninstall key are gone), runs MemryNote-win-Setup.exe --silent --verbose --log userData/logs/velopack-setup.log, and falls back to the already-downloaded NSIS installer if no Velopack Memrynote.exe exists afterwards, so the user is never left without an app. Every refusal falls back to the plain NSIS install: a release without the Velopack asset (info log only), a download failure (INSTALLER_HANDOFF_DOWNLOAD_FAILED through the updater error path) and a signature failure (INSTALLER_HANDOFF_UNVERIFIED). The next launch reads the marker: still the same version is the usual UPDATE_INSTALL_DID_NOT_APPLY with source: velopack-handoff; a new version running from the Velopack layout (..\Update.exe next to the app) is app_update_installed with action: migrated and source: velopack-handoff; a new version still running from an NSIS layout is INSTALLER_HANDOFF_DID_NOT_APPLY, meaning the update arrived but the migration did not. Memrynote.exe --cli migrate-installer path\to\MemryNote-win-Setup.exe drives the same verify-then-hand-off path against a local file (exit 0 when scheduled, 1 otherwise, Windows only), which is what the migrate job in .github/workflows/velopack-smoke.yml runs on a real NSIS install.

  • Expired GitHub signed asset URLs: GitHub serves a release asset by redirecting to a short-lived signed release-assets.githubusercontent.com URL. When the follow-up GET lands after that token expires, GitHub answers with the non-standard status 618 jwt:expired, and electron-updater has no retry on that path — builder-util-runtime's retryOnServerError is never called there, and its isServerError() covers 500-599, so it would not match a 618 anyway. One expired token therefore lost the whole update check (36 production exceptions across four releases, all in the check phase). The token is minted fresh on each redirect, so checkForUpdates() now retries twice, two seconds apart, on a 618 — or a 403 whose URL is the signed-asset host, host-gated so an unrelated 403 is never retried into a loop (isExpiredSignedAssetError in apps/desktop/src/main/updater-error-severity.ts). Only the attempt that still fails reaches the error handler: a recovered check does not flip the update surface to an error, does not advance the stuck-updater streak, and ships one app_log_recorded warn (update check hit an expired release-asset url, retrying) instead of an exception. A 618 that survives every retry is reported exactly as before — the severity classification above is unchanged.

  • Process lifecycle: the main process reports a child-process-gone fault with a composite type:reason:name error code (e.g. Utility:crashed:Embeddings). The worker label comes from Electron's details.name, not details.serviceName: Electron routes a fork's serviceNameoption to details.name, while details.serviceName holds the Mojo interface name — a constant (node.mojom.NodeService) that is identical for every utility fork and so cannot tell our workers apart. A fork that passes no serviceName option reports the default Node Utility Process. The exit status rides along in the line's exit_code field rather than inside the error code, so crashes still group by worker in PostHog Logs while the POSIX signal (11 SIGSEGV, 6 SIGABRT) stays visible. A utility worker's clean idle-shutdown (embeddings, image-processing, voice-model each exit after ~30s idle) is a lifecycle event, not a fault, so a clean-exit reason is skipped entirely — only a real fault produces an error event, mirroring the GPU crash guard. Note that child_process_gone is not throttled: the crash cadence is itself a diagnostic signal. When the owning module can resolve a lifecycle phase for the dead worker, the breadcrumb becomes child_process_gone_<phase> and the phase joins the exit status in the message (Embeddings utility process crashed (exit 6, idle_shutdown)). The error code stays phase-free: it is the Error Tracking fingerprint, so splitting it per phase would orphan the existing issue's history.

  • Embedding worker: phase is starting, in_flight, idle_shutdown, or idle. This is what separates a harmless teardown crash (idle_shutdown — the embedding was already delivered) from real user impact (in_flight — the user silently lost semantic-search indexing for that note).

    The phase is resolved by getEmbeddingWorkerCrashContext(details.name, details.reason) at the child-process-gone call site, not inside the worker's own exit handler. Electron's UtilityProcess exit event does not fire for a native crash — a SIGABRT out of the model runtime is neither a graceful exit nor the V8 FatalError the instance error event covers — so the bridge never learns its worker died. Production proved it: across 107 consecutive Utility:crashed:Embeddings events the bridge's own worker_exit_<phase> breadcrumb emitted zero, and so did embed_failed, while child-process-gone fired for all 107. Resolving the phase at the report that does arrive is what makes those events answerable. The worker_exit_<phase> breadcrumb is kept for the paths where exit does fire (a non-crash abnormal exit, a force-kill), and carries a different error code (EmbeddingWorkerExit) so the two are never confused.

    The report never arrives while the bridge still owns the worker. Resolving the phase from the live handle alone therefore answered null in 100% of production crashes (issue #1582): 76 events on 2026.817.1, none of them phase-suffixed. Every path that nulls the process handle either latches a teardown phase (stop() / reset()) or runs inside an exit handler that emits EmbeddingWorkerExit — and production has neither on that release. What remains is failProcess(), which forgets a worker that is still running (10s start timeout, fatal error). So the bridge keeps a last-worker record — pid, phase, how it was released, fork timestamp, model-cache state, stderr tail — bounded by a 60s TTL so an unrelated later crash cannot inherit it. failProcess() now also kills the worker it gives up on: the orphan was unreachable (every request goes through the process handle) yet kept a whole onnxruntime alive to abort later.

    Anything reading the phase from outside the exit handler must also survive the force-kill race: reset() latches idle_shutdown before killing, then nulls the process handle, so both the latch and the last-worker record are cleared when there is no process to kill — the latch now outlives the exit handler that used to clear it, and a stale one would make the next worker's crash read as a teardown it never had.

    The crash report also carries what the main process knew about that worker. Telemetry events accept at most one dimension (TelemetryDimensionsSchema) and log_action holds the phase, so the rest ships inside fields that already exist — no new event name, no new dimension key, no contract change, and therefore no sync-server deploy:

    FactWhere it shipsQuery it as
    lifecycle phasedimensions.log_actionlog_action (logs + events)
    platform exit statusmetrics.valueexit_code (logs) / value
    worker uptime at deathmetrics.durationMsduration_ms (events)
    crashes this sessionmetrics.retryCountretry_count (events)
    cached model sizemetrics.byteCountbyte_count (events)
    pid, reason, release, cache state, first-load vs reloaderror.message suffixmessage (logs)
    worker stderr tailerror.stackstack (logs)

    The message suffix is a bounded [reason=… pid=… uptime=…ms release=… cache=… cache_bytes=… load=… crashes=…] block appended only when a context was resolved, so every other child-process-gone family's message is byte-identical to before. The stderr tail is the closest thing to a stack trace this family can produce — the process that died is not the one reporting, so stack is otherwise always empty. It is captured into a bounded per-worker ring buffer, redacted on the device with the same salted redactText setup as Path A log shipping, capped at 2000 bytes, and every line is prefixed so it can never be mis-parsed back into a fabricated PostHog Error Tracking frame.

    Embedding generation failures emit embed_failed, throttled to one event per 5-minute window because a broken worker would otherwise fail once per note edit.

  • Desktop IPC envelopes: every { success: false } error envelope produced by the IPC layer (withErrorHandler / withDb) also emits an app_error_seen event, throttled in-memory to one event per action + error code per minute so an error loop can't flood the telemetry queue. The key must discriminate: keyed by error name alone and shared across all handlers, one benign recurring Error masked a genuine different Error from another handler for the whole window. The expected noVaultOpen envelope is not tracked, and the envelope's user-facing error string (which may contain note-derived text) never leaves the process — only the error code and redacted stack frames ship. Handlers that throw instead of returning an envelope report the same way: createHandler and createValidatedHandler (main/ipc/validate.ts) wrap the call, so canvas and calendar reads — which never pass through withDb / withErrorHandler and whose rethrow used to be their only trace — are now countable. A Zod validation failure is reported too (renderer↔main contract drift, not user error); only the ZodError name and stack ship, never the issue messages, which can echo input values. An expected condition is skipped before the throttle map, so a suppressed error can't claim the key and mask a real failure from the same handler. This throttle is main-side only: trackRendererError sends one app_error_seen per call, so a renderer loop calling one failing handler produces many renderer-sourced events against at most one main-sourced event per minute. Renderer and main counts for the same underlying failure are therefore not comparable — a large gap is the throttle, not a dropped main-side event.

  • Vault watcher: chokidar onError can burst per file (a permission-denied subtree), so watcher faults are sampled to one app_error_seen per minute (main/vault/watcher.ts). Every error still reaches the local log.

  • Unclaimed persistence tokens: a ((mention:…)) / ((date:…)) token the note-open normalize chain left as literal text, or a callout marker orphaned of its > prefix, is a block that will render broken with no error anywhere — the failure mode behind the #1843 round-trip epic, previously detectable only by a user emailing a screenshot. The renderer counts them after every normalize pass (renderer/…/content-area/unclaimed-token-telemetry.ts) and reports through app_error_seen as action: editor_unclaimed_token with errorCode: unclaimed_mention / unclaimed_date / unclaimed_callout_marker and the occurrence count in metrics.itemCount. First sighting emits immediately; after that, counts aggregate into at most one event per kind per minute, since the chain re-runs on every note open and remote update. Metric only — the user sees no toast or dialog, and no token content ever ships.

  • User-visible failures that previously left no trail: a renderer did-fail-load (the user is staring at a blank window) reports as DidFailLoad:<chromiumErrorCode> — the URL never leaves the process; a memry-file: protocol serve failure (a silently broken image/PDF/video embed) logs a MemryFile / serve_failed breadcrumb with the coded errno before answering 404; a window-close flush rejection (window refuses to close, edits possibly unsaved) reports as window_close_flush_failed; and quick capture reports QuickCapture / global_shortcut_register_failed when both the configured and the fallback global shortcut fail to register.

  • Expected conditions: some failures are normal states, not faults. They still surface to the UI as an error envelope, but the throw site marks them and error telemetry skips them, so they cannot drown real signal. Currently marked: an Ollama model-list fetch that is refused (ECONNREFUSED = Ollama is not running), and a calendar OAuth timeout (the user opened the consent screen and walked away). The suppression is deliberately narrow — a real Ollama misconfiguration (DNS failure, connection reset, or a bad HTTP status) is still reported.

  • Server errors: captureServerError pushes its redacted detail (operational message, stack, normalized path, error/status codes) as app="server" lines — level error for 5xx/unhandled, warn for handled 4xx.

  • Beyond route handlers: failures that never produce a failing HTTP response also reach PostHog Logs — scheduled cleanup-task failures (source="cron"), token-revoke failures during logout, Resend email send failures (RESEND_SEND_FAILED), and UserSyncState Durable Object alarm/websocket handler errors (source="user_sync_state_do", pushed directly via pushPostHogLogs since no route error handler ever sees them).

  • Attributes: app (desktop/server) and env are resource-level attributes shared by every line in a push; kind (error | log | report) and, when known, posthogDistinctId (the already-hashed identity, so PostHog can attribute the line to a person) are per-line attributes; level rides in severityText. Everything else — the JSON body under line in services/posthog-logs.ts — is the log body, not an attribute.

  • Retention: 14 days, per PostHog Logs' own retention policy.

  • Viewing: log lines are searchable in PostHog's Logs product, filterable on the service.name / deployment.environment / kind attributes and by distinct_id.

Canvas Rollout Panels (Grafana / Analytics Engine) ​

Panels for the spatial canvas rollout. Create these by hand in Grafana Cloud against the Analytics Engine dataset; they are recorded here so the rollout is reproducible and reviewable.

canvas_opened fires on every successful canvas load, so a tab-switch remount counts again. Read it as "canvas loads", not "distinct canvases opened".

The queries below are written against the real datapoint layout used by toDataPoint() / writeServerPoint() (apps/sync-server/src/services/telemetry.ts, apps/sync-server/src/services/analytics.ts): blob1 holds the event name and index1 holds the HMAC-hashed install id, against the dataset named in apps/sync-server/wrangler.toml (memry_product_telemetry_production / _staging / _dev). The column names are real; the exact SQL dialect accepted by the Analytics Engine SQL API through Grafana's Infinity datasource has not been verified by running these — treat them as a documented starting point, not verified panel queries, and adjust syntax as needed when building the panels.

PanelWhat it answersQuery
Canvas adoptionAre people turning it on and using it?SELECT toStartOfInterval(timestamp, INTERVAL '1' DAY) AS day, count() AS loads, uniq(index1) AS installs FROM memry_product_telemetry_production WHERE blob1 = 'canvas_opened' GROUP BY day ORDER BY day
Canvases createdIs creation growing or one-and-done?SELECT toStartOfInterval(timestamp, INTERVAL '1' DAY) AS day, count() AS created FROM memry_product_telemetry_production WHERE blob1 = 'canvas_created' GROUP BY day ORDER BY day
Conflict-copy rateIs last-write-wins hurting real users?SELECT toStartOfInterval(timestamp, INTERVAL '1' DAY) AS day, count() AS conflicts, uniq(index1) AS installs FROM memry_product_telemetry_production WHERE blob1 = 'canvas_sync_conflict_copy' GROUP BY day ORDER BY day
Oversized canvasesIs the size cap being hit?SELECT toStartOfInterval(timestamp, INTERVAL '1' DAY) AS day, count() AS blocked FROM memry_product_telemetry_production WHERE blob1 = 'canvas_too_large' GROUP BY day ORDER BY day
Unknown sync typesMixed-version tripwireSELECT toStartOfInterval(timestamp, INTERVAL '1' DAY) AS day, count() AS skipped FROM memry_product_telemetry_production WHERE blob1 = 'sync_skipped_unknown_type' GROUP BY day ORDER BY day

Swap memry_product_telemetry_production for _staging when checking a staging deploy. The event names are the stable part of these queries.

Go/no-go for flipping the canvas feature flag on by default is recorded in docs/superpowers/specs/2026-07-22-spatial-canvas-m7-rollout-design.md §7.

Server Configuration ​

Set these sync-server variables to enable PostHog capture and log shipping (unset in local dev, where both are a no-op):

bash
POSTHOG_KEY=...                          # wrangler secret (PostHog project token)
POSTHOG_HOST=https://us.i.posthog.com    # wrangler var (staging and production)

GITHUB_TOKEN is only used by the daily release download-count cron, and in practice it is required in staging and production. Without it the pull is unauthenticated and shares the 60-requests-per-hour-per-IP budget with every other Worker on the same Cloudflare egress address, which other tenants routinely exhaust before our one daily request arrives — GitHub then answers 403 and no release_download_count_snapshot event is emitted that day. A fine-grained PAT with public-repo read access is enough (this reads a public repo's Releases API and needs no write scope), and authenticated calls get 5,000/hour:

bash
wrangler secret put GITHUB_TOKEN --env staging
wrangler secret put GITHUB_TOKEN --env production

No workflow uploads it; it is set by hand, like every other Worker secret. It is deliberately not in the requiredSecrets fail-fast list — a missing token must not take the whole Worker down for a once-a-day measurement.

A failed pull throws and is reported as a release_download_counts cron failure rather than silently skewing the numbers, and the two failure shapes are separable. GitHub's own 403/429 raises GitHubReleasesRefusedError, which carries code: 'GITHUB_RELEASES_REFUSED' and the upstream status, so it reports as a handled 4xx logged at warn — expected upstream backpressure, with the message recording whether a token was in play. Anything else stays an unhandled 500. Nothing is corrupted either way: the throw happens before the D1 read/write, so the stored baseline is untouched and the next successful run emits the accumulated delta. The distortion is in the daily series (a zero, then a spike), not the running total.

Diagnostic Log Endpoints ​

Two additional endpoints feed the kind=log / kind=report streams. Both accept unauthenticated requests (no sign-in required), are rate-limited per user/IP, Zod-validated, and PostHog-Logs-only — neither writes a PostHog product event:

EndpointStreamRate limitPayload
POST /telemetry/logsPath A (kind=log)120 req / 60sDiagnosticLogBatchSchema — 1–50 redacted log lines
POST /diagnostics/reportPath B (kind=report)10 req / hourDiagnosticReportSchema — ≤200 redacted lines + a redacted device/sync snapshot + the triggering error

A malformed payload is rejected with 400 VALIDATION_ERROR (only the Zod path + issue code is logged, never values, same convention as /telemetry/batch). A valid payload always gets 202, including when POSTHOG_KEY is unset — the push inside pushPostHogLogs is a silent no-op in that case, so a dev build never error-spams.

/diagnostics/report attributes to an account when the desktop attaches a bearer (see Account identity); /telemetry/logs is still anonymous because the log shipper does not attach one yet. DiagnosticReportSchema.accountId is accepted for backward compatibility with older desktop builds but is deliberately ignored — a body field is client-asserted, and it would feed a distinct_id whose $identify merge is permanent, so identity comes only from the verified bearer. /telemetry/batch is unchanged by either endpoint.

Performance ​

trackTelemetry is debounced and batched. Calls during the first second of startup are deferred until after the vault is open so they never delay first paint. On the sync server, PostHog event capture and PostHog Logs pushes both run in waitUntil so neither can block the /telemetry/batch response.

Launch: note-readable mark ​

renderer/src/lib/launch-restore.ts stamps a performance.mark('memry:note-readable') when the note a launch restored has actually rendered its block tree, hooked to BlockNote's onEditorReady rather than to component mount, because mount only proves the CRDT binding started. scripts/launch-bench.mjs reads it over CDP and reports it as renderer_note_readable_ms.

The mark is deliberately narrow. It fires only for the note the launch restored, only once, and never for a note opened later in the session, so the metric cannot be inflated by ordinary navigation.

It is absent, rather than zero, whenever the launch restored something that is not a note. The restored tab is read from localStorage because the vault path is not available synchronously in the renderer, and the key with the newest savedAt wins. A machine with several vaults can therefore name a note from a vault the app is not opening: the speculative chunk prefetch is then wasted, and the mark never fires. Absent is the correct reading in that case, not a failure, but it does mean a median over launches must state how many runs carried the mark.

Login-shell PATH probe: failure classes ​

The probe runs $SHELL -ilc once per packaged launch to recover the user's real PATH, and reports one of four outcomes through login_shell_path_probe_failed:

  • marker_missing — the shell ran and exited cleanly but printed no marker, usually an rc chain that writes to stdout.
  • status_<n> — the shell ran and exited non-zero.
  • status_none — the probe hit its 3 s deadline and was killed.
  • spawn_error — the shell never started.

Two details are load-bearing and easy to undo by accident. The child's stdin is closed, not piped: an rc chain that reads stdin blocks forever on an open pipe, and every consumer awaiting whenLoginShellPathApplied() blocks with it. The deadline kills with SIGKILL, because an interactive shell ignores SIGTERM, so the polite signal would leave the promise pending for the whole session.

The unit suite injects a fake probe and therefore cannot catch either regression. A change to how the child process is spawned needs a real shell to verify.

Released under the GNU GPL v3.0.