Observability & Telemetry
Local logs for debugging plus a switchable, content-free telemetry stream for product metrics, launch diagnostics, and sanitized errors.
Logging
Use createLogger(scope) from electron-log everywhere in the desktop app:
import { createLogger } from '@/lib/logger'
const log = createLogger('Sync')
log.info('pull complete', { count, durationMs })
log.error('pull failed', err)- Never use
console.*. A pre-commit hook flags it. - Logs land in the OS-standard log directory and rotate automatically.
- Renderer and main process logs are separate files.
- Dev runs log at
debug; packaged installs are lowered toinfo(file) /warn(console) at startup based onapp.isPackaged, sinceNODE_ENVis undefined at runtime in packaged builds. - Important launch, renderer, and main-process errors are mirrored as telemetry events when product telemetry is enabled.
Choosing a Level
The redacted diagnostic log stream is floored at warn (see Error Logs in Grafana (Loki)), so warn and error are the levels an operator actually triages. Reserve them for conditions someone can act on, and log expected steady states at debug:
- Certificate pinning falling back to standard TLS when no pins are configured (
main/sync/certificate-pinning.ts) — a deliberate fallback, not a failure. sodium_mlock/sodium_munlockbeing absent in the WASM libsodium build (main/crypto/memory-lock.ts) — the expected state in Electron. An mlock/munlock call that is present but fails still logs atwarn.- Progress notes the embedding worker writes to stderr, such as transformers.js reporting an unknown content-length during a model download (
main/lib/embeddings.ts). A stderr chunk is only downgraded when every line in it matches a known-benign pattern, so a real failure interleaved with progress output still reacheserror. - Vector-clock bumps the
increment*ClockOfflinehelpers make while the sync runtime is down (packages/sync-client/src/offline-clock.ts) — the normal offline-edit path, one call per edited row (per changed field for tasks and projects). - An index rebuild triggered by a missing index DB (
emitIndexRecoveredinmain/vault/index.ts) —checkIndexHealthreportsmissingwhenindex.dbis not on disk, which is the expected first open of a vault: fresh install, newly linked device, or a deleted file. The index is a derived cache, so rebuilding it costs nothing but time. It is emitted atinfo, with the sameindex_recoveredaction anderrorCodeas before.corruptandmigration_failedstay atwarn— those are genuine data-corruption recovery.
A log payload is built before the transport decides whether to keep it, so a hot path pays for its arguments at every level. Keep these lines to identifiers: log the row id and the changed field names, not the resulting clock or field-clock objects.
Log Locations
| Platform | Path |
|---|---|
| macOS | ~/Library/Logs/memrynote/ |
| Windows | %USERPROFILE%/AppData/Roaming/memrynote/logs/ |
| Linux | ~/.config/memrynote/logs/ |
Older installs logged into an @memry/desktop directory (the raw package name). On startup the main process moves those files into the directories above (a name collision gains a legacy- prefix), then removes the emptied legacy directory. Dev profiles (MEMRY_DEVICE) log into a per-device memrynote-<device> directory instead.
The same applies to the app identity as a whole: production launches adopt the memrynote runtime app name and move userData (Application Support/@memry/desktop → …/memrynote, leaving a compatibility symlink for downgraded binaries and stored absolute paths). The macOS Safe Storage keychain item is copied to the new name so existing encrypted secrets keep decrypting; on Linux a populated safeStorage store keeps the install on the legacy identity (the keyring item cannot be carried over). See src/main/app-identity.ts.
Launch Phase Timeline
src/main/launch-timeline.ts stamps each startup milestone's offset (ms) from process start and emits them as one structured line, launch timeline, when the main window is revealed. Phases are recorded with recordLaunchPhase(phase), which also forwards each one as the per-phase app_launch_phase_completed telemetry event.
| Field | Meaning |
|---|---|
appReadyMs | Electron app.whenReady() startup work finished |
windowCreatedMs | main BrowserWindow constructed |
vaultOpenStartMs / vaultOpenReadyMs | vault restore started / reached isOpen |
rendererLoadedMs | renderer did-finish-load |
readyToShowMs | first ready-to-show (absent when it never fired) |
shownMs | window actually revealed |
reason | ready-to-show, fallback-timeout, or did-fail-load |
fallback | true when the 10s reveal fallback fired |
vaultOpenPending | vault open was still running at reveal — the prime suspect |
Vault-open timing stops at isOpen, not at the autoOpenLastVault() promise, which also waits on the first full sync. A launch onto the vault picker records no vault phase at all.
The line is logged at warn when the reveal came from the fallback or took ≥5s — only warn/ error records reach the diagnostic log sink — and at info otherwise, so healthy launches stay local instead of flooding the sink.
Post-Reveal Startup Queue
src/main/post-reveal.ts holds startup work the first frame does not depend on, so it cannot sit between window creation and the reveal and inflate shownMs. Register with onceWindowShown(name, task); the reveal calls schedulePostRevealTasks(), which drains the queue one second later. The delay is deliberate: ready-to-show means the renderer can paint, not that it has finished booting, and draining on the next tick would move main-thread contention rather than remove it.
name is the at-most-once key. A macOS dock reopen re-creates and re-reveals the main window, so a caller on that path registering again must not start a second copy. Work registered after the drain runs immediately. A task that throws or rejects is logged under the Startup scope and never stops the others. If the app begins quitting inside the delay the queue is left undrained, so a deferred task cannot re-arm a service the shutdown sequence has already torn down.
The updater is the first tenant: initializeUpdater() reconciles install health, reads its prefs, selects its backend and fires the startup-check update check, none of which is on the way to the first frame.
Login-Shell PATH Gate
A packaged app launched from Finder or the Dock inherits only the minimal system PATH, so which claude and which codex fail and Agent Chat greys out providers the user has installed. src/main/agent/cli/login-shell-path.ts recovers the real PATH by running the login shell, which sources the user's whole rc chain and costs anywhere from 180 ms to over a second.
That probe is not on the boot path. startLoginShellPathAugmentation() fires it at module scope and returns immediately; whenLoginShellPathApplied() settles once process.env.PATH has been merged. Anything that resolves an executable by name must await that gate rather than read PATH and hope. runBinaryCommand() in src/main/agent/cli/binary-detection.ts is where the agent CLIs do it, which covers which, the --version reads and, transitively, the per-turn spawns that run after detection. Headless --cli awaits it too, because it shells out before any window exists.
The probe still logs Augmented PATH from login shell for packaged launch when it changed PATH, and still reports login_shell_path_probe_failed with a failure class per breakage: status_<n> for a shell that exited non-zero, spawn_error for one that never started, marker_missing for output the parser could not read. A silent failure here is indistinguishable from an uninstalled CLI, so the breadcrumb is the only way to tell the two apart from a user's logs.
Telemetry
Telemetry is enabled by default in production builds and off by default in development builds. Users can turn it off via Settings → General → Privacy.
The build channel comes from MEMRY_ENV when set (dev/staging profiles) and otherwise from app.isPackaged, so packaged installs report production. A telemetry choice the user has already saved always wins over the channel default.
What Ships
Only enums and event metadata:
trackTelemetry('page_viewed', { surface: 'notes', action: 'viewed' })Recognized surfaces (TelemetrySurface in packages/contracts/telemetry-api):
app, home, onboarding, vault, notes, journal, tasks, inbox, calendar, search, graph, settings, sync, ai, voice, updater, canvas, projects, tags.
What Never Ships
- Note content
- Note titles
- Identifiers (note IDs, task IDs, project IDs)
- Search queries
- Tag names
- User file paths and note filenames
- Raw (unredacted) exception messages — a desktop error message can embed a note title or content
- Scraped page metadata — the
<title>, description or address of a URL the user clipped
The contract uses string-typed enums for surfaces and actions; arbitrary strings can't sneak through. The one exception is error diagnostics, which additionally ship a redacted stack trace of code locations and, optionally, a message — but only after the client has run it through redactText (packages/contracts/src/redact.ts), never the raw string — see Error Reporting.
The dimension allowlist
dimensions is the one free-form slot on an event, so its key namespace is closed. TELEMETRY_DIMENSION_KEYS in packages/contracts/src/telemetry-api.ts lists every key allowed to leave the device; sanitizeTelemetryDimensions keeps at most one entry, drops any key not on the list, and drops any value that fails the safe-value shape (no @, no ://, no slash, ≤ 64 chars, not UUID-shaped).
The shape check alone is not enough, which is why the allowlist exists: a scraped page title such as Divorce settlement calculator passes every one of those rules. Only an enumerable key namespace can tell a bounded enum from arbitrary user or page content.
Enforcement sits in createTelemetryClient's track() (apps/desktop/src/main/telemetry/client.ts) — the single chokepoint every event crosses before the durable queue and the network, whether it came from trackMainEvent, the renderer's IPC handler, or a direct runtime.track call. A new call site cannot opt out by skipping a helper. The renderer wrapper (src/renderer/src/lib/telemetry.ts) applies the same shared function so the UI drops rejected dimensions before the IPC hop.
Adding a key is the review gate. If the value cannot be enumerated ahead of time, it does not belong in a dimension — send a metric instead (a count, a duration, or a bucket label such as result_bucket).
The allowlist is deliberately not part of TelemetryDimensionsSchema. The sync-server validates /telemetry/batch with that schema and rejects a whole batch on one bad event, so narrowing it would make a newly deployed server 400 batches sent by already-shipped desktop builds. The allowlist is enforced on the client, where the data still is.
Failure detail on a failed request
An event that reports a failed HTTP request may also carry a failure object (TelemetryFailureDetailSchema) with three bounded fields:
| Field | Shape | PostHog property |
|---|---|---|
httpStatus | integer 100–599 | http_status |
serverCode | the server's SCREAMING_SNAKE error.code | server_code |
retryable | boolean — was the failure classified retryable | retryable |
This exists because sync_error used to ship one opaque server_error label covering 400, 403, 404, 409 and every 5xx alike (#1584): a permanent client-side contract bug and a transient edge 5xx were the same row, so no chart separated them and no alert threshold could be set. The label is unchanged — error_code still reads server_error — and these fields sit beside it, so error_code='server_error' AND http_status >= 500 now answers "is the backend down?" and retryable=false answers "how many of these will never succeed?".
An absent server_code on a 5xx is itself the signal: the response never reached the Worker's error handler, so it came from the edge and the backend has no record of it.
It is a field of its own rather than a dimension for two reasons. An event may carry at most one dimension and sync_error already spends it on transport; and unlike a dimension value, every field here is bounded by construction (a 3-digit range, an anchored enum-ish token, a boolean), so it can never become a free-text channel. sanitizeTelemetryFailure runs at the same createTelemetryClient.track() chokepoint as the dimension allowlist and drops any field that would fail validation — one bad field must never cost the whole batch.
The schema addition is optional and additive in both directions: an older desktop simply omits it, and a sync-server on older contracts strips the unknown key rather than rejecting the batch, so neither deploy order can 400 events from an already-shipped build.
On the desktop the fields are assembled in one place — apps/desktop/src/main/sync/sync-error-telemetry.ts — from what classifyError already computed, so a new sync_error call site cannot forward the category and forget the rest.
Because warn/error log records ship too (Path A, see What Ships and Error & Diagnostic Logs in PostHog), the same rule applies to log lines and thrown error messages in content-handling code: the inbox scraper (apps/desktop/src/main/inbox/metadata.ts) logs the clipped URL and scraped title at debug — below the ship floor — and keeps a content-free warn so the event stays countable.
Tracking Pattern
All telemetry calls are fire-and-forget — never await:
void trackTelemetry('onboarding_completed', {
surface: 'onboarding',
action: 'completed',
result: 'success'
})The void makes the call non-blocking and unfailable from the UI's point of view.
WebGL Availability
Every Sigma graph surface (graph-page, graph-canvas, local-graph-panel) probes hasWebGLSupport() (apps/desktop/src/renderer/src/lib/webgl-support.ts) before mounting. Chromium no longer falls back to SwiftShader for WebGL on its own, so main/index.ts appends --enable-unsafe-swiftshader before ready (enableSoftwareWebglFallback() in gpu-crash-guard.ts): a launch without GPU WebGL — hardware acceleration disabled by the crash guard, a blocklisted driver, Remote Desktop, a VM — renders the graph in software instead of showing the "Graph isn't available on this device" fallback. When even that fails, the probe records one app_log_recorded warn per session (source: WebGLSupport, log_action: webgl_unavailable) so affected installs can be counted in PostHog. The fallback is an expected device condition, never an app_error_seen exception.
Events Before the Runtime Exists
trackMainEvent used to no-op while getTelemetryRuntime() was still null, so anything that failed during early startup — the app-identity migration, the Safe Storage keychain carry-over, the GPU crash guard disabling hardware acceleration — was never reported. Early events are now buffered in memory (apps/desktop/src/main/telemetry/track.ts, capped at 100 events) and forwarded by drainEarlyMainEvents(), called from main/index.ts immediately after initializeTelemetryRuntime. Each event keeps its original occurredAt. The buffer is set to null once drained, so a runtime disposed during shutdown cannot quietly re-accumulate events nobody will ever flush.
registerMainDiagnostics() is registered before any other startup work for the same reason — the later call in the ready handler is an idempotent no-op — so a failure in the identity carry-over / pending-install window still produces a report.
Shutdown Drain
One flush sends at most TELEMETRY_BATCH_LIMIT events while a session can queue up to TELEMETRY_QUEUE_LIMIT, so the previous single flush('shutdown') silently dropped everything past the first batch. Shutdown now drains in bounded rounds (ceil(TELEMETRY_QUEUE_LIMIT / TELEMETRY_BATCH_LIMIT)), stopping at the first failed round so an offline quit never stalls the exit.
Event Categories
| Category | Events |
|---|---|
| Surface views | page_viewed — one per active-tab change; carries the tab type as objectType |
| App lifecycle | app_started, app_backgrounded, app_active_heartbeat, app_update_installed, deep_link_opened (coarse target only — open, billing_start, billing_complete, billing, oauth, pair, unknown; never the URL) |
| Onboarding | onboarding_started, onboarding_completed |
| Vault | vault_created, vault_opened |
| Notes | note_created, note_opened, note_updated (throttled 1/doc/5 min), note_deleted, note_imported, note_exported |
| Journal | journal_opened, journal_updated (throttled 1/doc/5 min) |
| Tasks | task_created, task_completed, task_reopened, task_updated, task_deleted |
| Projects | project_created, project_opened, project_updated, project_archived, project_deleted, project_item_linked |
| Tags | tag_created, tag_renamed, tag_deleted, tag_merged, tag_category_created |
| Canvas | canvas_created, canvas_opened, canvas_deleted, canvas_card_added, plus the rollout counters canvas_sync_conflict_copy, canvas_too_large, canvas_asset_uploaded, canvas_asset_dedup_hit, canvas_asset_gc_reaped |
| Inbox | inbox_captured, inbox_filed, inbox_archived, inbox_snoozed |
| Calendar | calendar_event_created, calendar_event_updated, calendar_event_deleted, calendar_google_connected, calendar_google_disconnected, calendar_google_sync_completed |
| Reminders | reminder_created (source = preset or custom; preset id in the value dimension), reminder_deleted — emitted by the reminder picker, so surface is the page it was opened from |
| Import | import_completed |
| Home | home_board_customized |
| Search | search_performed, search_result_opened (search_opened is defined but unused — it would duplicate command_palette_opened) |
| Command palette | command_palette_opened, search_result_opened (palette context) |
| Graph | graph_opened — on graph page mount |
| Voice | voice_recording_completed (duration + bytes), transcription_completed (success/failure + processing duration) |
| Settings | setting_changed — surface only, never the value |
| Agent chat | agent_chat_started, agent_chat_message_sent, ai_action_completed (turn result + duration) |
| Sync health | sync_enabled, sync_run_completed, sync_error (counts/status only, plus the failure detail) |
| Auth | signin_started, signin_succeeded |
| Diagnostics | app_log_recorded, app_error_seen, app_launch_phase_completed, app_crashed (see Crash & Unclean-Shutdown Detection) |
Crash & Unclean-Shutdown Detection
A hard crash — main-process abort, OOM kill, force quit — used to discard the in-memory telemetry queue, so the crash itself never shipped: the classic "it crashed and there are no logs" report. Two mechanisms fix that: a marker file that notices the crash, and a durable queue that keeps the resulting event alive long enough to send.
A marker file (apps/desktop/src/main/telemetry/crash-marker.ts):
session-marker.jsonis written intouserDataat startup with the session id,startedAt,lastAliveAt, and app version, then refreshed every 60s while the app is alive.clearCrashMarker()removes it once the shutdown cleanup chain completes.- A marker still present at the next launch means the previous session died uncleanly, and that launch emits
app_crashedon its behalf. Detection runs before the new session writes its own marker.
The marker's presence is the signal; its contents only enrich the event. An unparseable marker still reports the crash, just without the observed-uptime metric (metrics.durationMs, derived from lastAliveAt − startedAt). The previous session's app version ships as a prior_app_version dimension.
errorCode separates the failure modes:
errorCode | Meaning |
|---|---|
UNCLEAN_SHUTDOWN | no shutdown was attempted — hard crash, OOM kill, force quit |
SHUTDOWN_TIMEOUT_<STEP> | shutdown ran but its budget expired while <STEP> was still running |
SHUTDOWN_TIMEOUT | same, but the marker carries no step (written by an older build) |
SHUTDOWN_CLEANUP_FAILED | the cleanup chain rejected |
The last three are stamped by markShutdownFailure() immediately before the forced exit: the log line for that failure never flushes, but the marker survives to the next launch. Only the process that wrote a marker may remove one — a second instance that loses the single-instance lock shares userData and must not erase the primary's marker on its way out. Marker write failures are logged and swallowed; a read-only disk must never break startup.
The overrunning step rides in the errorCode rather than in a dimension, because an event ships at most one dimension and that slot already carries prior_app_version. The SHUTDOWN_TIMEOUT prefix is preserved so a query written against the old code still matches. Only a bounded kebab-case token is accepted from the marker; anything else degrades to the plain code.
app_crashed also carries an assembled message naming the shutdown failure, the overrunning step, the prior version, the observed uptime and whether the marker parsed — without it the Error Tracking issue was titled UNCLEAN_SHUTDOWN and held nothing else. The marker is a file on disk, so every string field is rejected outright unless it matches an enum-ish token (a character-substituted path still leaks its structure), and the assembled message is capped at 512 characters. That cap is not cosmetic: an over-length message fails TelemetryErrorDetailSchema at the sync-server, which rejects the whole batch with a 400, and the desktop client treats a 4xx as permanent — one corrupt marker field would otherwise discard up to 100 unrelated events on every launch until the marker cleared.
Shutdown Budget
before-quit runs its cleanup as an ordered list of named steps under one shared deadline (apps/desktop/src/main/shutdown-sequence.ts):
| Constant | Value | Role |
|---|---|---|
SHUTDOWN_BUDGET_MS | 8,000 ms | the whole graceful chain |
SHUTDOWN_LAST_CHANCE_MS | 1,500 ms | durability flush granted after the budget is gone |
SHUTDOWN_HARD_BACKSTOP_MS | 10,000 ms | timer outside the sequence; the process always exits by then |
The budget is derived from the bounded waits the chain contains, not guessed: 2,000 ms for the renderer flush handshake (windows in parallel) plus 3,000 ms for the voice, image-processing and embeddings utility stops — which run concurrently, so 3,000 ms together rather than 9,000 ms in a row — leaves 3,000 ms of headroom for the unbounded steps. Every step is also handed a cap() that clamps its own bounded wait to what is left of the shared deadline, so no set of waits can collectively overrun it.
Two rules keep a slow quit from becoming a lossy one:
- Order by durability. The window flush and
flushPendingWritebacks()run first, so a wedged teardown step behind them degrades a quit to slow rather than to lost edits. - Never force-exit with pending writes. When the budget expires, the step that overran is stamped into the marker, then
flushPendingWritebacks()andcloseAllDatabases()run inside the last-chance window beforeapp.exit(1).closeAllDatabases()matters because both SQLite files runsynchronous = NORMAL, which defers durability to the checkpoint thatclose()performs. The cleanup-error path does the same.
A quit where nothing is wedged still completes in milliseconds; these ceilings are only reached when a teardown step is genuinely stuck.
Durable Queues
The marker only detects the crash — app_crashed still has to survive long enough to be sent, and both telemetry queues flush on a 30s interval. A second hard crash inside that window would otherwise take the crash report with it, which is why both queues mirror to userData (telemetry/queue-store.ts):
| Queue | Mirror | Carries |
|---|---|---|
Event queue (telemetry/client.ts) | telemetry-event-queue.json | app_crashed, app_error_seen, all events |
Log-ship queue (telemetry/ship-queue.ts) | telemetry-log-queue.json | Path A redacted warn/error log lines |
Every enqueue reaches disk before it returns, and the mirror is rewritten after every flush, so what is on disk is what has not yet been accepted by the server. The next launch restores it and drains it; a drained batch is removed from the mirror, so nothing is sent twice.
Rules the mirror follows:
- Format is a journal: a
{"version":2}header line followed by one JSON item per line. An enqueue appends its own line rather than re-serialising the queue, which otherwise made each event cost a full rewrite of up to 500 objects — worst exactly during the error bursts and offline sessions that keep the queue pegged at its limit. The file is rewritten (compacted) on drains, on trims, and once the journal outgrows its bound, so it stays bounded. - Format changes cannot wedge startup. The previous
{"version":1,"items":[…]}format is still read, so an upgrading install keeps whatever its last session queued. A file whose version this build does not recognise is discarded rather than parsed, builds that predate the mirror never read it at all, and a build that predates the journal sees it as unparseable and discards it — so a downgrade costs one session's queue, never the launch. - Corruption is expected. The mirror is most likely to be truncated by exactly the crash it was written to survive, so an unparseable file is logged, deleted, and treated as empty.
- Write failures are non-fatal. A read-only or full disk costs durability, never logging or shutdown; the failure is logged once per streak rather than once per line.
- The limit applies to the restored set too (
TELEMETRY_QUEUE_LIMIT/SHIP_QUEUE_LIMIT), so a mirror written by a build with a larger limit cannot resurrect an unbounded queue. - Opting out deletes the mirror. Turning telemetry off clears the file, not just the in-memory queue, and a launch that starts with telemetry disabled discards the mirror instead of restoring it.
Restored events keep their own occurredAt; the batch envelope is stamped with the session that ships them, so an event resurrected from a dead session is attributed to the launch that sent it.
Native Crash Dumps
crashReporter.start({ uploadToServer: false }) runs before app.ready, so main, renderer, and utility processes all write minidumps for native crashes that no JS handler ever observes. The dumps stay in app.getPath('crashDumps') for the Path B diagnostic bundle the user submits deliberately. uploadToServer must stay false: a minidump is raw process memory — PostHog does not ingest minidumps, and there is no way to redact one, so uploading would breach the redaction model every other telemetry path is built around.
PostHog Event Capture
PostHog is the sole telemetry store (POSTHOG_HOST, https://us.i.posthog.com by default). The sync server transforms each accepted desktop event into a PostHog event (services/posthog-transform.ts) and posts it to the PostHog capture API's /batch/ endpoint (services/posthog.ts, capturePostHogEvents). Event names are preserved from the existing 50-event contract, with one rename: page_viewed → $pageview, which unlocks PostHog's native path-analysis and web-analytics views. Batch metadata (platform, arch, locale, app version, build channel, auth state, sync state, timezone offset) becomes person properties ($set); the event's own dimensions are flattened onto event properties first, then overwritten by server-derived keys (surface, action, environment, session_id, $session_id, $lib, $lib_version) so a client can never spoof a trusted key by naming a dimension after it. $session_id carries the same per-launch UUID as session_id: PostHog's session-scoped metrics and Error Tracking's sessions count read only $session_id, so without it every desktop event was session-less and all desktop traffic collapsed into a single session. $lib is memry-desktop ($lib_version is the app version), which is what Error Tracking's library filter matches on. Server-side business and error events (services/analytics.ts) post to the same capture API, tagged surface: 'server', so one PostHog project holds every event from both desktop and server.
Errors additionally become $exception events for PostHog Error Tracking (exceptionEvent in services/posthog-transform.ts), fingerprinted on our own errorCode when one is present so grouping follows the app's own error taxonomy rather than PostHog's pattern-hash default. They carry $session_id, $lib and $lib_version too, so an issue reports a real session count and is reachable through the library filter.
Generic constructor names do not group. toErrorCode falls back to the error's constructor name, so every bare new Error('...') reports errorCode: 'Error'. Pinning the fingerprint to that merged roughly thirty unrelated production failures — offline network errors, a Squirrel read-only-volume update failure, data.db failed PRAGMA quick_check, missing agent API keys — into one issue titled after whichever stack the first sample happened to carry (#2134). The transform therefore omits $exception_fingerprint for the built-in constructor names plus the UnknownError / StringError fallbacks (NON_DISCRIMINATING_ERROR_CODES), handing those back to PostHog's pattern hash so they split into real issues. A generic name still rides along as $exception_list[0].type; only the grouping key changes. Give a failure a typed code if you want it grouped as its own issue.
Three server-side guards keep this stream from flooding the PostHog quota:
- Warn-level log lines never reach Error Tracking. The desktop demotes expected failures to warn-level
app_log_recordedlines precisely so they stay out of Error Tracking (#1587);exceptionEventpromotes anapp_log_recordedevent only when itsactioniserror. Warn-level lines stay fully queryable as events and log records. - Legacy drop-tripwire noise is dropped at ingestion. Desktop versions before 2026.821 ship the
local_mutation_droppedtripwire once per polled row with no throttle or eligibility gate (#1579) — at peak 45% of the project's entire event volume, triple-billed as product event,$exceptionand log line.isLegacyMutationDropNoisedrops those events before all three sinks; fixed clients still forward their throttled diagnostic trickle. - Per-install hourly exception budget.
claimExceptionBudget(services/exception-budget.ts) caps$exceptionforwards at 60 per install per hour, reusing therate_limitstable. Only the$exceptionstream is trimmed — the product events and log lines for the same failures still forward, so a capped install stays diagnosable. The claim fails open on D1 errors: a flaky database costs extra PostHog events, never a swallowed crash report.
Stack frames
Error Tracking renders code locations only from $exception_list[].stacktrace — it never parses the exception's value. The desktop sends its stack as redacted text (that is the shape the client-side frame filter and redaction produce), so parseStackFrames in services/posthog-transform.ts turns each at fn (file:line:col) line back into a raw frame:
"stacktrace": { "type": "raw", "frames": [
{ "platform": "custom", "lang": "javascript", "function": "push",
"filename": "~/app/sync.ts", "lineno": 12, "colno": 5,
"resolved": true, "in_app": true }
] }Four rules that are load-bearing:
- Reversed. PostHog treats the last frame as the throw site; a JS stack string is innermost-first. The cap of 50 frames is applied before reversing, so a deep stack loses its outermost callers rather than the frame that actually failed.
platform: 'custom'. Claimingweb:javascriptenters PostHog's symbolification path, which needs uploaded source maps and a per-frame chunk id. We ship neither, so it would resolve to nothing;customframes render verbatim.in_app: falsefornode:*,internal/*,node_modulesandelectron/js2cframes. The UI hides non-in-app frames by default, falling back to showing all of them when an exception has none — so a fully-vendor stack is never blank.- Omitted, not empty. An exception with no parsable frame carries no
stacktracekey at all.frames: []would claim we resolved a stack and found nothing; utility-process crashes and log-derived errors genuinely have none.
value holds the redacted message alone — it is the issue title, and the stack belongs in frames. A React component stack is promoted to frames when there is no JS stack, and always ships intact as $exception_component_stack.
When there is no message the transform falls back to the error code, which makes the issue title identical to the code and tells an engineer nothing new. That fallback sets exception_message_missing: true on the event, so a message-less reporting site is countable on a dashboard instead of looking healthy:
SELECT properties.$exception_fingerprint AS fp, count() c, uniq(distinct_id) u
FROM events
WHERE event = '$exception' AND properties.exception_message_missing
AND timestamp > now() - INTERVAL 7 DAY
GROUP BY fp ORDER BY u DESCErrors that carry no JS stack by construction — child-process-gone for a crashed utility worker, where the process that died is not the one reporting — instead carry a synthesized message naming the worker, reason and exit status, so their issue page is not blank.
This transform is entirely server-side: it applies to batches from already-installed desktop versions as soon as the sync-server deploys.
No raw identifiers are stored: the install ID is HMAC-hashed server-side (TELEMETRY_HMAC_KEY, hashTelemetryId) and used as the PostHog distinct_id; server-side user_id/device_id/ vault_id are hashed the same way before they ride along as event properties.
Account identity
The desktop attaches its access token to /telemetry/batch and /diagnostics/report as an optional bearer. Neither route runs the auth middleware: resolveTelemetryAccountHash verifies the JWT if one is present and returns undefined for a missing, malformed or expired token, so telemetry is never rejected for auth reasons — that batch simply reports anonymously against its install hash.
The resolved account id is HMAC-hashed before it can become a distinct_id, exactly like the install ID. TransformContext names the field accountHash, and resolveDistinctId shape-checks it against hashTelemetryId's output (64 lowercase hex chars); anything else — most plausibly a raw account id — degrades to the install hash rather than reaching PostHog. This is deliberately strict: a PostHog $identify merge is permanent and cannot be undone or re-keyed, so a raw account id that reached a person profile could not be removed afterwards.
When a batch resolves to an account, a $identify event aliases the anonymous install person onto the account person. It fires once per app session, guarded by the telemetry_identify_sessions D1 table (claimIdentifySession, migration 0003, swept by the cron cleanup after 24h). Without the guard, the desktop's ~30s flush cadence would emit one identified event per batch. The guard fails open: a D1 error emits $identify anyway (idempotent in PostHog) rather than leaving the install unlinked.
Diagnostic reports resolve identity through the same resolveDistinctId path as events and logs, so a report lands on the same person profile as the events around it.
Usage segmentation
Two dimensions exist so "how many people use MemryNote" can be split without identifying anyone.
auth_state (anonymous | signed_in | signed_out) comes straight from the batch and is written as both an event property and a person property. The event property is the one to break down on: a person property holds the latest value, which answers "is this install signed in now" rather than "had yesterday's active users ever signed up".
plan and plan_status are person properties only, read from sync_entitlements by resolveTelemetryPlan — never sent by the client, which must not be trusted with a dimension like free vs pro. They resolve once per app session, behind the same claimIdentifySession claim that gates $identify, so the lookup costs one D1 row per session rather than one per ~30s batch. Every other batch omits both keys entirely rather than sending them as null, which would wipe what the session's first batch wrote. The status travels with the plan so a canceled pro cannot be counted as a paying user. resolveTelemetryPlan fails closed to undefined: a token whose account row is gone costs one person property, never the whole batch.
resolveTelemetryAccount returns the raw userId alongside accountHash purely so this lookup can read the account's own rows. It stays inside the worker — accountHash remains the only identity that reaches PostHog.
Note that an install which merely runs in the background still produces telemetry, so counting unique persons over "all events" measures installs that were running, not people who used the app. app_active_heartbeat (emitted only while a window is focused) or a real product event is the honest signal for the latter.
Known limitation: telemetry identity is verified but not revocation-checked. A revoked device's still-unexpired access token (≤15 min) can attribute telemetry until it lapses. Telemetry is not an authorization decision, so a per-batch device lookup is not worth the D1 read.
Environments are separated by an environment property on every event inside one PostHog project, not by separate projects.
Additional events in the same pipeline:
| Event | Source |
|---|---|
app_launch_phase_completed | Electron main/renderer startup milestones |
app_log_recorded | Sanitized desktop diagnostic breadcrumbs |
app_error_seen | Renderer, React boundary, and main errors |
server_error_seen | Sync-server request/background failures |
server_log_recorded | Structured sync-server diagnostic logs |
release_download_count_snapshot | Daily GitHub Releases download-count pull (not a user event) |
Batch Validation & Retry
Each /telemetry/batch payload is schema-validated on the sync server. A malformed batch is rejected with 400 VALIDATION_ERROR. The server logs the failing field paths (Zod path + issue code only — never the field values, which may hold the raw identifiers the schema is designed to strip) so rejections are diagnosable without leaking data.
The desktop client treats a permanent 4xx (any 4xx except 429) as unrecoverable and drops that batch, so one malformed event cannot wedge the queue head and replay the same rejected batch on every flush. Transient failures — 5xx, 429, and network errors — leave the batch queued for a later retry.
Token Lookup & Request Deadlines
The batch client asks the token manager for an access token only when the install is signed_in; an anonymous or signed_out install ships every batch without touching the secret store. Looking the token up regardless cost an OS keychain round-trip per flush, and on a machine whose keychain hangs (see Cryptography) that round-trip wedged the whole pipeline: events stopped while the log shipper — which never asks for a token — kept going. That "logs but no events" split is the diagnostic signature; the diagnostic report upload (sendIncidentReport) sat behind the same lookup and showed as Sending… forever.
Every net.fetch on the telemetry path — event batches, log shipping, incident reports — goes through boundedNetFetch (telemetry/bounded-net-fetch.ts) with a 30 s abort deadline. A request still in flight when Chromium's network service dies never settles on its own, and a flush loop awaiting it would otherwise stall for the rest of the run with every later batch queued behind it.
The keychain side of that hang is bounded twice over: OS keychain calls are serialized through one process-wide single-flight queue with a run-wide unavailability latch, so a wedged Secret Service can burn at most one libuv threadpool thread instead of all four, and the main process raises UV_THREADPOOL_SIZE to 16 before anything can use the pool. See Cryptography.
Autosave Event Throttling
note_updated and journal_updated events fired by the autosave path are throttled to at most one emission per document per 5-minute window (in-memory, resets on restart). This prevents high-frequency editor flushes from inflating event counts.
Body edits reach the note through the CRDT provider rather than the notes UPDATE IPC, so typing never registered as usage at all. trackNoteBodyEditThrottled (telemetry/diagnostics.ts, called from ipc/crdt-handlers.ts) now emits note_updated with source: 'editor_body' on the samenote_updated:<noteId> throttle key the UPDATE handler uses, so metadata saves and body edits share one 5-minute window per note. Only the throttle key ever sees the note id; the event itself carries no identifier.
The shared throttle map (telemetry/throttle.ts) is bounded to 1000 keys. Because the keys are per-document (note_updated:<noteId>, journal_updated:<date>, and the CRDT writeback keys), exceeding that inside one window is ordinary operation — a vault import or a writeback pass over a large vault does it — so the cap is enforced rather than advisory: keys whose window has elapsed are swept first, and if nothing has expired the oldest-inserted keys are dropped anyway. Dropping a key only forfeits its throttle, never an event.
Each entry records the window it was written under, and the sweep judges expiry per entry rather than by the window of whichever call happened to cross the cap. Callers do not share one window — the Google Calendar sync runner throttles on 60 seconds while the autosave keys use the 5-minute default — so a short-window caller must not be able to expire a still-live 5-minute entry and make note_updated re-emit early.
Release Download Counts
Downloads happen on GitHub Releases, where PostHog cannot see them — the landing site's download-click event measures intent, not a download. A daily cron on the sync server (services/release-downloads.ts, run from the scheduled handler at 04:00 UTC) reads GET /repos/memrynote/memry/releases and emits one release_download_count_snapshot event per asset.
It is a metric snapshot, not a user event. One emission per poll per asset, from a fixed service distinct_id — nobody downloaded anything at the moment it fired. Person and user counts on it collapse to the service identity and mean nothing, and it must never be used as a funnel or conversion step next to genuine user events like landing_download_click or app_started. The name says so: it was release_asset_downloaded until 2026-09-10, which read as a user action and was mistaken for one. The old name is marked deprecated and hidden in PostHog data management; events emitted before the rename still carry it, so any query spanning the cutover must ask for both names.
assets[].download_count is cumulative per asset. Emitting it raw would produce a monotonically increasing counter that is useless as an event stream — and it fails silently, producing meaningless numbers rather than an error. The last total seen per asset is therefore stored in D1 (release_download_counts, migration 0003) and only the delta is emitted:
- The first run for an asset seeds its row and emits nothing; a cumulative counter carries no meaningful delta until it has a baseline.
- A total that went down — GitHub recounting, or a replaced asset — reseeds the baseline rather than emitting a negative delta.
- The store is written before the events are captured. A D1 failure then throws, the cron reports it, and the untouched baseline makes the next run emit the full delta. Emitting first would double-count that delta after a failed write.
The pull rides its own cron entry (crons = ["0 */6 * * *", "0 4 * * *"]) so it runs once a day while the cleanup sweep keeps its 6-hourly cadence. The daily entry deliberately avoids the 6-hourly times — colliding entries collapse into one invocation.
Events are not person-scoped: an anonymous downloader has no identity to key on, so distinct_id is a fixed memry_releases_<environment>. Staging and production both poll the same public repo, so — as everywhere else — an insight that does not filter environment blends them.
| Property | Meaning |
|---|---|
release_tag | Release the asset belongs to (v2026-08-06) |
asset_name | Published filename |
platform | macos / windows / linux / unknown, derived from the filename |
asset_kind | installer, update_metadata, or update_package (see below) |
downloads | The delta — sum this, never cumulative_downloads |
cumulative_downloads | Total GitHub reported at pull time, for context only |
asset_kind is load-bearing. Some assets are only ever fetched by the auto-updater. The landing site never links to them (apps/landing/api/download.ts), so every count on one is an installed app updating itself, and counting it as a download swamps the number that matters:
asset_kind | Assets |
|---|---|
update_metadata | Polled on every update check: electron-builder latest*.yml and .blockmap, Velopack releases.<channel>.json and legacy RELEASES |
update_package | Downloaded when an update applies: Velopack .nupkg, and the macOS .zip Squirrel.Mac installs from |
installer | Everything else: .dmg, -setup.exe, MemryNote-win-Setup.exe, .AppImage, .deb, the portable -win.zip |
Filter to asset_kind = 'installer' for real downloads. Two caveats remain. -setup.exe and .AppImage are also fetched by electron-updater on NSIS and AppImage installs, and the filename cannot tell the two apart, so installer still carries some update traffic. And until 2026-09-28 the Velopack feed, .nupkg, and the macOS .zip were all labelled installer (Velopack assets also platform: unknown for RELEASES and .nupkg). releases.win.json alone put ~7,000 phantom Windows "downloads" into the week of 2026-09-21. Queries spanning that cutover must classify by asset_name, not asset_kind.
Downloads cannot be joined to activation. An anonymous downloader and a desktop install share no key. The funnel only works for people who sign up on the landing site and sign in on the desktop, where the identity merge puts both on one person. This is a limitation to state plainly, not to engineer around.
Landing Site Telemetry
The marketing site (apps/landing) runs analytics browser-side via posthog-js (apps/landing/src/lib/analytics.ts), ingesting through PostHog's reverse-proxy subdomain (https://e.memrynote.com) — session replay cannot be routed through the sync server, so landing traffic does not go through /telemetry/batch. The old POST /telemetry/web endpoint and its LandingTelemetryBatchSchema contract are gone.
- Client:
init()lazily configuresposthog-jsonce per page load, keyed onVITE_POSTHOG_KEYwithapi_host=VITE_POSTHOG_HOST(defaulting tohttps://e.memrynote.com),person_profiles: 'identified_only', and masked session recording (session_recording: { maskAllInputs: true }). It no-ops with nowindow(SSR/prerender) or no key configured.trackLandingPageViewfires PostHog's native$pageview;trackLandingEventfires one of a fixed set oflanding_*event names. - Payload: pages and targets are path-only — query strings and hashes are stripped client-side (
stripQueryAndHash) before either ever leaves the browser. UTM params (utm_source/medium/campaign/content/term) are read from the query string, trimmed, and capped at 120 characters. - Environment:
environmentis registered once viaposthog.register— Vercel'sVITE_VERCEL_ENVwhen present, otherwise a production/development split on Vite's buildMODE— so landing traffic is filterable apart from desktop/server events in the same PostHog project. - Scanner noise: a
before_sendfilter drops one$exceptionfingerprint —Object Not Found Matching Id:N, MethodName:update, ParamCount:4. That is Microsoft Office / Outlook SafeLinks pre-fetching a link out of an email, injecting its own scanner into the page, and then losing its own object handle;MethodName/ParamCountare a COM bridge's idioms and appear nowhere in this repo. It arrives in same-day bursts from a handful of readers a few times a quarter and has no type and no usable stack, so it is pure noise in the error list. The match is deliberately narrow — only that exact COM signature. A broad "drop every non-Error rejection" rule would hide real bugs.
Session replay coverage
Landing replay is not sampled. In PostHog project 412311, session_recording_opt_in is on and session_recording_sample_rate, session_recording_minimum_duration_milliseconds, session_recording_linked_flag, and the URL / event trigger configs are all unset — every session is eligible. The client passes only masking options (see #860 for the cookie-consent posture: no banner, no consent gate, so nothing client-side suppresses the recorder either).
$sdk_debug_rrweb_start_attempted is a per-event property, not a per-session one. It reports the recorder's state at the moment that single event was captured. analytics.ts captures the first $pageview immediately after posthog.init(), and the recorder only starts once PostHog's remote-config response comes back — so in almost every session the first event carries start_attempted = false (or no value at all) even when the recording starts a few hundred milliseconds later. Grouping sessions by their first $pageview therefore undercounts coverage badly; it is the wrong shape of query, not a coverage number.
Aggregate over the whole session instead — max(properties.$sdk_debug_rrweb_start_attempted = true) grouped by $session_id — or just count raw_session_replay_events. Measured over the 7 days to 2026-09-10 ($lib = 'web'): 350 of 408 sessions (85.8%) started a recording, and raw_session_replay_events holds stored replays for 365 sessions. The same window read first-$pageview-only says 49%.
Of the 58 sessions that never started a recording, 46 captured exactly one event — a bounce that ended before the remote-config round trip completed. That residual is structural: the posthog-js bundle is imported lazily (it is the largest dependency on the site; see the comment on load in analytics.ts), so init, remote config, and recorder start all happen after first paint. Eagerly importing it would shrink the gap at the cost of first paint, which is a trade the lazy import deliberately makes in the other direction.
Practical consequence for anyone reasoning about the landing funnel: replay covers ~86% of sessions with no sampling bias, but the uncovered slice is skewed towards the shortest sessions. Do not read replay as evidence about immediate bounces.
Error Reporting
Desktop error reporting follows the same product telemetry setting. Each captured error ships stable metadata — process area, component/source, action, phase, and the error's code (errorCode, a typed code where one exists — see Vault File Errors) — plus a redacted stack trace and, for React boundaries, the component stack.
errorCode prefers a typed code the error carries over its class name: a richer telemetryCode (NoteError's note error code plus the originating errno, see Vault File Errors), then a plain .code (better-sqlite3's error.code, a Node system code), walking the cause/AggregateError.errors chain to a bounded depth so a fetch failed TypeError still surfaces the underlying ECONNREFUSED. A note write failure therefore reports NOTE_WRITE_FAILED:EBUSY rather than collapsing every note fault to NoteError, and a locked database reports SQLITE_BUSY rather than an un-triageable SqliteError. A code is only trusted when it looks like an enum token (^[A-Za-z][A-Za-z0-9_.:-]{0,63}$); anything else — a path, an email, a URL, free-form prose — is rejected outright and the class name is used instead, because a character-substituted path (_Users_kaan_secret.md) still leaks its structure. The class name itself still passes through the safe-token rules (no @, ://, /, \, ≤64 chars).
An unhandled rejection can carry any value as its reason — a string, a plain object, or a cross-realm Error that fails instanceof Error — and those carry no stack, which previously landed in Loki as an unactionable bare Error with an empty stack. Reasons are normalized before reporting: a real Error passes through, a cross-realm error's own frames are adopted, and anything else gets a stack synthesized at the handler plus a code naming the reason's type (Rejection_string, Rejection_Object, Rejection_undefined). The reason's message rides along redacted, not dropped: it goes through the same redactText pass as any other error message before it leaves the device. Omitting it made every Rejection_* row in Error Tracking an issue titled after its own error code, with nothing inside to triage. A reason that crossed a structured-clone or IPC boundary keeps its .name but loses both its stack and its constructor; that name is preferred over the constructor name, so it reports Rejection_TypeError rather than collapsing to Rejection_Error. When the code is a Rejection_* name the stack is the handler's own frames, not the fault's — the code is the actionable part.
A window error does not always carry an error object: cross-origin scripts and some Chromium failure paths report only a message and a source location, which previously landed as StringError with an empty stack and nothing to triage. The error class is recovered from the message's leading token (Uncaught TypeError: … → TypeError, subject to the same enum-token rule) and the filename/lineno/colno are rebuilt into a stack frame, so the code location survives the same frame filter and redaction as a real stack. The message text ships too, redacted on the device by the same pass — without it a cross-origin failure was a WindowError issue titled WindowError.
Because a rejection reason or event.error can be any value — including a Proxy whose traps throw or an object with throwing getters — every property read in this path (including instanceof, which can trap getPrototypeOf) is individually guarded. A hostile value can no longer throw out of the diagnostics handler and destroy the report being built.
The free-form exception message was historically never sent at all, since on the desktop it can embed a note title, filename, or content. buildErrorDetail now ships it (TelemetryErrorDetailSchema.message, optional, capped at 512) after running it through redactText (packages/contracts/src/redact.ts) — the server re-runs redaction in mask mode as a backstop. This is what makes an issue readable: without a message PostHog titles every issue with the bare error code, which is how a whole family of production issues came to read StringError / StringError. redactText strips known-sensitive shapes (secrets, tokens, emails, ids, home-directory paths, content-file basenames) rather than proving the remaining prose is note-free, so this is narrower than the earlier all-or-nothing "no message field" guarantee. The stack is separately reduced to code-location frames only — the leading Name: message header line is stripped — so a crash's location shows up as, for example, TypeError at pushRecords (…/sync-engine.js:120). Frame file paths are app source/bundle locations (not user files); any home-directory prefix (/Users/<name>, C:\Users\<name>) is rewritten to ~, and emails, UUIDs, JWTs, and bearer tokens are scrubbed from anything that ships.
IPC Error Throttling
The action an IPC error reports is the channel it was registered on (notes:create), not the handler's function name. Handlers are registered as ipcMain.handle(Channel, createValidatedHandler(Schema, async (input) => …)), and an arrow passed straight in as an argument has an empty name — so every inline handler in the app used to collapse into one literal action, validated_handler, and a schema rejection could not be attributed to a channel from the wire data alone (the captured stack names only the bundled wrapper, and Zod strips its own frames). installIpcChannelLabels (main/ipc/lib/ipc-channel-labels.ts) records the pairing once at the ipcMain.handle boundary — the only place that knows both halves — and registerAllHandlers calls it before the first registration. A handler registered without it keeps the old generic label.
Every IPC envelope error becomes a telemetry event, so a handler stuck in a failure loop could flood the queue. trackIpcError (main/ipc/validate.ts) therefore emits at most one event per action:errorCode per 60-second window, in-memory and reset on restart. The key includes the action because keying on the error name alone let one handler's benign recurring Error mask a genuine Error from an unrelated handler for the whole window. An expected condition (Ollama not running, an abandoned OAuth flow) is skipped before the key is claimed, for the same reason — otherwise the suppressed error would keep refreshing a key it never reports on.
That key set is bounded to 1000 entries. Once past the cap, keys whose window has elapsed are swept; if a burst of previously unseen codes fills the map inside a single window with nothing to expire, the oldest-inserted keys are dropped instead. Dropping a key only forfeits its throttle — the next error for it is reported rather than lost.
Vault File Errors
A class name alone is often too coarse to act on: every failed note save reported NoteError, which cannot tell an antivirus or cloud-sync file lock apart from a full disk. NoteError therefore carries the originating fs error as its cause, and reports a composite errorCode of its note error code plus the errno — for example NOTE_WRITE_FAILED:EBUSY (locked) versus NOTE_WRITE_FAILED:ENOSPC (out of space). The errno is admitted by a strict allowlist (/^E[A-Z0-9]+$/), so the vault file path is never part of the code — paths are user data and stay out of telemetry, as above.
Writes to a locked file are retried a bounded number of times before failing (see withTransientFsRetry in main/vault/file-ops.ts). Each retry is written to the local log with its errno and attempt number — again never the path — so a slow or failed save is explainable from a user's log file even when telemetry is switched off.
Sync-server error reporting is server-side. Because the sync server is end-to-end-blind (it only ever holds ciphertext), its own error strings are operational: the redacted message and stack ship to Cloudflare's own first-party Workers console logs (with the raw userId/deviceId/vaultId attached, since that sink is trusted) and to PostHog Logs (with those same ids HMAC-hashed) — never to the PostHog event itself, which only gets the coded server_error_seen event with no message or stack; this is what makes sync failures debuggable without handing ciphertext-adjacent detail to a third party. The server_error_seen event itself carries no user, device, or vault id either — its distinct_id is a fixed memry_server_<environment> value, so the event alone cannot be traced to an account. Only the PostHog log record carries the HMAC-hashed userId/deviceId/vaultId (when the caller has them), which is what lets a failure be correlated to an account inside Logs without exposing the raw id. Dynamic path segments and query strings are normalized away. Expected handled 4xx responses (e.g. SYNC_PAYMENT_REQUIRED) are still counted as server_error_seen but are logged at warn rather than error level, keeping real failures distinguishable from expected noise.
Server Business Events
Server-side product events — user_signed_up, user_logged_in, device_registered, vault_registered, vault_deleted, and the Paddle subscription events — go through captureBusinessEvent in apps/sync-server/src/services/analytics.ts. Unlike server_error_seen, these have a real actor, so their distinct_id is the HMAC hash of the acting user's id (hashTelemetryId, same key and shape as the install hash). The raw id never reaches PostHog, and the hash is also kept in the user_id property so queries written against that property keep working. If a business event ever has no acting user, it falls back to the fixed memry_server_<environment> id rather than an empty one.
Before this, every server business event shared that fixed id, so uniq(person_id) over any of them evaluated to 1 and every unique-user funnel or retention metric touching a server event was wrong. Events emitted before the fix keep the old distinct_id and person history is not backfilled, so person-level metrics over server events are only correct from the fix forward. Event counts were unaffected in either period.
Error & Diagnostic Logs in PostHog
Error events also become searchable log lines in PostHog Logs. This replaced a self-hosted Grafana + Loki instance the sync server used to push to; PostHog Logs now carries the diagnostic detail (stacks, operational messages) that a PostHog event deliberately omits.
Transport: the sync server posts log lines to PostHog Logs' plain OTLP-JSON receiver (
{POSTHOG_HOST}/i/v1/logs,services/posthog-logs.ts,pushPostHogLogs) — no OpenTelemetry SDK is used — authenticated with the PostHog project token (POSTHOG_KEY) as a bearer. Pushes are fire-and-forget inwaitUntil: a missing key (local dev) is a silent no-op, and a failed push can never affect request handling. Records are grouped into oneresourceLogsentry perapp(desktop/server) soservice.nameanddeployment.environmentstay resource-level attributes rather than being duplicated onto every line.Desktop errors:
/telemetry/batchevents carrying anerrorCodeorerrordetail are forwarded asapp="desktop"lines containing the event name, error code, surface/action/source, app version, platform, a redacted message (TelemetryErrorDetailSchema.messageis optional and only accepted after the client has run it throughredactText; the server re-runs redaction as a backstop) and redacted stack frames,log_action— the operational breadcrumb that keeps log-type error events (which carry no stack of their own) identifiable — andexit_code, the platform exit status for process-lifecycle events (empty string when absent, since exit code0is itself meaningful). A failed request additionally contributeshttp_status,server_codeandretryable(see Failure detail on a failed request) — also empty string when absent, becauseretryable: falseis a real answer and must not read as "not reported".Redacted diagnostic logs (
kind=log, Path A, always-on): a main-process electron-log transport (apps/desktop/src/main/telemetry/log-ship.ts, installed once frommain/index.ts— never fromlogger.ts, which must stay electron-free for worker bundling) intercepts everywarn/errorrecord and redacts it viaredactLogLine(packages/contracts/src/redact.ts) before anything leaves the device, using a per-install salt (diagnosticsSalt, persisted intelemetry.json) and the active vault root. Redacted lines batch (queue 500 / batch 50 / flush 30s, drop-4xx-except-429) toPOST /telemetry/logs, gated on the telemetry toggle (getTelemetryRuntime().getSettings().enabled) and disabled outright in dev builds. Repeated identicallevel|scope|messagelines within a 3s window are throttled into one line with arepeatCountfield. The Path B ring those lines also feed is a fixed-capacity circular buffer (200 slots, 5-minute window) that stores each line's epoch ms at push time and evicts oldest-first through a head index, so onewarn/errorcosts O(1) instead of re-parsing every retained timestamp — the error path has to stay cheap when something is looping. A record's arguments are flattened byparseRecord: the first string wins themessageslot, plain objects merge into the fields, and anErrorcontributeserrorNamepluserrorMessage. That last part matters —logger.error('updater error', err)used to ship{"errorName":"Error"}and nothing else, because the label had already claimed the message slot and the Error's own message was dropped. Worker processes (embeddings, image processing, voice transcription) forward their ownwarn/errorrecords to main overprocess.parentPort(apps/desktop/src/main/lib/log-forward.ts, electron-free) for the same redaction + ship pass, taggedorigin: 'worker'andworkerName. The server re-runsredactLogLinein mask mode (no salt) as defense-in-depth before writing to PostHog Logs (desktopLogRecordinservices/posthog-logs.ts) — the client-side redaction is primary; the server pass is a second net, not the source of truth.Incident reports (
kind=report, Path B): on a real error, the app offers a one-time "Send diagnostic report" action (the tab error boundary, IPC-error toasts, and a Settings entry — available independent of the telemetry toggle). The tab error boundary sends the report automatically, with no dialog and no button, when Settings > General > Privacy > "Automatically Send Error Reports" is on. That flag isautoSendDiagnosticsintelemetry.json(telemetry:setAutoSendDiagnostics); a missing key means on, so fresh installs and upgrades default to on and only an explicitfalseopts out. When it is off, or the automatic send fails, the boundary shows the Send button and the consent dialog instead. Thediagnostics:previewReport/diagnostics:sendReportIPC calls build aDiagnosticReportvia the same purebuildIncidentReportfunction: a generatedincidentId(MEMRY-XXXXXXXX, random base32), the last ≤200 redacted lines from the Path A ring buffer (≤5 min), a redacted device/sync snapshot (app version, platform, locale, uptime, sync/auth state, queue depth — no content), and the triggering error's redacted stack (header line dropped, frames only). Preview and send call the same builder with the sameincidentId, so the consent dialog's preview is byte-identical to what ships.sendReportposts toPOST /diagnostics/report; the server writes one summary line plus one line per log entry, all taggedincident_id, underkind=report.Redaction guarantees: same non-negotiables as What Never Ships — no note content, titles, attachment filenames, absolute home/vault paths, emails, JWTs/tokens, vault/device keys, or IPs.
kind=log/kind=reportmessage text runs through the sameredactTextas the/telemetry/batcherror message above;redactLogLineadditionally redacts each structuredfieldsentry: secrets are dropped first, paths collapse to~//<vault>/, note/attachment basenames are salted-hashed to[name:hash8].ext, known id fields (noteId,deviceId,installId, …) are salted-hashed, and emails become[email:hash8]. IPs are masked to<ip>. UUID-shaped ids in free text are salted-hashed on the client (correlatable, like the id fields above); the server's re-redaction pass has no salt, so it masks them to a fixed<id>instead. A fixed field allowlist (level, scope, action, errorCode, appVersion, buildChannel, platform, arch, origin, workerName, reason, phase, mode, status, kind, result, plus numeric metric keys likedurationMs/itemCount) ships verbatim; most other field values run through the same redaction as the message.Updater backends:
initializeUpdater()picks one backend at init and keeps it for the session. Velopack whenprocess.platform === 'win32', the app is packaged, andnew UpdateManager('https://github.com/memrynote/memry')constructs; electron-updater in every other case, which covers macOS, Linux, and Windows installs made by the older NSIS installer, where Velopack's constructor throwsThis application is not properly installed. Velopack reads the GitHub releases feed itself (releases.win.jsonplus the.nupkgassets) and has noapp-update.yml. Both backends drive the same state machine inapps/desktop/src/main/updater.ts, so thephasevalues, the severity classification, the install marker and the install-health streak below are backend-neutral. What differs is the library's own log lines, which carry scopeElectronUpdateron one path and scopeVelopackon the other.Updater failures: every failure in
apps/desktop/src/main/updater.tslogs adescribeUpdaterError()field bag alongside the raw error, so a silent auto-update failure is diagnosable from Loki alone:phase(startup-check,scheduled-check,auto-check-enable,auto-download-enable, or — for electron-updater's ownerrorevent, which carries no phase — one ofcheck/download/downloaded/installinferred from the status at the time), pluserrorName,errorMessage,errorCode,httpStatus,url,errorCauseand the top stack frames aserrorStack. The field names are chosen against the redaction allowlist above:phaseanderrorCodeship verbatim,urlis path-redacted (query string stripped), the rest are text-redacted and capped.Updater severity classification: the local
main.logline and the user-facing error state are unchanged — every updater failure is still logged aterrorand still flips the UI to the error state. What is classified is the telemetry severity (apps/desktop/src/main/updater-error-severity.ts). A failure during a check (check/startup-check/scheduled-check/auto-check-enable) whose message or cause chain carries only allowlisted Chromium transport codes —net::ERR_NAME_NOT_RESOLVED,ERR_INTERNET_DISCONNECTED,ERR_NETWORK_CHANGED,ERR_TIMED_OUT,ERR_CONNECTION_TIMED_OUT,ERR_CONNECTION_RESET,ERR_CONNECTION_CLOSED,ERR_CONNECTION_REFUSED,ERR_CONNECTION_ABORTED,ERR_ADDRESS_UNREACHABLE,ERR_ADDRESS_INVALID,ERR_NETWORK_ACCESS_DENIED,ERR_NETWORK_IO_SUSPENDED,ERR_HTTP2_PROTOCOL_ERROR,ERR_HTTP2_SERVER_REFUSED_STREAM— ships as anapp_log_recordedwarninstead of anapp_error_seenexception. Being offline is a normal state for an offline-first app, and those events were 33.2 % of every exception in the product. The cause chain matters because electron-updater'sGitHubProviderwraps a transport failure in a parse-shapedERR_UPDATER_INVALID_RELEASE_FEED; a feed that is genuinely malformed has no network cause and stays an exception. The set is an allowlist, never anet::ERR_prefix test:net::ERR_CERT_*/net::ERR_SSL_*are security signals, and anything unrecognised fails closed toerror.ERR_CONNECTION_ABORTED,ERR_ADDRESS_UNREACHABLE,ERR_ADDRESS_INVALIDandERR_NETWORK_ACCESS_DENIEDwere added by #1994, measured on the only population that can evidence a gap in this set — builds already carrying this classification (2026.822.1and newer), since an older build reported every code as an error regardless. An upstream 5xx stays an exception deliberately: a 504 on the releases feed is GitHub failing for everyone at once, the one check-phase shape meaning the whole fleet has stopped receiving updates, unlike the per-device transport codes above that scale with the number of flaky networks. Everything else is untouched — HTTP 4xx/5xx (including theHTTP_ERROR_618jwt:expiredon GitHub's pre-signed asset URLs), signature failures, install-phase errnos,ENOENT … app-update.yml, and any failure in thedownload/downloaded/installphases, where a network drop can leave a half-applied update. Reclassified events are never dropped: same error code, same redacted message and stack, and each one carriesretryCount— the consecutive-failed-check streak — so a cross-install signal can separate one laptop on a train from many installs failing in a row. An install that has not completed a single check in 24 hours and has failed at least 6 checks in that time raises one exception (latched until the next successful check), so a genuinely stuck updater is still loud.Post-exit teardown aborts: a
child-process-gonereport for a worker the owning module released asgraceful_stopis demoted towarn. That release is only ever recorded on an observedcode === 0, so the pairing is proof the worker already exited cleanly and the report is the native runtime aborting while unwinding afterwards — not a failure the user felt. This was 70% of macOS installs (#1990): the embeddings worker handled its shutdown message with a bareprocess.exit(0), which skips JS cleanup and runs onnxruntime's static destructors with sessions still live, aborting with SIGABRT. Node reported 0 and Electron reported 6 for the same process, which is why both codes appear. The worker now disposes the pipeline and drops its lastmessagelistener — Electron'sParentPortpauses itself onremoveListener, releasing the handle that holds the loop open, and there is noclose()to call — with an unref'd fallback that must stay belowSHUTDOWN_TIMEOUT_MSinembeddings.tsso a wedged disposal is never force-killed into the very teardown death this removes. The demotion keys on the recorded release alone, never the phase:idle_shutdownis also reachable from a force-kill, where no exit code was ever observed. Demoted reports keep their error code, message, stderr tail and metrics and stay queryable asapp_log_recorded; they only leave Error Tracking.Repeated install failures: a failed check is loud in telemetry, but a failed install was silent to the user. On macOS, Squirrel.Mac stages a downloaded update into its own ShipIt copy, and when that copy fails (
ditto: Could not lstat …,No space left on device) the error lands in a session that then quits normally. Nothing survived the quit, so the next launch re-served the same cached zip and failed the same way — two production installs sat on2026.817.1through four releases with no signal of any kind (#1999).apps/desktop/src/main/updater-install-health.tspersists the streak inupdate-install-health.jsonunderuserData, keyed on the pair (running version, target version), and after three consecutive failed attempts setsinstallFailedso the existing manual-download dialog appears. It counts attempts, not launches: electron-updater re-serves an already-validated cached zip on every auto-check, so a genuinely stuck install surfaces within ~30 minutes. The streak clears the moment the app boots as a different build, which is the only honest evidence an install applied — display versions cannot be ordered, so "newer than the failing one" is not a test that can be written. Every field is re-validated on read and a corrupt file degrades to "no streak"; an install that has never failed has no file and behaves exactly as before. Both backends share this streak and the pending-install marker (update-install-attempt.json, reported asUPDATE_INSTALL_DID_NOT_APPLY). On Velopack the marker is written before the hand-off toUpdate.exeand read on the next launch, and aninstall-phase failure advances the streak exactly as above. Velopack does no download-time staging, so it never produces thedownloaded-phase attempts Squirrel.Mac does. The marker also records which installer was handed off to (installer:electron-updater,velopack, orvelopack-handoff), and the silent install-on-quit path writes it too, so a Windows update that applies on a normal quit and never comes back is no longer invisible.NSIS to Velopack hand-off: a Windows install made by the older NSIS installer keeps electron-updater for checking and downloading. Once the NSIS
setup.exeis on disk,apps/desktop/src/main/installer-handoff.tslooks forMemryNote-win-Setup.exeon the same release tag (HEAD, then a streamed download intouserData/installer-handoff/, folded into thedownloadingstate) and verifies it withGet-AuthenticodeSignature: statusValidand a signer subject containingCN=Open Source Developer Kaan Karaca, or the file is deleted and never run. The install step then spawns a detached batch script (rendered byinstaller-handoff-script.ts, so its exact command lines are pinned by tests) that waits for the app's PID to exit, runs a copy ofUninstall MemryNote.exe /S _?=INSTALLDIR(no--updated, so the tolerant removal path runs, and the copy blocks until the old install, its shortcuts and its uninstall key are gone), runsMemryNote-win-Setup.exe --silent --verbose --log userData/logs/velopack-setup.log, and falls back to the already-downloaded NSIS installer if no VelopackMemrynote.exeexists afterwards, so the user is never left without an app. Every refusal falls back to the plain NSIS install: a release without the Velopack asset (info log only), a download failure (INSTALLER_HANDOFF_DOWNLOAD_FAILEDthrough the updater error path) and a signature failure (INSTALLER_HANDOFF_UNVERIFIED). The next launch reads the marker: still the same version is the usualUPDATE_INSTALL_DID_NOT_APPLYwithsource: velopack-handoff; a new version running from the Velopack layout (..\Update.exenext to the app) isapp_update_installedwithaction: migratedandsource: velopack-handoff; a new version still running from an NSIS layout isINSTALLER_HANDOFF_DID_NOT_APPLY, meaning the update arrived but the migration did not.Memrynote.exe --cli migrate-installer path\to\MemryNote-win-Setup.exedrives the same verify-then-hand-off path against a local file (exit 0 when scheduled, 1 otherwise, Windows only), which is what themigratejob in.github/workflows/velopack-smoke.ymlruns on a real NSIS install.Expired GitHub signed asset URLs: GitHub serves a release asset by redirecting to a short-lived signed
release-assets.githubusercontent.comURL. When the follow-up GET lands after that token expires, GitHub answers with the non-standard status 618jwt:expired, and electron-updater has no retry on that path —builder-util-runtime'sretryOnServerErroris never called there, and itsisServerError()covers500-599, so it would not match a 618 anyway. One expired token therefore lost the whole update check (36 production exceptions across four releases, all in thecheckphase). The token is minted fresh on each redirect, socheckForUpdates()now retries twice, two seconds apart, on a 618 — or a 403 whose URL is the signed-asset host, host-gated so an unrelated 403 is never retried into a loop (isExpiredSignedAssetErrorinapps/desktop/src/main/updater-error-severity.ts). Only the attempt that still fails reaches theerrorhandler: a recovered check does not flip the update surface to an error, does not advance the stuck-updater streak, and ships oneapp_log_recordedwarn(update check hit an expired release-asset url, retrying) instead of an exception. A 618 that survives every retry is reported exactly as before — the severity classification above is unchanged.Process lifecycle: the main process reports a
child-process-gonefault with a compositetype:reason:nameerror code (e.g.Utility:crashed:Embeddings). The worker label comes from Electron'sdetails.name, notdetails.serviceName: Electron routes a fork'sserviceNameoption todetails.name, whiledetails.serviceNameholds the Mojo interface name — a constant (node.mojom.NodeService) that is identical for every utility fork and so cannot tell our workers apart. A fork that passes noserviceNameoption reports the defaultNode Utility Process. The exit status rides along in the line'sexit_codefield rather than inside the error code, so crashes still group by worker in PostHog Logs while the POSIX signal (11 SIGSEGV, 6 SIGABRT) stays visible. A utility worker's clean idle-shutdown (embeddings, image-processing, voice-model each exit after ~30s idle) is a lifecycle event, not a fault, so aclean-exitreason is skipped entirely — only a real fault produces an error event, mirroring the GPU crash guard. Note thatchild_process_goneis not throttled: the crash cadence is itself a diagnostic signal. When the owning module can resolve a lifecycle phase for the dead worker, the breadcrumb becomeschild_process_gone_<phase>and the phase joins the exit status in the message (Embeddings utility process crashed (exit 6, idle_shutdown)). The error code stays phase-free: it is the Error Tracking fingerprint, so splitting it per phase would orphan the existing issue's history.Embedding worker: phase is
starting,in_flight,idle_shutdown, oridle. This is what separates a harmless teardown crash (idle_shutdown— the embedding was already delivered) from real user impact (in_flight— the user silently lost semantic-search indexing for that note).The phase is resolved by
getEmbeddingWorkerCrashContext(details.name, details.reason)at thechild-process-gonecall site, not inside the worker's ownexithandler. Electron'sUtilityProcessexitevent does not fire for a native crash — a SIGABRT out of the model runtime is neither a graceful exit nor the V8FatalErrorthe instanceerrorevent covers — so the bridge never learns its worker died. Production proved it: across 107 consecutiveUtility:crashed:Embeddingsevents the bridge's ownworker_exit_<phase>breadcrumb emitted zero, and so didembed_failed, whilechild-process-gonefired for all 107. Resolving the phase at the report that does arrive is what makes those events answerable. Theworker_exit_<phase>breadcrumb is kept for the paths whereexitdoes fire (a non-crash abnormal exit, a force-kill), and carries a different error code (EmbeddingWorkerExit) so the two are never confused.The report never arrives while the bridge still owns the worker. Resolving the phase from the live handle alone therefore answered
nullin 100% of production crashes (issue #1582): 76 events on2026.817.1, none of them phase-suffixed. Every path that nulls the process handle either latches a teardown phase (stop()/reset()) or runs inside anexithandler that emitsEmbeddingWorkerExit— and production has neither on that release. What remains isfailProcess(), which forgets a worker that is still running (10s start timeout, fatal error). So the bridge keeps a last-worker record — pid, phase, how it was released, fork timestamp, model-cache state, stderr tail — bounded by a 60s TTL so an unrelated later crash cannot inherit it.failProcess()now also kills the worker it gives up on: the orphan was unreachable (every request goes through the process handle) yet kept a whole onnxruntime alive to abort later.Anything reading the phase from outside the exit handler must also survive the force-kill race:
reset()latchesidle_shutdownbefore killing, then nulls the process handle, so both the latch and the last-worker record are cleared when there is no process to kill — the latch now outlives the exit handler that used to clear it, and a stale one would make the next worker's crash read as a teardown it never had.The crash report also carries what the main process knew about that worker. Telemetry events accept at most one dimension (
TelemetryDimensionsSchema) andlog_actionholds the phase, so the rest ships inside fields that already exist — no new event name, no new dimension key, no contract change, and therefore no sync-server deploy:Fact Where it ships Query it as lifecycle phase dimensions.log_actionlog_action(logs + events)platform exit status metrics.valueexit_code(logs) /valueworker uptime at death metrics.durationMsduration_ms(events)crashes this session metrics.retryCountretry_count(events)cached model size metrics.byteCountbyte_count(events)pid, reason, release, cache state, first-load vs reload error.messagesuffixmessage(logs)worker stderr tail error.stackstack(logs)The message suffix is a bounded
[reason=… pid=… uptime=…ms release=… cache=… cache_bytes=… load=… crashes=…]block appended only when a context was resolved, so every otherchild-process-gonefamily's message is byte-identical to before. The stderr tail is the closest thing to a stack trace this family can produce — the process that died is not the one reporting, sostackis otherwise always empty. It is captured into a bounded per-worker ring buffer, redacted on the device with the same saltedredactTextsetup as Path A log shipping, capped at 2000 bytes, and every line is prefixed so it can never be mis-parsed back into a fabricated PostHog Error Tracking frame.Embedding generation failures emit
embed_failed, throttled to one event per 5-minute window because a broken worker would otherwise fail once per note edit.Desktop IPC envelopes: every
{ success: false }error envelope produced by the IPC layer (withErrorHandler/withDb) also emits anapp_error_seenevent, throttled in-memory to one event per action + error code per minute so an error loop can't flood the telemetry queue. The key must discriminate: keyed by error name alone and shared across all handlers, one benign recurringErrormasked a genuine differentErrorfrom another handler for the whole window. The expectednoVaultOpenenvelope is not tracked, and the envelope's user-facingerrorstring (which may contain note-derived text) never leaves the process — only the error code and redacted stack frames ship. Handlers that throw instead of returning an envelope report the same way:createHandlerandcreateValidatedHandler(main/ipc/validate.ts) wrap the call, so canvas and calendar reads — which never pass throughwithDb/withErrorHandlerand whose rethrow used to be their only trace — are now countable. A Zod validation failure is reported too (renderer↔main contract drift, not user error); only theZodErrorname and stack ship, never the issue messages, which can echo input values. An expected condition is skipped before the throttle map, so a suppressed error can't claim the key and mask a real failure from the same handler. This throttle is main-side only:trackRendererErrorsends oneapp_error_seenper call, so a renderer loop calling one failing handler produces many renderer-sourced events against at most one main-sourced event per minute. Renderer and main counts for the same underlying failure are therefore not comparable — a large gap is the throttle, not a dropped main-side event.Vault watcher: chokidar
onErrorcan burst per file (a permission-denied subtree), so watcher faults are sampled to oneapp_error_seenper minute (main/vault/watcher.ts). Every error still reaches the local log.Unclaimed persistence tokens: a
((mention:…))/((date:…))token the note-open normalize chain left as literal text, or a callout marker orphaned of its>prefix, is a block that will render broken with no error anywhere — the failure mode behind the #1843 round-trip epic, previously detectable only by a user emailing a screenshot. The renderer counts them after every normalize pass (renderer/…/content-area/unclaimed-token-telemetry.ts) and reports throughapp_error_seenasaction: editor_unclaimed_tokenwitherrorCode: unclaimed_mention/unclaimed_date/unclaimed_callout_markerand the occurrence count inmetrics.itemCount. First sighting emits immediately; after that, counts aggregate into at most one event per kind per minute, since the chain re-runs on every note open and remote update. Metric only — the user sees no toast or dialog, and no token content ever ships.User-visible failures that previously left no trail: a renderer
did-fail-load(the user is staring at a blank window) reports asDidFailLoad:<chromiumErrorCode>— the URL never leaves the process; amemry-file:protocol serve failure (a silently broken image/PDF/video embed) logs aMemryFile/serve_failedbreadcrumb with the coded errno before answering 404; a window-close flush rejection (window refuses to close, edits possibly unsaved) reports aswindow_close_flush_failed; and quick capture reportsQuickCapture/global_shortcut_register_failedwhen both the configured and the fallback global shortcut fail to register.Expected conditions: some failures are normal states, not faults. They still surface to the UI as an error envelope, but the throw site marks them and error telemetry skips them, so they cannot drown real signal. Currently marked: an Ollama model-list fetch that is refused (
ECONNREFUSED= Ollama is not running), and a calendar OAuth timeout (the user opened the consent screen and walked away). The suppression is deliberately narrow — a real Ollama misconfiguration (DNS failure, connection reset, or a bad HTTP status) is still reported.Server errors:
captureServerErrorpushes its redacted detail (operational message, stack, normalized path, error/status codes) asapp="server"lines — levelerrorfor 5xx/unhandled,warnfor handled 4xx.Beyond route handlers: failures that never produce a failing HTTP response also reach PostHog Logs — scheduled cleanup-task failures (
source="cron"), token-revoke failures during logout, Resend email send failures (RESEND_SEND_FAILED), andUserSyncStateDurable Object alarm/websocket handler errors (source="user_sync_state_do", pushed directly viapushPostHogLogssince no route error handler ever sees them).Attributes:
app(desktop/server) andenvare resource-level attributes shared by every line in a push;kind(error | log | report) and, when known,posthogDistinctId(the already-hashed identity, so PostHog can attribute the line to a person) are per-line attributes;levelrides inseverityText. Everything else — the JSON body underlineinservices/posthog-logs.ts— is the log body, not an attribute.Retention: 14 days, per PostHog Logs' own retention policy.
Viewing: log lines are searchable in PostHog's Logs product, filterable on the
service.name/deployment.environment/kindattributes and bydistinct_id.
Canvas Rollout Panels (Grafana / Analytics Engine)
Panels for the spatial canvas rollout. Create these by hand in Grafana Cloud against the Analytics Engine dataset; they are recorded here so the rollout is reproducible and reviewable.
canvas_opened fires on every successful canvas load, so a tab-switch remount counts again. Read it as "canvas loads", not "distinct canvases opened".
The queries below are written against the real datapoint layout used by toDataPoint() / writeServerPoint() (apps/sync-server/src/services/telemetry.ts, apps/sync-server/src/services/analytics.ts): blob1 holds the event name and index1 holds the HMAC-hashed install id, against the dataset named in apps/sync-server/wrangler.toml (memry_product_telemetry_production / _staging / _dev). The column names are real; the exact SQL dialect accepted by the Analytics Engine SQL API through Grafana's Infinity datasource has not been verified by running these — treat them as a documented starting point, not verified panel queries, and adjust syntax as needed when building the panels.
| Panel | What it answers | Query |
|---|---|---|
| Canvas adoption | Are people turning it on and using it? | SELECT toStartOfInterval(timestamp, INTERVAL '1' DAY) AS day, count() AS loads, uniq(index1) AS installs FROM memry_product_telemetry_production WHERE blob1 = 'canvas_opened' GROUP BY day ORDER BY day |
| Canvases created | Is creation growing or one-and-done? | SELECT toStartOfInterval(timestamp, INTERVAL '1' DAY) AS day, count() AS created FROM memry_product_telemetry_production WHERE blob1 = 'canvas_created' GROUP BY day ORDER BY day |
| Conflict-copy rate | Is last-write-wins hurting real users? | SELECT toStartOfInterval(timestamp, INTERVAL '1' DAY) AS day, count() AS conflicts, uniq(index1) AS installs FROM memry_product_telemetry_production WHERE blob1 = 'canvas_sync_conflict_copy' GROUP BY day ORDER BY day |
| Oversized canvases | Is the size cap being hit? | SELECT toStartOfInterval(timestamp, INTERVAL '1' DAY) AS day, count() AS blocked FROM memry_product_telemetry_production WHERE blob1 = 'canvas_too_large' GROUP BY day ORDER BY day |
| Unknown sync types | Mixed-version tripwire | SELECT toStartOfInterval(timestamp, INTERVAL '1' DAY) AS day, count() AS skipped FROM memry_product_telemetry_production WHERE blob1 = 'sync_skipped_unknown_type' GROUP BY day ORDER BY day |
Swap memry_product_telemetry_production for _staging when checking a staging deploy. The event names are the stable part of these queries.
Go/no-go for flipping the canvas feature flag on by default is recorded in docs/superpowers/specs/2026-07-22-spatial-canvas-m7-rollout-design.md §7.
Server Configuration
Set these sync-server variables to enable PostHog capture and log shipping (unset in local dev, where both are a no-op):
POSTHOG_KEY=... # wrangler secret (PostHog project token)
POSTHOG_HOST=https://us.i.posthog.com # wrangler var (staging and production)GITHUB_TOKEN is only used by the daily release download-count cron, and in practice it is required in staging and production. Without it the pull is unauthenticated and shares the 60-requests-per-hour-per-IP budget with every other Worker on the same Cloudflare egress address, which other tenants routinely exhaust before our one daily request arrives — GitHub then answers 403 and no release_download_count_snapshot event is emitted that day. A fine-grained PAT with public-repo read access is enough (this reads a public repo's Releases API and needs no write scope), and authenticated calls get 5,000/hour:
wrangler secret put GITHUB_TOKEN --env staging
wrangler secret put GITHUB_TOKEN --env productionNo workflow uploads it; it is set by hand, like every other Worker secret. It is deliberately not in the requiredSecrets fail-fast list — a missing token must not take the whole Worker down for a once-a-day measurement.
A failed pull throws and is reported as a release_download_counts cron failure rather than silently skewing the numbers, and the two failure shapes are separable. GitHub's own 403/429 raises GitHubReleasesRefusedError, which carries code: 'GITHUB_RELEASES_REFUSED' and the upstream status, so it reports as a handled 4xx logged at warn — expected upstream backpressure, with the message recording whether a token was in play. Anything else stays an unhandled 500. Nothing is corrupted either way: the throw happens before the D1 read/write, so the stored baseline is untouched and the next successful run emits the accumulated delta. The distortion is in the daily series (a zero, then a spike), not the running total.
Diagnostic Log Endpoints
Two additional endpoints feed the kind=log / kind=report streams. Both accept unauthenticated requests (no sign-in required), are rate-limited per user/IP, Zod-validated, and PostHog-Logs-only — neither writes a PostHog product event:
| Endpoint | Stream | Rate limit | Payload |
|---|---|---|---|
POST /telemetry/logs | Path A (kind=log) | 120 req / 60s | DiagnosticLogBatchSchema — 1–50 redacted log lines |
POST /diagnostics/report | Path B (kind=report) | 10 req / hour | DiagnosticReportSchema — ≤200 redacted lines + a redacted device/sync snapshot + the triggering error |
A malformed payload is rejected with 400 VALIDATION_ERROR (only the Zod path + issue code is logged, never values, same convention as /telemetry/batch). A valid payload always gets 202, including when POSTHOG_KEY is unset — the push inside pushPostHogLogs is a silent no-op in that case, so a dev build never error-spams.
/diagnostics/report attributes to an account when the desktop attaches a bearer (see Account identity); /telemetry/logs is still anonymous because the log shipper does not attach one yet. DiagnosticReportSchema.accountId is accepted for backward compatibility with older desktop builds but is deliberately ignored — a body field is client-asserted, and it would feed a distinct_id whose $identify merge is permanent, so identity comes only from the verified bearer. /telemetry/batch is unchanged by either endpoint.
Performance
trackTelemetry is debounced and batched. Calls during the first second of startup are deferred until after the vault is open so they never delay first paint. On the sync server, PostHog event capture and PostHog Logs pushes both run in waitUntil so neither can block the /telemetry/batch response.
Launch: note-readable mark
renderer/src/lib/launch-restore.ts stamps a performance.mark('memry:note-readable') when the note a launch restored has actually rendered its block tree, hooked to BlockNote's onEditorReady rather than to component mount, because mount only proves the CRDT binding started. scripts/launch-bench.mjs reads it over CDP and reports it as renderer_note_readable_ms.
The mark is deliberately narrow. It fires only for the note the launch restored, only once, and never for a note opened later in the session, so the metric cannot be inflated by ordinary navigation.
It is absent, rather than zero, whenever the launch restored something that is not a note. The restored tab is read from localStorage because the vault path is not available synchronously in the renderer, and the key with the newest savedAt wins. A machine with several vaults can therefore name a note from a vault the app is not opening: the speculative chunk prefetch is then wasted, and the mark never fires. Absent is the correct reading in that case, not a failure, but it does mean a median over launches must state how many runs carried the mark.
Login-shell PATH probe: failure classes
The probe runs $SHELL -ilc once per packaged launch to recover the user's real PATH, and reports one of four outcomes through login_shell_path_probe_failed:
marker_missing— the shell ran and exited cleanly but printed no marker, usually an rc chain that writes to stdout.status_<n>— the shell ran and exited non-zero.status_none— the probe hit its 3 s deadline and was killed.spawn_error— the shell never started.
Two details are load-bearing and easy to undo by accident. The child's stdin is closed, not piped: an rc chain that reads stdin blocks forever on an open pipe, and every consumer awaiting whenLoginShellPathApplied() blocks with it. The deadline kills with SIGKILL, because an interactive shell ignores SIGTERM, so the polite signal would leave the promise pending for the whole session.
The unit suite injects a fake probe and therefore cannot catch either regression. A change to how the child process is spawned needs a real shell to verify.