Sync Protocol
Encrypted payloads move between devices through a Cloudflare Workers API backed by D1 and R2.
Storage Split
| Storage | Holds |
|---|---|
| D1 | Sync item metadata: id, type, vector clock, blob key, size, content hash, timestamps |
| R2 | Encrypted payload blobs (avoids the 1 MB D1 row limit) |
Splitting metadata from blob saves cost and lets the server reason about ordering without ever touching ciphertext.
Entitlement Gate
Every /sync/* route is authenticated and paid-gated before record, CRDT, WebSocket, or blob logic runs. Paddle webhooks write the active sync_entitlements row for the user, and the server copies the plan limits into quota enforcement:
| Plan | Storage limit | Vault limit | File limit | Version history |
|---|---|---|---|---|
| Plus | 1 GB | 1 | 5 MB | 30 days |
| Pro | 10 GB | 10 | 200 MB | 365 days |
| Believer | 50 GB | Unlimited | 200 MB | 365 days |
Inactive, past-due, paused, canceled, or expired entitlements return SYNC_PAYMENT_REQUIRED before sync data is read or written. Vault and file-size limits return SYNC_VAULT_LIMIT_EXCEEDED and STORAGE_FILE_TOO_LARGE.
SYNC_PAYMENT_REQUIRED and SYNC_VAULT_LIMIT_EXCEEDED are both HTTP 402 and mean opposite things: the first says there is no active plan, the second says the plan is active and this vault is one more than it syncs. Clients must branch on the error code, never on the status alone — a vault-limit 402 reported as a billing failure tells a paying user to pay again. The desktop classifies them as the separate sync_payment_required and sync_vault_limit_exceeded error categories, and only the former is an entitlement statement.
A request with no X-Memry-Vault-Id header resolves to an existing vault and never creates a second one: an account with no vaults (or one that already holds the legacy default vault) is served default as before, an account with exactly one vault is served that vault, and an account with several is answered 400 VALIDATION_ERROR rather than guessing and consuming a vault slot.
The desktop client mirrors this gate locally to avoid pointless round-trips that can only return 402. Handlers for paid-only endpoints check the cached entitlement first and return their empty value (GET_STATUS → local_only, GET_STORAGE_BREAKDOWN → null) when the cache says the user is on the free plan. Only a known-unpaid entitlement is gated — an unknown/uncached entitlement (fresh install, before the first status call) still calls the server, so the gate can never lock a paying user out on stale local state. The server-side gate remains authoritative.
The sync runtime's own gate additionally expires that negative: a cached "not paid" older than one hour (or written without a timestamp) is re-fetched from GET /auth/billing at startup rather than trusted. Nothing refreshes the cache in the background, so without the expiry an upgrade bought on the web or on another device would leave the user in local-only until they opened Settings → Account. A cached paid verdict needs no expiry — the server re-verifies it on every request.
Development sync servers can seed a dev_seed Believer entitlement for configured local admin accounts during sign-in, billing checks, reconcile, and paid-sync middleware access. This path is guarded by ENVIRONMENT=development; production and staging rely on Paddle webhooks, explicit admin overrides, or billing reconcile only.
Desktop checkout is account-owned. The app requests /auth/checkout-token, opens memrynote.com/pricing with the token in the URL fragment, and the landing page passes that token to the Paddle checkout transaction API. After payment, Paddle webhooks are the primary entitlement writer. Desktop can also call /auth/billing/reconcile with the returned transaction id; the server fetches the Paddle transaction, verifies the embedded memrynote user id, and provisions the entitlement only for completed transactions.
Billing status and customer management stay on authenticated account routes:
| Path | Purpose |
|---|---|
GET /auth/billing | Return current plan, status, limits, usage, expiry, portal flag |
POST /auth/billing/reconcile | Reconcile an optional Paddle transaction id into entitlement |
POST /auth/billing/portal-session | Create a temporary Paddle customer portal URL |
Portal URLs are temporary authenticated links from Paddle and are never cached. Refund and chargeback automation is intentionally out of scope; support handles those from email and the Paddle dashboard.
Client Identification and the Per-Platform Write Gate
A client may identify itself with a x-memry-client: <platform>/<semver>[+<build>] header (ios, android, or desktop). The header is optional: a request without it is a legacy desktop client and keeps full access, unchanged. A malformed header is treated as absent and logged rather than rejected — a parser bug must never be able to lock a user out of their own vault. Pre-release versions (1.0.0-beta.1) count as malformed, so a beta can never satisfy a floor its release does not.
The server keeps one client_policies row per platform, holding a semver write floor and a kill switch. It is consulted on writes only; reads are never gated, so a device dropped to read-only can still open every note it owns.
| Condition | Server behaviour |
|---|---|
| No header | Allow (legacy desktop) |
No row, or min_write_version is NULL | Allow |
| Version at or above the floor, writes on | Allow |
| Version below the floor | 426 with CLIENT_UPGRADE_REQUIRED and the required minVersion |
writes_enabled = 0 | 403 with PLATFORM_WRITES_DISABLED |
The kill switch is evaluated before the floor: when writes are off for a platform, telling users to upgrade would send them chasing a release that cannot help. Every uninterpretable policy — absent row, NULL floor, unparseable floor — resolves to allow, so an unreadable policy table degrades to today's behaviour rather than to an outage.
GET /sync/status echoes the caller's own policy as an optional clientPolicy field whenever the request identified itself, so a device learns about a flipped switch on its next foreground poll instead of by attempting a write and being rejected. Header-less clients get byte-identical status responses.
On receiving either rejection a client enters explicit read-only mode, parks its outbox (queued writes preserved, attempts stopped), polls the policy on foreground, and resumes automatically once clear.
Write attribution
Item writes are stamped with the calling platform and version on sync_items, crdt_updates, and crdt_snapshots. NULL means the row was written by a client that predates the header — which is every desktop build shipped so far; there is no backfill. The CRDT tables are included because a note's body lives there, and that is the payload most likely to need a targeted rollback after a mobile incident. Attribution records the latest writer, not the creator, so a desktop rewrite clears an earlier mobile stamp. No read path depends on these columns.
Sync Items
Every domain object syncs as a sync_item. The server sees:
{
id: string
user_id: string
device_id: string // last writer
type: 'note' | 'task' | 'agent_conversation' | 'agent_message' | ...
vector_clock: VectorClock // doc-level
blob_key: string // R2 path
size_bytes: number
content_hash: string
created_at: timestamp
updated_at: timestamp
deleted_at: timestamp | null
signature: bytes // Ed25519 over the metadata + blob hash
crypto_version: int
}The blob is the encrypted body. The server can reason about order, dedupe, and authorize writes — but the contents stay opaque.
Blob key layout
Item ids are human-readable and may repeat across types: the default project id is inbox, a tag_definition id is the lowercased tag name, and a folder_config id is the folder path. R2 keys for sync-item payloads therefore include the item type, and since items-v3 also the payload's content hash — new pushes write to <user>/vaults/<vault>/items-v3/<type>/<id>/<content-hash>, so a project and a tag both named inbox own separate objects, and every push writes its own immutable object instead of mutating a shared per-item one. Content-addressing is what makes concurrent pushes of the same item safe: with a shared mutable key, two devices racing on one id (external calendar events have deterministic ids, so every device pushes the same ids) could interleave blob and row writes such that the surviving row carried one push's signature over the other push's bytes — the item then failed Ed25519 verification on every pull until re-pushed. After a replacing push commits its row, the previous version's object is deleted best-effort; a delete that loses a race merely leaks a bounded orphan object. Rows written before these layouts keep their legacy items/<id> or items-v2/<type>/<id> keys; every read path resolves the blob_key stored on the row rather than re-deriving it, so old rows continue to work without a migration. (The untyped layout let same-id items of different types overwrite one shared object, which permanently broke the losing row's signature.) A pull that finds a row whose object is missing skips that row instead of failing the page: a replaced item re-arrives at a later cursor, and a dangling row must not wedge every puller behind one broken item.
Per-item bookkeeping and retry semantics
Because ids repeat across item types, every piece of client-side per-item bookkeeping — the signature-failure quarantine, the corrupt-item re-fetch tracker, the within-run apply dedup, and the manifest diff — keys on the (type, id) pair, never the bare id. A permanent quarantine on one type does not block its same-id sibling of another type, and a re-fetch that asks for one (type, id) pair ignores the sibling rows the server returns for the same id.
The within-run apply dedup also compares cursors. For each (type, id) a pull run applied, it keeps the highest cursor at which a changes page listed it: the ref's serverCursor, or the page's nextCursor for a deleted id or a server that sends no ref cursor. A later page that lists the item above that cursor carries a newer version (an edit or delete committed during the run), and it goes through the normal clock-guarded apply. Only a listing at or below the recorded cursor is skipped.
Retry semantics: the pull cursor only advances past pages that were actually applied. A page the client refused (all items failed crypto, or the key was mid-transition during sign-in/recovery) does not move the cursor, so a manual Retry lands on the same page instead of skipping it and reporting a clean sync. Persisted quarantine entries expire after 7 days — if the underlying server row is still broken the item re-quarantines within a few pulls, and if it was repaired server-side the item flows again without an emergency wipe. The manifest-check throttle (30 minutes) persists in sync state, so engine restarts and vault switches cannot re-arm an immediate check.
Both of those in-memory ledgers are bounded, because a server-side incident can brand a very large number of items in a single session. The corrupt-item tracker holds at most 5,000 cooldown entries and sheds the coldest ones first; the coldest entry is also the one closest to its one-hour cooldown lapsing, so the most an eviction can cost is one extra re-fetch for an item that was about to become eligible anyway. Expired entries are also swept at the end of every pull, not only when a later pull happens to touch the same item.
The quarantine ledger applies the same 7-day expiry to live entries that it already applied to persisted ones, so a long-running session behaves like a restart. Its 10,000-entry cap is deliberately soft: only entries that have not yet reached the permanent threshold can be evicted, because those are attempt counters and a still-broken item simply re-quarantines. Permanent quarantines are the record that keeps a failed-signature item out of the vault and are never dropped to satisfy the cap.
A payload this build's schema refuses, usually a newer peer's shape, is not dropped. The item returns 'schema_invalid' and goes into a persisted ledger (schemaInvalidItems sync-state key) with the app version that refused it, and the cursor moves on, because one global cursor cannot wait for one item. Once another app version runs, the next pull re-fetches every ledger entry by id, in 100-id requests, and applies it again in FK order. An entry the new build still refuses is kept with the new version; one that applies, or that the server no longer has, leaves the ledger. A pulled item that fails the envelope schema goes into the ledger too, while its page-mates apply; because that is usually a server fault, it is also retried after the one-hour corrupt-item cooldown. Ledger entries are listed with the quarantined items and count as known to the manifest check, so they never trigger a full re-pull. Orphan repair never tombstones a child whose parent the server still has, even when this build cannot apply the parent.
A changes page of up to 500 refs is pulled and applied in slices of at most 100 ids, one transaction per slice, and LAST_CURSOR never moves before the last slice. Up to PULL_SLICE_FETCH_WINDOW (2) slice POSTs are in flight at once, so the next slice downloads while the current one applies. Slices still apply in page order, and a stop leaves at most one fetched slice unread. When nothing has to run after the last slice commits, the cursor is written as the last statement of that slice's transaction, so the cursor and the page's last rows commit together or not at all. A page with a by-id re-fetch pending, notes whose CRDT bodies are fetched after the commit, an item deferred for a retry, or a transaction that could not be opened writes the cursor after that work instead. The page's notes are flagged as holding unmerged CRDT state inside the slice transaction, before any cursor write. A crash before the cursor commits pulls the whole page again; the rows that already committed come back with an equal clock and an identical payload and are skipped without a row write or a renderer event (a still-dirty syncedAt is stamped, and missing canvas assets and note attachments are requested again). An equal clock with a different payload still applies, because that is how two devices whose merge re-pushes collided converge (protocol 06 §6.5.2 P4). Renderer events raised while a slice applies are held until its transaction commits and dropped if the slice or the item rolls back, so no window is told about rows that never landed.
A tombstone past the user's version-history window loses its payload on the server but keeps its row as a marker, so a device that was offline longer than that window cannot push the item back: the marker refuses the stale write per item exactly as a fresh tombstone does, for every client. The exception is a re-create of an item whose id comes from what the user sees (a journal's date, a tag or property name, a folder path, a bookmark target, a provider calendar or event) or from a file the user can restore (a note's frontmatter id, restored from a backup or the trash): a create over its marker is accepted, so re-creating or restoring it reaches the other devices. Every device remembers the clock of each delete it makes or applies for such an id (desktop: sync_tombstone_clocks), and a re-create is minted with that clock merged in and ticked, so it happens strictly after the delete: the server accepts it inside retention too, a stale copy from a device that missed the delete is refused or merged rather than applied over it, and the device's own late tombstone does not delete it. At runtime start the desktop replays the deletes it still owes (sync_pending_deletes), except for such an id that has a live local row again: that delete is retired instead, so a re-create made while sync was off, or before its delete's clock was recorded, is pushed live rather than deleted on every device. A clockless note or journal row is queued as that re-create by the retire itself. The start-up seed (seedUnclocked) does not push a clockless note or journal row whose id has a recorded tombstone clock: without a pending delete it cannot tell a re-create from a row the delete left behind, such as the rows an older build's pack write-back resurrected, and pushing one would bring the deleted note back on every device. The row and its file stay on the device, logged once per id. The next user edit of that note mints its clock past the tombstone and re-creates it. A delete raised with no snapshot is clocked from the local row's own clock; with no row, or a row that was never clocked, it is not pushed at all, because a clock minted from nothing is refused by the server as a replay.
Only a client that declares purged_tombstones in X-Memry-Sync-Types sees a marker on the read side; for every other client (older desktops, iOS) it is invisible, as the old hard delete was. The desktop declares it. /sync/changes then lists the marker's id, except on a request from cursor 0, and /sync/pull returns the delete in a purgedTombstones list next to items, without ciphertext. The entry carries no record signature, since the shed deleted the payload it covered. Instead, a delete pushed by a current client carries a delete attestation: an Ed25519 signature by the deleting device over {purpose, id, type, deletedAt, clock} (protocol 04 §4.8.4). The server verifies it on push, rejecting a bad one per item as SYNC_INVALID_SIGNATURE, and keeps it through the shed. It serves the attestation with the entry as signerDeviceId and deleteAttestation. The desktop applies an entry only if the request asked for that id, the type carries a required clock, the entry has one, and the attestation verifies under the signer's key over the entry's own type, id, clock and deletedAt. It then goes through the same delete path as a signed tombstone: a local row whose clock happens strictly after the tombstone is kept. An entry with no attestation, an unknown signer, or a signature that does not verify is logged and never applied. That covers every marker shed before attestations existed. A failure to fetch the signer's key fails the page and the cursor holds. It never applies one in a pull run that started from cursor 0, over a local row with no clock (a restored vault folder), over a re-created or restored item that is still queued or was touched after the delete, or when the clock names a device this account does not know. An applied purged tombstone is an ordinary delete inside the slice transaction and adds no post-commit work; a refused or skipped one is not applied at all, so neither changes when the cursor commits.
A live row whose payload the server lost comes back in a blobMissing list instead of vanishing. The desktop applies nothing for it, keeps any local row, and records it in the schema-invalid ledger (retried after the one-hour cooldown, never counted server-only). Orphan repair tombstones a child only when the server positively says its parent is gone, never because a parent was not returned. An unattested purged parent counts as gone, the same as a parent the server omits. Nothing on the desktop deletes a local row because the server does not list it, and the manifest check re-uploads local rows the server lacks only after a pull that delivered in the same run. Tombstones the server hard-deleted before markers existed stay unprotected. Protocol 05 §5.12.3 and §5.12.4 have the full rules.
A /sync/pull body that is not a pull envelope at all is a server contract regression: the cursor holds, the run is refused and sync shows a server error, so the page re-arrives once the server answers correctly.
The first /sync/changes page of a pull outside a full sync (a socket wake or the periodic pull) asks for inline=1. The server then clamps the page to 100 refs and returns, in inline, the exact /sync/pull item for every id it could inline: rows up to 64 KiB, tombstones included, and only ids whose rows on the page all fit. The client pulls only the page ids that no inline item names, so a wake that delivers one small change makes no /sync/pull call. Inline and pulled items apply as one page, in one transaction and one apply order, and the cursor rules do not change. Later pages and full syncs (startup, Sync now, first sync) keep 500-ref pages without inline. Protocol 05 §5.11.2 has the full rules.
Each pulled page applies inside one SQLite transaction. Handlers run synchronously inside it; the files they produce (a note or journal markdown file, a rewritten .memry/properties.md after a remote property-definition delete) are recorded in a crash journal just before the commit and written after it. A crash between the commit and the file writes is healed at the start of the next pull. A handler that wrote its rows from an unawaited promise could land them after the page committed, or inside the next page's transaction, with no crash-journal record for its file. Synced journal entries were applied that way until #2284.
A remote note or journal delete goes through the same journal (#2385). Its row delete commits with the page and the file is removed after the commit. Before, the file was removed by an unawaited unlink with no journal record, so a crash or a failed unlink left a file with no row. The vault indexer could adopt that file as a new local note and push it back. On replay, a journaled delete removes the file unless the file's modification time is later than the moment the delete was journaled, so a file re-created on this device after the crash is kept. A delete entry has no content field, and replay in older builds ignores entries without one. An older build that finds the journal therefore skips the delete and replays only the writes. Each path keeps only its latest journal entry, so no earlier write to a deleted path is left behind for an older build to restore.
When a push starts
A local mutation asks for a push, and the push goes out at once when the last requested push started more than 300 ms ago (PUSH_DEBOUNCE_MS). A request inside that window arms one timer for the rest of it, and every other request in the window rides on that timer, so continuous editing costs at most one POST /sync/push per 300 ms. That was a flat 2-second trailing timer, which made up most of the delay between an edit and its arrival on another device.
A request that finds a sync cycle running does not arm a timer. It is held, and the push runs once when the cycle ends: after the scheduled cycle's promise settles, or, for a cycle started directly (the first full sync at startup, a manual sync, the final push on shutdown), when the engine releases its sync lock or the full sync returns. The old code re-armed the 2-second timer while the cycle ran, so a push raised during a long pull waited up to 2 seconds more after it. Nothing requested before or during stop() pushes after it.
Note bodies follow the same rule per note with a 1-second window: the first update after a quiet second is flushed at once, updates inside the window go out in one trailing flush, and an update that lands while that note's push is in flight is flushed once when the push settles. The window stays at 1 second because every CRDT push route shares one 300-per-minute crdt_push bucket per device.
Body updates are queued durably in the same sync_queue, as append-only note_body rows the record push never dequeues (#2298). They are not coalesced like record rows: coalescing overwrites the payload, which would keep only the last of a note's unflushed Yjs updates. The note-body outbox merges a note's rows at flush time and deletes exactly the rows a push carried once it succeeds, so updates survive a crash or a quit while offline. Queued note_body rows survive sign-out; every other sync_queue row is cleared. See CRDT & Notes Sync.
Push acknowledgements and in-flight mutations
The push queue coalesces: a new mutation for an item that already has an unattempted row overwrites that row's payload instead of inserting a second one, and dequeue is a plain read that leaves no in-flight marker. A row handed to a push therefore stays a valid coalesce target for the whole flight — worker encryption, the round trip, and every retry — and the user can rename or re-tag the item at any point in that window.
An acknowledgement is consequently conditional: the push remembers the payload each row held when it was dequeued and only deletes rows that still match. A row that changed under the push is left queued and goes out on the next iteration. Deleting unconditionally would drop the newer mutation permanently, because the local clock advances at mutation time: the item would sit ahead of the server with nothing queued, and every later pull would resolve skip rather than repair it.
Enqueue-time coalescing only folds into a row that has not been attempted yet, so a failed or rejected push leaves the next edit to open a second row for the same (type, itemId). Both rows can then land in one batch, and the push collapses them again before encrypting. That batch-level collapse keeps the newest row: the batch arrives oldest-first, and the older row's payload is by definition the stale one. Keeping the newest row also keeps the row that is still unattempted, which is the row a concurrent local edit would coalesce into — so the conditional acknowledgement above continues to guard it.
The superseded row's operation is folded into the retained row with the same precedence enqueue uses: a later delete wins outright, and an unacked create survives a newer update, because the server has never seen the item and an update for an unknown id is not an equivalent request.
Collapsing to the newest row matters on its own terms rather than as an optimisation. The pushed payload is normally rebuilt from local state at push time, but that rebuild hook is optional on the handler interface and settings does not implement it — settings live in config.json and the preferences cache, not in a sync table there is anything to rebuild from. For any such type the frozen queue payload is the entire push, so retaining the older row published older state and discarded the newer edit through the success path, with no error surfaced.
A batch the server refuses shrinks instead of ending the run
The acknowledgement rules above all assume a per-item verdict. A request that never reaches the handler has none: Cloudflare terminates an oversized POST /sync/push at the edge — invocation outcome exceededCpu, an empty 503 with no server error code — so nothing is accepted, nothing is rejected, and no attempt is charged. Resending that request unchanged cannot succeed, because what tripped the limit is the batch itself: a 100-item push costs roughly 50 ms of Worker CPU, about 80% of it the per-item Ed25519 signature check.
The push loop therefore reads a 5xx as a statement about the batch's shape rather than a transient blip. It halves the batch and keeps going — from PUSH_BATCH_SIZE (100) down to MIN_PUSH_BATCH_SIZE (1) — so the queue drains at whatever size the server can take. withRetry is given retryOn5xx: false for any batch that can still be split: retrying an identical request first spends the backoff budget before the loop can adapt, then dead-letters the batch and ends the run. That is how one vault sat at 2914 pending changes for five days, with POST /sync/pull and GET /sync/manifest taking collateral 503s from the same isolate. A one-item batch has no smaller shape, so its 5xx is retried with the standard backoff instead of ending the run.
The size that worked is remembered on the coordinator rather than re-derived per run, so a vault refused at 100 does not re-spend those doomed requests every cycle. It is not permanent: after three consecutive full-size pushes that got a response the size doubles, back up to the configured one. A vault still too big at the doubled size pays one refused request per raise and halves again.
A per-item STORAGE_QUOTA_EXCEEDED rejection refuses only that item. The rest of the same response is still acked, and the run keeps dequeuing, because the server refuses only items that grow storage: a delete or a shrinking update still commits and is how a vault gets back under its quota. POST /sync/push has no request-level quota check: the Worker reserves storage for the growing items only and answers the ones that do not fit per item. Before #2303 it first compared the JSON length of the whole request against the quota and refused the batch with a 413, which blocked those deletes too. The paid-plan check is unaffected; it runs on every /sync/* route.
The per-item attempt budget is spent per sync cycle
A push response can accept some items and reject others. A rejected row keeps its payload and is charged one attempt; after five it stops being dequeued and no longer counts as pending.
That budget is spent at most once per push cycle. One push() call loops until the queue drains, and the loop has no backoff, so a row re-sent inside the same call would consume all five attempts across a handful of back-to-back requests — turning a few seconds of transient server trouble into a permanently parked edit that the UI reports as nothing-to-sync. Rows the server rejects are therefore excluded from the remaining dequeues of that call and picked up again on the next cycle, which is where the delay between attempts comes from. Other queued rows are unaffected: a rejected row is skipped, not treated as a barrier, so healthy items behind it still go out in the same cycle.
The exclusion lives only in memory for the duration of the call. It is keyed on row id rather than on the persisted lastAttempt timestamp, so a backwards system-clock jump can never hide a pending edit. A row that changed while in flight is a different case and is not excluded — that re-send carries the newer payload and costs no attempt.
Migration 0047 resets attempts to 0 for rows that an earlier build had already exhausted, so edits stranded by the old in-cycle spend are retried again after upgrading. It preserves the row, its payload, and its recorded error.
An item the schema refuses is rejected on its own, not with the batch
The push body used to be validated as a single value: one item that failed RecordPushItemSchema failed all 100, and the server answered 400 VALIDATION_ERROR with the first Zod issue as the message. That message names no item id. The client cannot map it to a queue row, so it marked nothing, charged no attempt, and dequeued the identical batch again on the next cycle. One note that left a device without clock metadata stopped a vault from syncing at all — 29 failed pushes in two hours, every one the same request and the same verdict.
The route now parses only the envelope (the 1-to-100 bound) as a request, then validates each item separately. A refused item comes back as one rejected[] entry with reason SYNC_INVALID_ITEM and the failing issue appended, while every valid item in the same batch still commits. An item too malformed to yield an id and a known type cannot be named in the response at all; it is dropped, and the client's existing rule — an id it sent that appears in neither accepted nor rejected counts as failed — is what retires that row.
The client stopped relying on the server for this verdict as well. Before sending, it validates each encrypted item against the same contract and retires the ones that can never be accepted, so a row the server is certain to refuse costs no round trip. The related repair that stamps a missing clock onto an outgoing payload used to give up when the payload was not readable as JSON — the exact input it exists for. A delete is now rebuilt around a first clock, which is already the fallback shape for a tombstone. A create or update is not: inventing a body for one would push an empty record over the server's copy and blank every field it holds, so those are retired instead.
A rejected signature is a device-identity failure, not a per-row failure
SYNC_INVALID_SIGNATURE is the one rejection that says nothing about the row. It means the key this install signs with is not the public key its device id is registered under, so every queued row is condemned equally and no payload change can help. Charging it as a normal rejection spent one attempt per row across the whole queue and then dead-lettered every edit silently — 497 rejections inside 20 minutes on one account, followed by edits that never synced again.
The push loop therefore treats it as a device-level verdict: the run stops at the first one, no row is charged an attempt (the queue keeps its full budget for after the repair), and the engine reports the non-retryable device_key_mismatch category once.
The repair is re-registration, because the server holds the public key the device registered with and never rotates it in place. The desktop client used to "self-heal" a key mismatch by rewriting its local sync_devices row to match the keychain, which only made local state agree with a key the server rejects. Both detection points — the startup integrity check and the runtime signing-key read — now tear the session down instead, so the ordinary sign-in flow re-registers the device under the key it actually holds. Both stand down while key material is in flux (sign-in, recovery, linking), where a transient mismatch is expected.
An attachment manifest signed by such a device is unverifiable for every reader, forever. The download path classifies ManifestSignatureError as a permanent failure so the re-driver stops probing it, rather than re-fetching the same unverifiable manifest every hour.
Dead-letter purge and the pause flag are kept off the enqueue path
Rows that exhausted their budget are purged once at least 50 of them are older than ERROR_RETENTION_DAYS. Nothing reads those rows and they can only accumulate at the pace of a failing push, so the threshold is probed on the first enqueue and every fiftieth after that rather than on every one — and the probe is a bounded existence check, not a count of the table. A bulk import therefore no longer scans sync_queue once per queued mutation. The purge itself is unchanged: same threshold, same retention window, same deletion.
Whether sync is paused is asked on every WebSocket message, every enqueue and every pull tick, so the answer is held in memory. Only the not paused answer is cached. Pause and resume write through the state manager and flip the cached answer with the row, so they take effect on the next call. Paused always re-reads, because sync_state rows are also removed outside the manager — the emergency wipe, session teardown and device re-registration all delete them — and a removed syncPaused row can only mean "not paused". A cached false can never go stale that way; a cached true could. That extra read costs nothing, because paused is exactly the state in which every caller stops early.
Recovering pushes that never landed
Items expose a "the server has this state" stamp (syncedAt) that advances on a confirmed push as well as on an applied pull. Anything modified after its stamp — or never stamped at all — is re-queued on the next full sync for tasks, projects, notes and journals alike.
Recovery re-sends the item's stored clock rather than bumping it. An item that is genuinely in step is then replay-detected by the server, costs one round trip, and is stamped clean; only an item that really is ahead of the server changes anything. Scope is limited to items the server already knows: clock-less rows belong to the initial seed.
Notes and journals share a table but not a sync service, so they are swept separately and each re-push is handed to the service that owns it. The journal sweep also carries the entry's date, which its payload builder needs to find the file on disk — recovering a journal without one would fail before the builder's own error handling and take the rest of the sweep down with it.
Because recovery never advances a clock, a change made while the sync runtime is down has to advance its own at write time or the re-push would be dismissed as a replay. Records park that tick under a placeholder device that their sync service rebinds on the way out. Rebinding runs before every record push, create as well as update: a row that was created and edited while signed out is recovered as a create, and until that path rebound too it shipped the placeholder device id to the server — an id every install claims, which makes two devices' clocks compare equal for edits that are genuinely concurrent. Notes and journals have no rebinding step, so their fallback bumps under the current device directly and does nothing when no device is registered (the same thing the online path does). It also clears the sync stamp, because metadata-only writes — recording an uploaded attachment or editing a journal's tags, say — deliberately leave modifiedAt alone and would otherwise be invisible to the "modified after its stamp" test above. A row that never leaves the device, and one with no clock yet, are both left alone.
Foreign-key parents and orphan repair
Some rows carry foreign keys — a task references its project and its status — and the data DB enforces them. Server cursor order is last-update order, not dependency order, so pulled items are sorted so FK parents apply before their children, and anything that still fails is retried once after every page has landed. The retried items' ITEM_SYNCED events go out after the whole retry loop, not one per item inside it. Each event can make the renderer refetch over IPC, and in one fresh-device bootstrap the retry of 126 items took 7 seconds, almost all of it spent waiting at the loop's yields.
That covers a parent that simply arrived late. It does not cover a parent that is gone, which is what a cascade delete produces: deleting a project removes its tasks locally through SQLite ON DELETE cascade, and a cascade is invisible to sync unless each child is tombstoned explicitly. Project deletion therefore pushes a tombstone for every task it cascades away, including completed and archived ones. Without that, the child rows stay alive on the server, every device re-pulls them, the FK insert fails, the item is skipped, the next manifest check still sees it server-only, and the cycle repeats forever.
For installs already holding such orphans, the end of a pull run repairs them. The missing parent is re-fetched by id, which is authoritative in a way the cursor window is not:
- the server still returns the parent → apply it, then the child lands normally.
- the server no longer returns it → the parent is gone everywhere, so the child is a confirmed orphan and is tombstoned. That is what the cascade should have pushed originally, and it ends the re-pull loop on every device.
Deletion is gated on that second condition alone; a child whose re-apply fails for any other reason is left untouched and retried on the next cycle. A dangling status_id is not an orphan at all — the FK is ON DELETE SET NULL, so the reference is simply cleared rather than failing the apply.
That tombstone is stamped with this device's clock before it is queued. The payload it is built from is the one just pulled, so its clock is the server's own clock for that row, and the server rejects any push whose clock has no entry greater than the one it already holds. Sent back unchanged the delete is answered SYNC_REPLAY_DETECTED, the queue row is cleared as already applied, the next pull serves the same orphan again, and the repair runs again — the loop it exists to end, running forever. A normal delete never hits this: it is built from a local row by the domain layer, which stamps the clock on the way out, and the push path sends delete payloads verbatim apart from the last-resort clock stamp described in Initial seeding, which only fires when the payload has no clock at all and so cannot lift a server clock past itself. An orphan has no local row, which is what makes it an orphan, so nothing else can stamp it. Without signing keys the stamp is impossible, and the orphan is left for the next pull rather than spending a push that would only be refused again.
Sync Type Negotiation
Clients declare the record sync item types they understand via an X-Memry-Sync-Types header (comma-separated), sent on authenticated sync calls alongside the existing X-Memry-Vault-Id. The value is RECORD_SYNC_ITEM_TYPES joined with commas. The server (/sync/changes, /sync/manifest, /sync/pull) binds only the negotiated types into its item_type IN (...) SQL filter.
| Header | Resolves to |
|---|---|
| Absent | The frozen LEGACY_RECORD_SYNC_ITEM_TYPES list (15 types) |
| Present, nothing recognized | An empty list — serves zero rows |
| Present, some recognized | The recognized subset, deduped and intersected with the server's supported types |
No header means the client predates negotiation and never declared anything, so it gets exactly the frozen legacy list — the property that protects binaries already in users' hands. This list is never edited when a new sync item type is added; adding to it would hand that type to clients whose parsers reject it, which is exactly the bug this feature exists to prevent.
A header that is present but names nothing recognized is a different situation and resolves differently: the client did negotiate, so it must never be handed types it didn't declare. Empty types short-circuit before any DB query, and getChanges returns the incoming cursor unchanged so nothing advances.
Requested types are deduped and intersected with the server's supported set, bounding the bind-parameter count against D1's 95-parameter ceiling.
Why this exists: the desktop client does not runtime-validate /sync/changes, does not filter item refs by type before pulling, and validates a pull page with a single whole-page safeParse. One unknown item type fails the entire page, the client drops it without throwing, and its cursor still advances past it — silently losing convergence for every note and task on that page, not just the unrecognized item. Published binaries cannot be patched, so the server is the only place this can be fixed.
Deploy order: the sync-server change must reach production before any desktop build carrying a new item type.
Vector Clocks (Doc-Level)
Used by the server to order changes across devices. The server itself never inspects fields — it sees a single clock per document and uses it to pick the correct write on conflict.
Field-Level Merge (Tasks & Projects)
Inside the encrypted blob, tasks, projects, and agent conversations carry per-field vector clocks (field_clocks).
- Concurrent edits to non-overlapping fields merge cleanly.
- Concurrent edits to the same field resolve last-writer-wins by the sum of device ticks (
tickSum). On a tie the incoming (remote) write wins, so which value survives depends on which device happens to merge. - A merged row is re-queued and pushed back under the union of both clocks, and a pulled row whose clock equals the local one is applied rather than skipped. Together these are what converge two devices that both merged the same concurrent pair: the first re-push is accepted, the second is refused as a replay, and the refused device takes the accepted row on its next pull.
See packages/sync-client/src/field-merge.ts for the merge implementation and docs/protocol/06-vector-clocks-and-field-merge.md for the normative rules. TASK_SYNCABLE_FIELDS is 15 fields; PROJECT_SYNCABLE_FIELDS is 8; agent conversations merge title, backend, backendModel, trustList, and pinned.
Property Definitions
property_definition is an encrypted record sync item type whose id is the property name. It carries the property's type plus its options column verbatim, as the opaque JSON string the row stores — a bare option array, or { categories } for a status property. It is deliberately opaque so a newer client's per-option field survives a round trip through an older one.
.memry/properties.md stays the human-readable file, and it is local to one machine. The data DB row is what replicates. Two consequences worth knowing before touching either:
PropertyDefinitionsService.reload()rebuilds the table from that file, and the pull coordinator calls it after every pull. The rebuild carries each row's clock across, or every definition would look unclocked andseedUnclockedwould re-push the whole set on the next sync.- A pulled definition exists only as a row until the file is written, so
reload()unions the clocked rows into the cache and persists when the union gained something — including on a device that has no file yet. A remote delete reconciles the file too, or the next reload reads the definition straight back in. - A
relationdefinition never reaches the file. The file schema has no member for it, and one such entry fails the parse for the whole file. The union skips a syncedrelationrow, the writer skips one in the cache, andreload()drops and rewrites arelationentry an older build already wrote. - No device creates or pushes a
relationdefinition row either. Note indexing does not save a definition for a property whose value resolves torelation, andseedUnclockedskips an unclockedrelationrow an older build or the CLI left in the data DB. Rows already on the server stay there; receivers drop them on reload.
property_definition is not in LEGACY_RECORD_SYNC_ITEM_TYPES; clients that predate it negotiate it away via X-Memry-Sync-Types and never see it.
Agent Chat Items
Agent chat adds two encrypted record sync item types:
| Type | Merge behavior |
|---|---|
agent_conversation | Field-level merge for title, backend, backend model, trust list, and pinned state |
agent_message | Append-only by message id; duplicate ids are idempotent |
Conversation titles, message bodies, and attachments are stored as purpose-bound encrypted JSON envelopes before sync encoding. Streaming messages are not eligible for sync until they reach a terminal status.
Cursors
server_cursor_sequence holds one counter per user. Every accepted push row takes the next value as its server_cursor. Each device keeps its own pull cursor, LAST_CURSOR, which means "every row at or below this value is applied here". Pull is incremental: fetch everything strictly after the cursor, apply it, advance, repeat.
Two rules keep that paging lossless.
- The server reserves a push's cursors inside the transaction that commits its rows. When the reservation ran in a separate D1 batch, a later push could take a higher range, commit first, and let a reader page past the lower range before those rows existed. The reader then never saw them.
- Only the pull moves
LAST_CURSOR. The push response'smaxCursorsays where this device's own rows landed, not that the rows below it were pulled. A desktop build that movedLAST_CURSORtomaxCursorskipped every peer row that was still unpulled at that moment.
Installs that ran a build with either bug may already have skipped a range. On the first full sync of a fixed build, the desktop resets LAST_CURSOR to 0 once, under the sync lock, and re-pulls its whole history. Rows it already holds at the same or a newer clock change nothing. The cursorSkipRepair sync-state key tracks the repair: pending:<cursor> after the reset, done after a pull delivered. An interrupted repair resumes from the persisted cursor on the next full sync, so a page the pull refuses every time costs one pass, not one per sync. A device with no cursor yet pulls from 0 anyway and records done without a reset.
The server fix has to be live before a desktop build runs the repair: a repair pull that races a peer push on an old Worker can skip the range again and still record done.
Note bodies on the same cursor
crdt_updates and crdt_snapshots rows also take a server_cursor from the same per-user sequence, reserved in the batch that writes the row (migration 0011). A snapshot takes a new cursor every time it is rewritten. Rows written before the migration keep NULL and never enter the feed.
A client that adds note_body to X-Memry-Sync-Types gets a noteBodies array on every /sync/changes page: update and snapshot entries in cursor order, read in the same D1 batch as the record rows so one page never skips a row committed between two reads. An update up to 4 KiB carries its bytes inline; a larger one and every snapshot are refs the client fetches from the CRDT routes. A page that carries bodies is capped at 100 rows. note_body is not a record type: it never reaches the manifest, /sync/pull, bootstrap or /sync/push, and a client that does not declare it gets the same response as before. The desktop declares it (#2297); see CRDT & Notes Sync for how it applies bodies.
End-to-End Latency Trace
A record change crosses four hops: device A queues it and pushes it, the server commits it, the user's Durable Object broadcasts changes_available, and device B pulls and applies it. Each hop logged on its own, so "where did this item spend 4 seconds" had no answer. The row's server_cursor is now the join key at every hop (#2280).
| Hop | Where | Emitted |
|---|---|---|
| origin queue | desktop A, PostHog | sync_run_completed action=push_lag, durationMs, value = maxCursor |
| push accept | Worker log Record sync push processed | vaultId, cursorRange: [min, max], itemCount |
| broadcast | Worker log Record changes broadcast (DO) | vaultId, cursor, sent, itemsSent when the push carried socket items |
| receiver apply | desktop B, PostHog | sync_run_completed action=e2e_latency, durationMs, value = cursor |
e2e_latency carries source=pull for rows a pull applied and source=socket for rows applied from socket items (below). The socket path emits one event per frame, with the frame's cursor as value, and borrows the clock offset of the latest pull; before the first pull it emits nothing.
The join: B's value falls inside the push line's cursorRange and is at most the broadcast cursor; A's value is the push response's maxCursor, the top of the same range. Cursors are per account, so PostHog events join on the person plus the cursor, and Worker lines on vaultId plus the cursor. Neither log line carries a user id, device id or item id.
The two metrics:
push_lagis push accepted minus row enqueued, both on device A's clock.sync_queue.created_atis epoch seconds, so the queue keeps a millisecond enqueue time in memory for rows queued this session and falls back tocreated_atfor older rows. This half covers the push debounce.e2e_latencyis applied minuscommittedAtMs, the server's millisecond time of the push batch that wrote the row (sync_items.committed_at_ms, migration0010). The apply time is moved onto the server clock by an offset estimated fromserverTimeMson/sync/changesand the request's round-trip midpoint; the lowest-RTT sample of the run wins, because a prefetch that overlaps a page apply reads its response late.
The product number is their sum. Both are capped at 20 events per pull or push run, skip anything older than 10 minutes (an offline backlog or a first sync is not propagation), and report a negative estimate as 0. e2e_latency counts only items applied or merged as a conflict and signed by another device: the feed serves a device's own writes back; at an equal clock with an identical payload they are skipped, otherwise they apply like a peer write. Both reuse sync_run_completed because a new event name would fail the whole telemetry batch on an older server; a chart counting sync runs must filter on action.
Compatibility: every field is optional. An old desktop strips the new ref fields, an old server sends none and the desktop then emits no e2e_latency, and rows written before migration 0010 keep a NULL commit time that the change feed omits.
Pull Scheduling and Hang Recovery
A periodic tick fires every 60 seconds; WebSocket changes_available and connected messages schedule additional pulls in between. The interval is armed before the first full sync, and a failure in that first sync is logged rather than propagated, so one transient error at startup cannot leave a session without a pull cycle.
changes_available wakes are filtered and coalesced. A wake whose cursor is at or below LAST_CURSOR is dropped: cursors are assigned in commit order and only the pull moves LAST_CURSOR, so every row it announces is already applied. The wake's cursor is only compared, never stored. A wake that arrives while a wake-driven pull is queued adds nothing, and any number that arrive while a pull runs queue exactly one trailing pull, so a peer pushing N requests in a burst no longer costs N serial pulls. The stale-lock watchdog clears the queued flag, so a pull chained behind an abandoned sync cannot swallow later wakes.
The tick does not always pull. Its pull exists to heal a changes_available broadcast that never arrived, so when the socket has been continuously connected since the previous tick — same connectionGeneration, still connected — the request is skipped: the socket pings every 25s and terminates itself after 31s of silence, so a half-open connection reports disconnected before a tick would trust it. Any drop between ticks bumps the generation and restores the every-tick pull, and a reconnect pulls on its own. A 5-minute floor caps the skipping, because a server that stops broadcasting is indistinguishable from a quiet vault from the client side. The stale-lock watchdog runs on every tick regardless.
Network status feeds the same scheduling. Electron exposes no main-process event for net.online, so it is polled — every 5 seconds while offline, every 30 seconds while online, dropping back to the fast cadence the moment the status goes offline. A returning network is therefore always detected within ~5 seconds (plus a 2-second debounce), and powerMonitor resume polls immediately rather than waiting for a tick; suspend applies offline right away, so a machine waking from sleep is already on the fast cadence.
Three guards keep a wedged sync from lasting until restart. Every sync HTTP request carries a 60-second abort timeout, so a black-holed socket (suspend/resume, NAT teardown) surfaces as a retryable network error instead of pinning the sync lock forever. If the lock is still held after 15 minutes anyway, a watchdog on the periodic tick force-releases it, aborts the in-flight run, and lets the next pull proceed. Skipped periodic pulls log Periodic pull skipped with the blocking flags, which is the first thing to look for when a device shows stale data.
Socket items
The last hop after a wake is still one HTTP round trip, for bytes the Worker held in memory when it broadcast. A socket can opt in to receiving them: desktop sends X-Memry-Socket-Items: 1 and the same X-Memry-Sync-Types its HTTP requests carry on the WebSocket handshake. When a record push commits at most 64 KiB of items, that socket's changes_available frame also carries items (each exactly what POST /sync/pull returns for the row) and committedAtMs, filtered to the types the socket declared. A larger push, a socket that did not opt in, a socket whose token has expired, and every socket accepted before this shipped get the old hint-only frame byte for byte. The Worker decides from the payload sizes before it builds any item, so an over-budget push costs nothing extra, and SYNC_SOCKET_ITEMS_MAX_BYTES="0" turns items off without a client release. A device revoked while the revoke call fails can keep receiving items until the revocation alarm closes its socket, at most a minute later.
The items are advisory. The wake pull is scheduled for every frame, with or without items, and it still owns LAST_CURSOR, quarantine, the schema-invalid ledger, corrupt re-fetch and the breaker. The socket applier only applies, one frame at a time in arrival order and at most 50 items per frame. It skips a frame while paused, offline or in a full sync, and when the frame's cursor is at or below the larger of LAST_CURSOR and the owned-through mark. That mark is the highest nextCursor of a changes page a pull has read, kept until LAST_CURSOR reaches it, even when the run stops mid-page: a pull commits a page's rows before its cursor, so without it an older socket item could re-create a row a tombstone had just deleted, and nothing would re-deliver the tombstone. A frame that meets a local push waits for it to settle, because a conflict requeue coalesced into a row an in-flight push dequeued would be deleted by that push's ack. Only the push that set that gate clears it, and the stale-lock watchdog resets it for a push it abandons.
The reverse order is an accepted transient: a frame that deletes a row after the pull fetched an older version of it, but before the page applied, is followed by the page re-creating the row. The frame's own wake pull re-delivers the tombstone within one cycle.
It decrypts and verifies signatures with the pull's batch decrypt and applies what verified through the pull's ItemApplier, so pending local deletes, vector clocks, conflict push-back, tombstone clock recording on deletes and the sync-intent deferral behave as in a pull. The frame runs in its own page session like a pull page: each item's row, conflict requeue and body debt commit together on a savepoint or not at all, and note file writes and unlinks are journaled before the commit and landed in the same synchronous run. A failed unlink (a file held open on Windows) stays journaled and the next pull's replay removes the file, so the indexer never re-adopts it as a new note. Renderer events go out after the commit and only for rows that changed. Anything that fails (signature, decrypt, schema, a missing FK parent, an undrained sync intent) is dropped without a record and reaches the device through the pull. The re-delivery of an item that did land is an equal-clock, identical-payload skip.
A frame never carries a note body or a purged tombstone: both are feed-only, and a purged tombstone applies only from POST /sync/pull with its delete attestation. A note or journal record applied from a frame and changed (applied or merged) owes its whole body exactly as a pulled one: a durable record debt plus the unmerged flag, so no snapshot push claims coversThrough past a body this device has not merged. A frame item skipped as already applied owes nothing. The wake pull re-delivers the record and its CRDT batch pays the debt.
The apply must not race a pull page. A page commits its rows and then writes note files in an async flush, so a socket write to the same file in between would leave the older file under the newer row. The applier loops until no page transaction is open, every journaled file op has landed and no push is in flight (at most 5 seconds, then it drops the frame and logs once per session with the cause), and the last check runs in the same synchronous run as the apply: nothing is awaited between it and the last file write. The journal replay at the start of each pull clears the ops it healed from memory, so one failed flush does not turn the applier off until a restart.
Socket latency events are one per changed row, like the pull's, capped at 20 per minute. The #2300 merge gate reads the combined e2e_latency p50 of both sources, not the socket source alone: the socket source sees only the frames it applied.
Runtime Emitters and Listener Budgets
Three main-process objects in the sync runtime are EventEmitters: NetworkMonitor (status-changed), WebSocketManager (message, connected, device_revoked, certificate_pin_failed, error) and SyncEngine.
Their subscriber counts are small and fixed. NetworkMonitor has three status-changed subscribers — the SyncEngine, the sync runtime itself and the attachment UploadQueue. WebSocketManager has at most one listener per event name. Nothing in the main process subscribes to SyncEngine: its status reaches the renderer through emitToRenderer, not through listeners.
Each one calls setMaxListeners(10), Node's default. That is deliberate: the ceiling has to stay close enough to the real count that an accumulating-subscriber bug trips MaxListenersExceededWarning instead of hiding behind a generous budget. src/main/sync/emitter-budget.test.ts pins both the budget and the observed counts, so raising either needs a test change and an explanation.
Every subscriber is detached on teardown: the engine removes its own in stop(), the runtime keeps a reference to its status-changed handler so stopSyncRuntime() can remove it, and the attachment UploadQueue is disposed with the runtime that built it (see "Upload queue lifetime" under Note Attachments). A subscriber left attached does more than leak: it keeps the dead note-body outbox and provider reachable for the rest of the session.
Manifest Integrity
Desktop periodically compares /sync/manifest with local syncable records. Notes and journals are matched from canonical note_metadata first, with the rebuildable index cache as a fallback, so a freshly pushed note is not treated as server-only while indexing catches up.
The full-sync log line says whether a manifest was actually diffed. fullSync: manifest check complete { rePullNeeded, serverOnlyCount } appears only after a real fetch and diff. A check that did not run logs fullSync: manifest check skipped { reason, nextEligibleAt } instead, where reason is throttled (the 30-minute window has not elapsed), no-token or error, and nextEligibleAt is the ISO time the next check may run. A skipped check never reports serverOnlyCount: 0, so it cannot be read as a verified clean result.
The comparison reads ids only. Repair payloads are built one row at a time, and only for a record the server manifest is actually missing, so the usual clean check never materializes or serializes a single row body — the cost of the check scales with the size of the disagreement, not with the size of the vault. The bytes a repair pushes are unchanged: the lazy build runs the same full-row select through the same serialization the eager pass used.
A server item counts as server-only, and so resets LAST_CURSOR to 0 for a full re-pull, only when the check can prove it is missing. Calendar events, sources, bindings and external events, folder configs, tag categories and agent conversations and messages are listed from their local tables for that direction only; they are never re-uploaded from the manifest check. A server item of a type this build does not list is never counted. Neither is an id the device declined on apply because the thing it describes is already held under another id: a second inbox project, or a calendar_source whose (provider, kind, remote_id) belongs to an existing row. Those ids are kept in sync_state under declinedSyncRefs and dropped once the server stops listing them. Before this, every live row of the unlisted types counted as missing, so each full sync outside the 30-minute throttle re-pulled the whole vault.
Manifest pagination
GET /sync/manifest pages opt-in via limit (with an optional cursor). A param-less request — which is every client shipped before this — keeps the original complete single response, nextCursor field and all absent. A cursor without a limit is a malformed pagination attempt and is rejected rather than answered with the full manifest, which would silently duplicate the pages already served.
Pagination keys on server_cursor: it is unique per user and only ever grows, so pages can neither skip nor split rows. A row updated between pages gets a new cursor greater than any page already served, so it may appear twice across the run — once at its old cursor, again at its new one — never zero times, and the client's (type, id) map dedups the repeat. nextCursor is taken from the last row kept, unsupported types included, so the type filter cannot open a gap the next page skips over. MAX_MANIFEST_PAGE_LIMIT is 1000, and the +1 row the query fetches probes for another page without a COUNT.
The sync_manifest ceiling moved from 10/min to 30/min alongside this: a paginated client spends ceil(rows / page) requests per integrity check instead of 1, so the old ceiling would have stalled any vault past 10 pages. Each paged request is a bounded indexed scan — strictly cheaper than the single unbounded full scan the old ceiling was budgeted for.
Note Attachments
Files embedded in a note (images, PDFs) live on disk under the vault's attachments/<noteId>/ folder and are uploaded to the blob store as encrypted chunks with a signed, encrypted manifest. Three mechanisms make them portable across devices:
- Reference sync — each note's payload carries
attachmentReferences, the ids of the blobs it embeds. When a device applies a note and is missing a referenced file, it downloads the blob into its ownattachments/<noteId>/folder; the filename comes from the decrypted manifest (sanitized, skipped when already materialized at the same size). Older clients parse payloads in strip mode and ignore the field. - Cross-device path remap — note blocks store the origin machine's absolute
memry-file://local/<path>URL. The protocol handler resolves a path that is outside this device's allowed roots by remapping itsattachments/<noteId>/<file>tail onto the local vault (traversal-guarded), so notes written on another OS render without rewriting note content. - Upload queue lifetime — the in-memory
UploadQueueis a module singleton owned by the IPC layer, but its lifetime is scoped to the sync runtime. It bindsgetNetworkMonitor()once, at construction, and only unsubscribes indispose(), so a queue reused across a runtime restart (vault switch, sign-out/in) would stay attached to the previous monitor. That monitor is stopped, which clears its poll timer: itsonlineflag is frozen and it can never emitstatus-changedagain, so the reconnect wake-up that clears the network backoff would be dead for the rest of the session — and a frozen offline flag makes every retry burn the full five-minute offline wait before failing.resetSyncServiceSingletons()therefore disposes the queue and the attachment service on both teardown paths (stopSyncRuntime()and the startup-failure cleanup), so the next runtime builds them against the live monitor and vault A's queue can never serve vault B. The IPC layer registers its disposer throughattachment-outbox, which is already the seam between the sync runtime and this singleton, so no import cycle is introduced. Uploads pending at dispose are rejected rather than carried over — the outbox below is what makes that safe. - Durable upload outbox — the upload intent is persisted in the data DB (
attachment_upload_queue, migration 0039) before the transfer starts and cleared only after the server accepts the file. Failed or quit-interrupted uploads are retried on every sync runtime start instead of being lost with the in-memory queue. Recording the reference enqueues a note push so peers learn the blob exists; if that lands while the runtime is down — an upload finishing during quit, a vault switch, re-auth — the note is marked for recovery instead, so the push happens at the next runtime start rather than waiting for an unrelated later edit. - Durable download verdicts — a download that does not succeed is recorded in the data DB (
attachment_download_failures, migration 0051), keyed by (note, attachment). Only the outcome writes here: the request itself no longer counts as a result, so a 404 and a success are no longer the same event. A transient failure (5xx, network, decrypt) keeps its retry on an exponential backoff from one minute. A404/410— the server saying it does not have this blob — is re-probed at most once a day, three times, and then never automatically again. Session state (in-flight claims and this session's successes) is still in memory and still cleared on runtime teardown, because the file is on disk and a re-request is cheap; the verdicts deliberately are not, since surviving a sync stop/start is the whole point. - Reference pruning —
attachmentReferencesmerges union-only, so a reference to a blob that is genuinely gone would otherwise live in the note's payload forever and be handed to every device that ever pulls it. On the pull that applies a note, ids whose 404 probes are exhausted are dropped from the stored list. The rule is positive evidence only: this device watched the server 404 that exact manifest three times, and the note has no pendingattachment_upload_queuerow (bytes still on their way up are never pruned). An id that has merely not uploaded yet has no verdict and is never touched, and nothing on disk is ever deleted. The way back is that attachment ids are minted per upload — re-inserting the file produces a new id with no verdict against it — and a local upload of the same id clears its row outright.
attachmentReferences is the only signal that tells another device a note embeds a file — the markdown link alone points at a path that exists nowhere but the authoring machine. It is sync bookkeeping, not file state, so the canonical note upsert leaves it (and the sync stamp) untouched when a caller has nothing to say about it. Ordinary vault writes — a content save, a rename, a move, a re-index — carry file state only, and must not erase it.
On the server, dereferencing only lowers a chunk's ref_count; a scheduled sweep reaps chunks at zero. It deletes the row first, and only while it is still unreferenced, then deletes the objects of the rows it removed, skipping any key a retried upload claimed in the meantime. A failed object delete only leaks storage. The sweep works in batches of 90 ids (D1 binds at most 100), 300 chunks per tick.
Tombstones
Deletions include deleted_at inside the Ed25519-signed payload — preventing a hostile server from forging deletions.
A tombstone body carries no user content. The receiving side never decodes it: ItemApplier short-circuits on operation === 'delete' and calls applyDelete(ctx, itemId, clock), and SyncItemHandler.applyDelete has no parameter that could accept the body. Handlers resolve whatever they need from the local row instead — the journal handler, for example, reads the journalled day from noteMetadata.journalDate.
So note and journal tombstones ship { clock, createdAt, modifiedAt } and nothing else: no title, no journal date. Anything more is encrypted and uploaded on every delete for no reader, and sits in plaintext in the local sync_queue row until the push drains.
clock is the one field a tombstone must keep. PushCoordinator.extractPayloadMetadata parses it back out of the payload string to stamp the server-side item version, so dropping it would break delete ordering across devices.
Payload schemas therefore mark these fields optional (NoteSyncPayloadSchema.title, JournalSyncPayloadSchema.date) — a tombstone legitimately omits them. Where a field is still required for a create or update, the handler enforces it: journal-handler.applyUpsert skips an upsert that arrives with no date.
Account Vault Directory
An account can hold several vaults (subject to the plan's vault limit). The directory lets any signed-in device see every vault on the account and pull one it does not have locally yet.
Each vault registers itself in the sync_vaults table, keyed UNIQUE (user_id, vault_id). The server stores only the ciphertext of the vault's display name:
| Column | Holds |
|---|---|
vault_id | The vault UUID that scopes all sync data for the vault |
encrypted_name, name_nonce | XChaCha20-Poly1305 ciphertext of the display name + nonce |
Names are encrypted client-side by encryptVaultName (AAD bound to the vault UUID) and decrypted locally; the server never sees a plaintext vault name. Registration is authenticated but does not require the vault to have synced any items, so a freshly created vault still appears in the directory.
Every authenticated sync call — and the WebSocket handshake — stamps the active vault's UUID into X-Memry-Vault-Id. That UUID is a single vault_metadata row, so it is resolved once and cached against the open data-database handle rather than re-read per request. The cache lives with the resolver itself, so every consumer shares it: the request header, device registration, vault-key derivation, canvas reconcile and per-attachment uploads all read the row once per open vault instead of once per operation. Opening, closing or switching a vault installs a new handle and therefore misses the cache on its own; the one rewrite that keeps the same handle — a linked device adopting the initiator's vault identity — invalidates it explicitly, so device registration and the first sync bind to the adopted vault, never the pre-adoption one.
Desktop reads the directory over IPC:
| IPC method | Purpose |
|---|---|
vault.listAccount() | Returns AccountVaultInfo[] (uuid, decrypted name, item count, local path, suggested download path) |
vault.downloadRemote(vaultUuid, parentPath?) | Clone a cloud-only vault into a local folder and open it |
The renderer surfaces this as an in-account switcher section plus a download dialog where the user picks the destination folder. A name that fails to decrypt is shown as null rather than blocking the list.
Vault account binding
Sign-out keeps every vault on disk. Before bindings, the next account to sign in on the machine synced those vaults into its own account, and sign-in with a non-empty vault open silently adopted the account's largest vault. Each vault now records the account it syncs with, and startSyncRuntime checks it before any sync service is built (apps/desktop/src/main/sync/vault-account-binding.ts).
The binding is a settings row in the vault's data DB, sync.account-binding.v1 = { userId, mode: 'sync' | 'local' }, so it travels with the folder. Older app versions ignore the key. The store's vault entry mirrors it as accountBinding for vaults that are not open.
The gate decides in this order, with userId taken from the session token's sub:
- Bound to this account:
syncstarts,localholds aslocal-only. - Bound to another account for sync: holds as
foreign. - No local content (no files outside dot-directories, no tasks, no inbox items): binds and starts, even offline.
GET /sync/vaultsfails: holds asunknown, stays sync-eligible, retries after 60 seconds.- The account already lists the vault uuid: binds and starts. This is how vaults from older versions bind without a prompt.
- A clock on a task, inbox item or note carries a device id other than this install's (or the offline placeholder): holds as
foreign. Every sign-in registers a new device id, so this is history from another session. - Otherwise: holds as
needs-decisionand the renderer asks once.
The renderer reads the state with sync:get-vault-binding, listens on sync:vault-binding-changed, and answers with sync:resolve-vault-binding:
| Choice | Allowed from | Effect |
|---|---|---|
sync | needs-decision, local-only | Binds for sync; the vault registers as its own account vault |
merge | needs-decision with a target | adoptVaultLocally onto the account's largest vault, then binds |
local | needs-decision | Records local; not asked again for this account |
foreign accepts nothing: its items carry another account's clocks and would never be seeded, so syncing one needs a fresh identity and a full re-push. Binding also drops a current-device row and cursor left in the vault by a previous session, so ensureDeviceRowForVault cannot adopt a stale device id.
Two related paths follow the binding. adoptAccountVaultIfAbsent at device registration only adopts an empty vault. refreshVaultDirectory only self-registers vaults whose store entry is bound to the signed-in account for sync, so another account's vaults never reach this account's directory.
Endpoints
| Path | Direction | Purpose |
|---|---|---|
POST /sync/push | up | Upload new sync items (metadata + blob refs) |
POST /sync/pull | down | Fetch updates since cursor |
POST /sync/crdt/updates | up | Incremental Yjs binary updates |
GET /sync/crdt/updates | down | One note's incremental updates (note_id, since, limit query params) |
POST /sync/crdt/updates/batch | down | Incremental updates for up to 100 notes in one request, plus snapshotMeta |
POST /sync/crdt/snapshot | up | Full Yjs document baseline; prunes stored updates it covers (coversThrough, else watermark) |
POST /sync/crdt/snapshot/batch | up | Up to 50 full baselines in one request; same store-and-prune semantics, reported per note |
GET /sync/crdt/snapshot/:noteId | down | The note's snapshot baseline and its revision, applied before its incrementals |
GET /sync/vaults | down | List the account's registered vaults |
POST /sync/vaults | up | Register or update a vault's encrypted name |
PUT /sync/vaults/:vaultId/icon | up | Set or reset a vault's encrypted icon; last writer wins on the client's change time |
POST /sync/bootstrap | mixed | Open an elevated bootstrap window; returns a token, the first manifest page and a tail cursor |
POST /sync/bootstrap/renew | mixed | Slide the window's TTL under the same session id |
POST /sync/bootstrap/close | up | Release the window and its per-user session slot (idempotent) |
GET /sync/packs | down | List this vault's compaction packs, newest-first, with presigned URLs when available |
POST /sync/attachments/presign-batch | down | Presigned R2 GETs for attachment chunks the caller already owns |
POST /auth/* | mixed | OTP, sign-in, refresh, sign-out |
POST /auth/oauth/google/native | mixed | Trade a platform-issued Google ID token for a setup token (mobile) |
GET /auth/key-verifier | down | Account key verifier for an established session (vault-key mismatch detection) |
GET /auth/devices | down | Signing-key directory; includes revoked devices (revokedAt set) |
POST /devices/* | mixed | Linking, listing, revoking |
POST /keys/* | mixed | Key sealing during link, rotation |
GET /auth/devices is the key directory every client verifies item and CRDT signatures against, not the device-management list (GET /devices, which hides revoked devices). It keeps revoked devices because push rejects a revoked signer, so everything signed under a revoked key was accepted before revocation and must stay verifiable. Without them, a device revoked from Settings left its whole history unappliable on every install, failing with No public key for signer.
The five /sync/crdt/* routes are the only ones that carry a note body; the record feed above them moves metadata only. A device reads a body by applying the baseline from GET /sync/crdt/snapshot/:noteId and then replaying incrementals from GET /sync/crdt/updates. POST /sync/crdt/updates/batch is the whole-vault form of that second step — up to 100 notes per request, each with its own since — but it batches incrementals only, so the baselines stay one GET per note and dominate a first-sync sweep's request count.
Native OAuth on mobile
GET /auth/oauth/google starts the browser flow and only accepts a redirect_uri that is a 127.0.0.1 loopback (the desktop app) or the configured web origin. iOS has neither, so mobile does not use it. The app signs in with Google's own SDK and posts the resulting ID token to POST /auth/oauth/google/native, which skips the authorization-code exchange and rejoins the browser flow at token validation. Everything after that point — user lookup, entitlement, setup token, analytics — is the same code path, so the two entry points cannot drift into two account models.
Two consequences worth knowing:
- The route validates the ID token against
GOOGLE_IOS_CLIENT_ID. When that binding is unset it answers501rather than falling back to the web client, because validating against the wrong audience would accept a token minted for a different application. - The client ids are also what gate the button. A mobile build without them omits the Google option from the sign-in screen entirely rather than showing one that fails when tapped.
Snapshot revisions
A snapshot baseline carries a revision: an opaque token the server replaces on every snapshot write, so a client can tell whether the server's baseline moved without downloading it.
GET /sync/crdt/snapshot/:noteIdreturnsrevisionalongsidesnapshot,sequenceNumandsignerDeviceId, so a device learns the token for the baseline it just merged. When the note has no server snapshot the response is{ snapshot: null, sequenceNum: 0, signerDeviceId: null, revision: null }.POST /sync/crdt/updates/batchreturns a top-levelsnapshotMetamap besidenotes:{ "<noteId>": { "sequenceNum": 42, "revision": "…", "signerDeviceId": "…" } }. A note absent from a present map has no server snapshot at all; an absentsnapshotMetakey means the server predates this field.POST /sync/crdt/snapshotreturns{ sequenceNum, revision }, and each accepted entry ofPOST /sync/crdt/snapshot/batchreturns{ noteId, accepted: true, sequenceNum, revision }. The token is the one that write stored, so a device that pushes a baseline records the same revision a later read would advertise instead of leaving it unset until the next pull.
The token is deliberately not sequenceNum. A replacement snapshot keeps the note's existing sequence number so later incrementals stay pullable, so the number does not move when the blob does. It is random per write rather than a counter, so a row deleted and recreated — vault deletion, account recreation — cannot reuse a token a client still holds.
Rows written before the field existed are not backfilled; the server derives a deterministic token for them at read time from the row's identity, creation time and size, and the next snapshot push replaces it with a real one. Both read paths return the same token for the same row.
All of these fields are additive: an older client reads these responses through an unvalidated cast and ignores the extra keys, and a client talking to a server that predates the push-response field stores no revision for its own push rather than inventing one.
snapshotMeta is read on the same round trip as the incrementals, as extra statements inside the batch the pull already sends. D1 refuses any single query carrying more than 100 bound parameters and answers the whole request with an error, and the metadata read binds the user and vault ahead of one parameter per note, so it is split into several statements sized under that ceiling rather than one statement naming every note in the chunk. A request-sized chunk of 100 notes would otherwise cross the line — and only a full chunk does, which is why a first sync on a new device was the one caller that hit it.
The vault sweep's conditional baseline
The paced drain (the legacy sweep and durable debts, #2421) uses snapshotMeta to stop re-downloading baselines a device already holds. It runs each chunk in two phases:
- Probe. One
POST /sync/crdt/updates/batchfor the chunk withlimit: 1, asking only whether anything moved. No document is opened, no snapshot is fetched, nothing is decrypted. A note whose baseline is unchanged and whose update list comes back empty is finished here. - Apply. Only the notes with work take the existing open-document, baseline, apply path, and a note whose baseline the probe proved redundant skips the
GET /sync/crdt/snapshot/:noteIdand resumes its incrementals from the sequence it already applied.
A baseline is skipped only when both of these hold:
- the note's
revisioninsnapshotMetaequals the revision this session actually merged, and - the sequence this session applied is at or above the note's
sequenceNuminsnapshotMeta.
The second condition is not implied by the first. sequenceNum is the server's prune watermark, and pruneUpdatesBeforeSnapshot has already deleted every update at or below it — a pull starting under that line is answered with silence rather than an error, so the note would go quietly stale.
Everything else falls through to a fetch: an absent snapshotMeta key (an older server), a note the server left out of the response, and any note this session holds no merged revision for. A missed skip costs one request; a wrong skip costs a note body.
The bookkeeping — { appliedSequence, snapshotRevision } per note, the snapshot watermark — is persisted inside the per-vault CRDT store, so the warm path survives an app restart and a fresh sign-in rather than paying a full cold sweep on every launch. It is a y-leveldb document meta key, which puts it in the same LevelDB, in the same directory, behind the same handle as the note's updates: there is no way to read a watermark without the store that holds the document it describes.
That location is forced, not chosen. A watermark that outlives its document makes the sweep skip that baseline forever against a body that never had it. Because a meta key sits inside the key range clearDocument wipes, purging a note drops its watermark in the same operation; quarantining, rebuilding or re-pathing the store moves or destroys every watermark with every document; and a store that could not be opened at all leaves the provider in memory-only mode with no handle to read a watermark through. The watermark is never written for a snapshot that was not actually received — GET /sync/crdt/snapshot/:noteId answers null when the D1 row exists but its R2 blob is gone.
The key is additive. A store written by a build that predates it has no record, which reads as unknown → fetch, never as "sequence 0 → skip", so the first sweep on such a store costs exactly what it cost before. A newer store read by an older build is inert: the older build never asks for the key. No protocol change, no D1 schema change, no IPC contract change.
Two properties this does not change. The drain stays exhaustive — every note in a chunk is still named in the probe; this changes what a note costs, never whether it is visited. And the single-note pull path (GET /sync/crdt/snapshot/:noteId then GET /sync/crdt/updates) stays unconditional: it reports whether the server's state was fully merged, and the pending-note replay turns that report into a snapshot push, which prunes peers' updates.
Pushing a snapshot is an assertion, not just a write: once the snapshot is stored, pruneUpdatesBeforeSnapshot deletes every crdt_updates row for that note at or below the snapshot's sequence number. A snapshot that does not already contain those updates does not merely fail to add them — it removes them from the server.
Batched snapshot uploads
POST /sync/crdt/snapshot takes one note per request, and the server spends six serial D1/R2 round trips on it: the note's current watermark, its existing snapshot row, a storage reservation, the R2 put, the upsert, and the prune. That is fine for the editing path, where snapshots trickle out behind a 30-second debounce, and expensive for seeding, where a fresh vault pushes a baseline for every note it owns. On a 1000-note vault each 100 bodies cost 15 seconds, essentially all of it server time.
POST /sync/crdt/snapshot/batch carries up to 50 snapshots and pays those costs once per request instead of once per note — the metadata reads become one db.batch chunked at the bind-param ceiling, the reservations become one call for the batch's summed growth, the R2 puts run concurrently, and the upserts and prunes become one batch each. Semantics per note are identical to the single-note endpoint, including the rule that a note's snapshot watermark stays where it is once a snapshot exists.
The response reports each note separately:
{
"results": [
{ "noteId": "…", "accepted": true, "sequenceNum": 42 },
{ "noteId": "…", "accepted": false, "reason": "…" }
]
}Entries are returned in request order, and one note failing does not fail the others. A storage quota that the batch as a whole cannot satisfy still rejects the whole request, exactly as the single-note path does — nothing partial lands.
Two notes never ride a batch:
- A note with unmerged remote state. The batch endpoint prunes stored updates below the new watermark, so a device that merged around server state it could not verify must not assert "I contain everything up to here". Those notes take the non-pruning
POST /sync/crdt/updatesroute, the same way the single-note path already routes them. - Anything pushed to a server that predates this endpoint. Such a server answers 404 here and 200 for the single-note route. The first 404 latches for the session and every later chunk falls back to one request per note, so an old server costs one wasted request rather than one per batch. A 413 also falls back per note but does not latch: the aggregate body being too large says nothing about which note is oversized, and only the per-note path can name it to the user.
Deferred snapshots ride the batch too
The desktop's snapshot scheduler (30 s quiet period, 120 s ceiling) sends the notes that come due within a 2-second window as one call, which the provider packs into batch requests. A doc the LRU evicts hands its pending snapshot to that scheduler instead of pushing it on the spot.
Both used to send one POST /sync/crdt/snapshot per note. Eviction is one doc per open, so a vault written in bulk (an import, a seed) pushed a snapshot for nearly every note it wrote: 200 notes spent about 120 requests in ten seconds, on the same per-device crdt_push budget (300/min) as /sync/crdt/updates, and drew 429s. The same 200 notes now take a handful of batch requests.
Nothing is weakened by deferring. The body already reached the server through /sync/crdt/updates; the snapshot is only a compaction point, and a close-time push that failed was dropped anyway. A close an editor or teardown asks for still pushes before it returns, and the shutdown flush clears the deferral first, so an eviction during teardown pushes directly.
CRDT update sizing
POST /sync/crdt/updates is bounded by two different server limits, and the client (src/main/sync/crdt-payload.ts) plans every batch against both:
- Per update. Each update is stored as a BLOB inside a D1
crdt_updatesrow, so an update can never exceed D1's 1 MB row limit. The route's own 5 MB check is not the binding one. - Per request.
/sync/*bodies are capped at 8 MiB, which limits how many updates one POST can carry.
A batch that exceeds the request budget is split across several POSTs rather than truncated. A single update too large for a D1 row cannot use the incremental path at all; the client pushes the note's full document to POST /sync/crdt/snapshot instead, which is R2-backed and therefore not subject to the row limit. The local Y.Doc already contains those operations, so the snapshot carries them, and every client version already applies the snapshot as its baseline before pulling incrementals. If that fallback fails the push rejects, leaving the batch buffered for the next flush and — on quit — recorded for replay. No path discards an update.
The two caps do not compose cleanly at the top end. A payload the route would reject with the precise Snapshot exceeds 5MB limit never reaches the route once its encoded body passes 8 MiB: the body-limit middleware answers 413 first, and with it the route's snapshot_rejected event — the one carrying totalBytes — was lost. The middleware therefore emits that event itself for /sync/crdt/snapshot and /sync/crdt/updates, with reason: 'body_limit_exceeded' and the observed encoded body size (base64 plus the JSON envelope, roughly 4/3 of the decoded payload). An oversized CRDT payload is diagnosable whichever check catches it. On the client a bare 413 with no storage code maps to note_too_large, which names the note in its toast — it is not, and must not read as, a storage-quota problem.
Retried CRDT pushes are stored once
A push that times out after the server committed it is retried with the same bytes. The server stores the SHA-256 of each update (migration 0012), and a unique index on the note and that hash makes the retried insert a no-op. The response carries the sequence number the stored row already has, so the client gets the answer the first attempt would have given, and storage is charged once. Every update carries a fresh encryption nonce, so identical bytes always mean a retry. Rows written before the migration have no hash and are never matched. Quota is reserved only for the bytes the note does not already hold, so a retry is answered even when the quota filled up after the first attempt.
CRDT write notifications
Both CRDT write paths notify peers the same way. Once the write is durable, the server broadcasts crdt_updated carrying the note id to every socket on that vault except the pushing device. The frame also carries the highest server_cursor the write reserved, omitted when it stored nothing new; clients must not use it as their pull cursor. A desktop whose legacy body sweep is done treats a frame with a cursor as a wake and pulls the change feed, which serves the body (#2421); before that, or for a frame without a cursor, it pulls that one note. A body write that does not broadcast stays invisible until the receiving device's next pull.
The symmetry matters most for POST /sync/crdt/snapshot, which is not only the oversized-update fallback. Edits made while signed out do enter the local Y.Doc, but with no session they are never enqueued as incremental updates, so on the next sign-in that whole accumulated state can only leave the device as a snapshot. Before the snapshot path broadcast, such a push was stored and announced to nobody: the peer's socket stayed connected and silent, its 60-second tick pulled only the record feed, and the backlog appeared solely once someone typed one more character — that produced an incremental, which did broadcast, and the pull it triggered picked up the snapshot's content too.
The snapshot broadcast is issued after the snapshot is stored and after pruneUpdatesBeforeSnapshot has run, never between them, so no peer is told to pull while the superseded updates are still being removed. Delivery is best-effort on both paths: the write has already succeeded, so a failed broadcast is captured in the background rather than returned to the client, which would otherwise retry a write that already landed.
Rate limiter mechanism
Every rate-limited endpoint shares one fixed-window middleware (createRateLimiter), and its counter lives in a RateLimiter Durable Object — one instance per bucket:identifier key — not in D1. The previous implementation paid a 2-statement db.batch against a single rate_limits row per bucket on every request, which meant write contention exactly when a fresh device hammered the API. The DO keeps identical semantics: a request older than the bucket's window starts a fresh window at count 1, anything else increments, and the middleware compares the count against the bucket ceiling. Nothing client-visible changed — same bucket names, ceilings, and windows, the same 429 body, and Retry-After still reports the exact seconds left in the window.
Two properties of the split matter. The DO only counts; the ceiling comparison stays in the middleware, which is what lets a request-scoped elevation hook (getElevatedLimits, the seam for bootstrap-session elevation) widen the effective ceiling without touching the counter. And failure stays fail-closed: a missing binding or DO error blocks the request with a 500, exactly as a D1 error did before.
The record push bucket, sync_push, allows 300 requests per 60 seconds per device, the same ceiling and key as crdt_push. It used to be 60 per 60 seconds per account, so every device on the account shared one budget: three devices editing at once split 60 pushes a minute. A request without a deviceId falls back to the userId → IP key. sync_push gets no bootstrap elevation. sync_changes is still 60 per 60 seconds per account; it can move to a per-device key once wake-driven pulls are coalesced.
The rate_limits table still exists: previously-deployed code writes it during a deploy window, and the OTP per-email limiter plus the telemetry exception budget still use it. It is dropped only after those two migrate.
CRDT rate limits
The three CRDT limiters (crdt_push, crdt_pull, crdt_batch_pull) key their buckets by deviceId, not by account. Body sync is device-local work: each device pulls the note bodies it does not already hold, so a second device on the same account is normal use rather than contention. Under the default per-user key the two devices split one budget, and a legitimate first sync on device B made device A's ordinary syncing start failing with 429s. A request that arrives without a deviceId keeps the existing userId → IP fallback, so nothing becomes less strict.
crdt_push allows 300 requests per 60 seconds. Batched snapshot uploads changed what that budget buys: a push iteration dequeues at most 100 creates and a batch carries 50, so an iteration yields at most two snapshot requests, and iterations are serial. Seeding a 1000-note vault therefore costs roughly 20 requests rather than 1000, which is why crdt_push receives no bootstrap elevation — BOOTSTRAP_ELEVATION_MULTIPLIERS stays pull-only, and pushes keep their abuse ceilings.
crdt_pull allows 600 requests per 60 seconds, which is sized for one device pulling an entire vault's bodies after a fresh sign-in. That sweep costs two GETs per note — snapshot plus incrementals — so a 121-note vault spends roughly 242 requests within a few seconds, and the ceiling leaves room for a vault twice that size plus the editing traffic running alongside it. The client paces and batches the sweep itself; this limit is the safety margin for when that pacing is wrong or missing, not the mechanism that shapes the traffic.
Server base URL
Every path above is appended to a single resolved base URL. resolveSyncServerUrl() (src/main/sync/sync-server-url.ts) is the only resolver — sync HTTP, OAuth sign-in, canvas assets and attachment transfers all call it, so one env var cannot end up with two policies.
Two properties of that resolver are load-bearing:
- Resolved per call, never at import time. The main process applies
.env.<environment>via dotenv inindex.tsafter the IPC handler modules are imported, so a module-levelconst URL = process.env.SYNC_SERVER_URL || …freezes to the fallback before the env file lands. Indevthe fallback happens to equal the configured value, which hides the bug; indev:stagingit silently pinned sync and OAuth to localhost. - Trailing slashes are stripped. Callers build paths as
`${base}${path}`, so a slash-terminatedSYNC_SERVER_URLyieldshttps://host//sync/push. Cloudflare Workers routes the doubled slash as a different path, so the request 404s instead of reaching its handler. Only trailing slashes are normalized — scheme, host, port and any base path are left verbatim so a typo still fails loudly rather than being rewritten into something that "works".
SYNC_SERVER_URL is required. The http://localhost:8787 fallback applies only when NODE_ENV is development or test — the unpackaged dev server and the vitest/Playwright harnesses. A packaged build has NODE_ENV undefined and gets an explicit configuration error rather than a silent dial to a localhost port nothing is listening on. Packaging cannot legitimately omit the value: scripts/build-packaged-app.js refuses to build without apps/desktop/.env.production and asserts the value is a non-local HTTPS URL.
Bootstrap Sessions
A device that has never pulled a vault has to move the whole thing. Every steady-state rate ceiling on this server is sized for a device that already has its data and is exchanging deltas, so a fresh device's first sync is the one workload those ceilings actively fight. A bootstrap session is a time-boxed, per-device window in which the pull-side ceilings are widened — and nothing else changes.
| Route | Purpose |
|---|---|
POST /sync/bootstrap | Open a window. Returns the session token, the first manifest page, tailCursor, and (when the deployment can presign) a first page of attachment chunk hashes |
POST /sync/bootstrap/renew | Slide the TTL under the same session id |
POST /sync/bootstrap/close | Drop the ledger row and free the per-user slot; idempotent |
They are mounted as their own router under /sync/bootstrap because Hono's use('*') does not leak between routers and these routes need their own stack: authMiddleware, clientGateMiddleware, paidSyncMiddleware, syncTypesMiddleware — the same gates as the rest of the vault pull path. Their own limiter bucket, bootstrap_session, allows 30 requests per minute keyed by device, since eligibility is per device and open/renew/close are once-per-run operations rather than hot-path traffic.
The token
An HMAC-SHA256 signature over a base64url JSON payload carrying {v: 1, userId, deviceId, vaultId, jti, iat, exp} — the same shape as the checkout token, keyed by its own secret binding. Per-purpose HMAC secrets are this codebase's pattern (OTP_HMAC_KEY, WEBHOOK_HMAC_KEY, TELEMETRY_HMAC_KEY): one shared key across token classes would let a leak in any of them forge all of them.
Verification reads no database. exp is inside the signature, so an expired token fails verification with zero I/O on the hot path, and the token can be checked on every elevated request without a round trip. verifyBootstrapSession returns null on any failure — wrong secret, tampered payload, expired, malformed — and never throws: an invalid bootstrap header must not be able to fail an unrelated sync request.
iat is carried in the payload but never inspected. The only temporal check is if (payload.exp <= nowSeconds) return null, so a token whose iat is in the future verifies normally as long as its exp has not passed. That is harmless — iat is informational, only the signing key's holder can set either claim, and exp alone bounds the window — but a not-before check is not something this token has, and the function's own doc comment claiming "future-dated" among its rejections is wrong.
The one thing statelessness cannot express is a cap on concurrent sessions, which is what the bootstrap_sessions ledger (migration 0007_bootstrap_sessions) exists for. It is never on the verification path. It is written at issuance, renewal and close, deleted per (user, vault) by vault-deletion revocation, and pruned two more ways: issuance runs a lazy DELETE FROM bootstrap_sessions WHERE user_id = ? AND expires_at < ? for that user only before its insert, and the 6-hourly cron runs cleanup_expired_bootstrap_sessions, which deletes every expired row account-wide. An abandoned session therefore costs one row until its expiry passes, never a live cap slot.
What elevation actually does
The limiter middleware asks a request-scoped hook (getElevatedLimits) for a multiplier and compares the count against ceiling × multiplier. The Durable Object still only counts; the comparison stays in the middleware, which is exactly what lets elevation widen a ceiling without touching the counter — see Rate limiter mechanism.
| Bucket | Steady state | Multiplier | Elevated | Cost per request |
|---|---|---|---|---|
crdt_pull | 600/min | ×5 | 3000/min | One indexed D1 read, at most one R2 read |
crdt_batch_pull | 30/min | ×5 | 150/min | One indexed query |
blob_download | 600/min | ×5 | 3000/min | One indexed D1 row + one R2 class-B read |
sync_pull | 120/min | ×3 | 360/min | Up to 100 R2 reads per POST (pullItems caps concurrency at 25) |
sync_changes | 60/min | ×3 | 180/min | One indexed scan per page of 500, refs only |
sync_manifest | 30/min | ×3 | 90/min | Bounded indexed keyset scan per page |
Everything else — every push, every upload, status, vaults, the socket, and presign issuance itself — stays at its steady-state ceiling whether or not a valid token is present. Bootstrap is a pull problem, and elevating a write path would only widen an abuse ceiling.
sync_pull gets the smaller multiplier because it is the expensive one: 6 requests/second against bursts of 25 concurrent R2 reads is roughly 150 R2 operations/second worst case. That is bounded by the TTL and the two-session cap, and is still below what a single warm push batch spends.
Nothing is granted — the client assumes ×5
The multipliers above are the server's table and never travel. POST /sync/bootstrap returns no elevationFactor: the field is absent from BootstrapOpenResponseSchema (packages/contracts/src/bootstrap-api.ts) and the route never sets it. The client probes the open response for one anyway and, finding none, always falls through to DEFAULT_ELEVATION_FACTOR = 5 (src/main/sync/bootstrap-session.ts). Every client-side pacing site therefore runs at a hard-coded ×5, including against sync_pull, sync_changes and sync_manifest, which the server only widens ×3.
That is safe rather than tuned, because the client's pacing sites and the ×3 buckets barely overlap. The CRDT sweep — the one site that paces continuously — spends crdt_batch_pull and crdt_pull, both ×5 buckets, and crdtSweepChunkDelayMs divides each of its three slice terms by the factor, so its own derivation's "at most 50% of the bucket" property survives elevation exactly. The pull path is not client-paced by this factor at all: grep getBootstrapElevationFactor and every call site is one of the three rows in the table below, or the session module notifying its own listeners. The pull is bounded by page size and the server's own limiter, which is where a ×5 assumption against a ×3 ceiling would surface: as a 429, which http-client turns into a RateLimitError carrying Retry-After. The pull coordinator rethrows that rather than swallowing it, so the run ends and the next sync cycle retries — a delay, not a lost pull.
The probe is worth keeping: it is the seam a future server-negotiated factor would arrive through without a protocol change. As shipped it does nothing, and the number in the client is not the number in the server.
The factor is read at charge time in exactly one place
| Pacing site | How it takes the factor | Reverts when the session ends |
|---|---|---|
CRDT sweep per-chunk delay (full-sync-runner) | crdtSweepChunkDelayMs(cost, getBootstrapElevationFactor()), called inside the chunk loop | On the very next chunk |
Attachment download queue (sync-attachment-handlers) | Seeded once, then push-updated through onBootstrapElevationChange(...) | On the next notification, before the next request |
Pack download pacer (full-sync-runner) | pacer.setMultiplier(getBootstrapElevationFactor()) once per bootstrap run, no subscription | Not within the run — the multiplier is cached until it ends |
Only the sweep reads it at charge time. The pack pacer's caching is deliberate and commented as such: a session that closes mid-run only narrows the factor, pack transfers are tens of large objects rather than thousands of small ones, and those transfers are presigned GETs direct to R2 that spend no Worker bucket at all.
Requests carry the token in X-Memry-Bootstrap-Token. Identity is re-bound to the authenticated context of the request it arrives on, so a token replayed from another device or another vault elevates nothing.
Eligibility, TTL and the lifetime cap
POST /sync/bootstrap is only granted to a device that has never completed a pull for this vault: device_sync_state.last_cursor_seen is absent, NULL or 0. updateDeviceCursor only writes when a changes page actually delivered items, so that is a genuine never-pulled device — the same signal the client's own LAST_CURSOR gate uses. An already-synced device gets BOOTSTRAP_NOT_ELIGIBLE (409).
Any GET /sync/changes page that delivers an item therefore spends eligibility, including one that is not part of a pull. The desktop launch probe for a revoked device (checkDeviceStatus) used to read /sync/changes?limit=1, so every fresh desktop device got the 409 on its first full sync. It now reads GET /sync/status, which sits behind the same authMiddleware revocation check and writes nothing. It is not a single-row read: getSyncStatus also counts the vault's sync_items rows above the device cursor, which for a fresh device is the whole vault, served by idx_sync_user_cursor. A client that sends identification headers also costs one getClientPolicy read. Desktop builds that still send the old probe keep losing the session; that costs them elevated limits, not data.
| Constant | Value | Why |
|---|---|---|
BOOTSTRAP_SESSION_TTL_SECONDS | 60 minutes | Per-token lifetime; slides on renewal |
MAX_BOOTSTRAP_SESSION_LIFETIME_SECONDS | 6 hours | Absolute ceiling from the ledger row's created_at |
BOOTSTRAP_RENEW_LEAD_SECONDS | 5 minutes | How far before expiry the client renews |
MAX_CONCURRENT_BOOTSTRAP_SESSIONS | 2 | Per user, counted over unexpired rows only |
Renewal is sliding and identity-bound: a still-valid token exchanges for a fresh TTL under the same jti, so the ledger row simply extends rather than churning through insert/delete. An expired or revoked token is refused — renewal is not resurrection — and a token presented by any other user, device or vault is refused with BOOTSTRAP_IDENTITY_MISMATCH (403) instead of being extended. Bearer possession alone never renews anything.
The absolute lifetime exists because the per-token TTL slides. Without a hard ceiling a client that never finishes its pull, and never fires its close sweep, could hold an elevated window open indefinitely. Six hours comfortably covers any real full-vault pull — the pacing budgets assume hours, not days. Past the ceiling, /renew answers BOOTSTRAP_SESSION_EXPIRED (403) and drops the ledger row: the session can never come back, so keeping it would only pin a cap slot until its TTL passed. The client treats this like any other failed renewal — close locally, revert to steady-state pacing.
The concurrency cap is taken atomically. The whole decision lives in one statement — the COUNT subquery is evaluated as part of the INSERT's own execution, and D1 serialises write statements, so there is no interleaving point left between check and take. changes === 0 means the WHERE refused, i.e. the user is at cap, answered BOOTSTRAP_SESSION_LIMIT (429). The earlier prune-then-count-then-insert sequence had a TOCTOU window where two concurrent opens could both see one free slot.
Errors are logged with the caller's identifiers only, never the mismatched token's, so the warning line cannot be used to enumerate other users' identifiers.
Documented limitation: a closed token still elevates for up to 60 minutes
Elevation is stateless by design — it reads the signature and exp, never the ledger. A token whose session was closed through /sync/bootstrap/close, or one that was stolen, therefore keeps elevating from its own device until its exp, which is at most 60 minutes away.
This is accepted rather than fixed, and the reasoning is what bounds it:
- The residual is hard-capped by the TTL and cannot be extended or moved. Since renewal and close became identity-bound, no other user, device or vault can renew it — or even close it. It can only be waited out.
- Every elevated bucket is pull-only. The worst case is a device reading its own vault's ciphertext faster than steady state allows, for under an hour.
- Making it stateful would put a D1 read on the hot path of every rate-limited request, which is the cost the token was designed to avoid.
Closing early is still worth doing, because /close frees the per-user session slot for the account's other devices.
The 501 degradation
BOOTSTRAP_SESSION_HMAC_KEY is optional. When it is absent:
- All three endpoints answer a typed
BOOTSTRAP_UNAVAILABLE(501), mirroringSTORAGE_PRESIGN_UNAVAILABLE. bootstrapRateLimitElevationreturns null everywhere, so every bucket keeps its steady-state ceiling.
That is byte-for-byte today's behaviour. The client's failure discipline makes it invisible: a 404 from an old server, a 501 from an unconfigured one, a 409 from an already-synced device, a 429 at the cap, a malformed body, or a plain network error all resolve to "no bootstrap here" — the client clears local state and syncs exactly as it did before the feature existed. The token can only ever widen a ceiling, so losing it can never lose data.
Completing the bootstrap and releasing the session are two decisions
They used to be one, and #1835/#1837 pulled them apart because every path where completion is legitimately impossible was leaking the per-user session slot.
Marking the bootstrap complete is a claim that every note body the server holds is now current on this device. maybeMarkBootstrapFullText() fires only when four things hold at once: sweepSettledOnThisEngine, bootstrapPullSucceeded, no paced CRDT chunk in flight, and both the paced queue and the pending set empty.
sweepSettledOnThisEngine is set once a full sync ends online with a CRDT store, after queuing the legacy sweep if it was owed (#2421 removed the throttled vault sweep it used to wait on). Offline does not settle it: it means "nothing is fetchable", never "nothing is outstanding".
bootstrapPullSucceeded is the other half, and is why PullCoordinator.pull() and SyncEngine.pull() return Promise<boolean> rather than Promise<void>. On a fresh device an empty index DB makes every drain trivially empty whether or not the pull failed, so "queue empty" only becomes "bodies delivered" once a pull has reported that it actually delivered.
Releasing the elevated session is a resource concern, and takes any of these paths:
| Path | Trigger | Reason logged |
|---|---|---|
| Completion | maybeMarkBootstrapFullText() — all four gates hold | completed |
| Stalled drain | Empty paced queue, nothing in flight — whether the drain finished or ended owing notes back | idle |
| Blocked drain, terminal | No crdtProvider; fixed for the engine's life, so the block cannot lift and the dwell is skipped | idle |
| Blocked drain, transient | Blocked continuously for BOOTSTRAP_DRAIN_BLOCKED_DWELL_MS = 2 minutes | idle |
| Run threw | The catch in run(), which also abandons the telemetry window if no pull had resolved yet | failed |
| Vault switch / runtime restart | dispose() | vault_switch |
The dwell is what separates the two blocked cases. A network block lifts by itself constantly, so cutting the session on a first blocked tick would revert pacing in the middle of a bootstrap that is about to resume; a missing CRDT provider never lifts, so waiting buys nothing. An active fullSyncActive is exempt from the blocked check entirely.
dispose() also abandons the bootstrap telemetry window — but only one this runner owns (bootstrapWindowOwned). The window is a module global and beginBootstrap no-ops while one is set, so an unconditional abandon during a vault switch would delete the incoming vault's window: downloadRemoteVault arms it before selectVault closes the outgoing engine.
On every path, local state is cleared first — the token is captured, then clearBootstrapSessionState() runs and notifies the factor listeners, so pacing reverts in the same tick. The POST /sync/bootstrap/close call follows best-effort with the captured token, and its failure is logged at debug: the token dies on its own TTL, so a lost close request is harmless. The reason is local; it is logged and never sent.
What the open response carries
Besides the session, POST /sync/bootstrap answers with:
manifest— the first page of the paginated manifest service (MAX_MANIFEST_PAGE_LIMIT), never the whole vault.tailCursor— the currentMAX(server_cursor)for the vault, so the client knows when its pull has caught up.attachments— present only when the deployment can presign. ItschunkHashesis the first keyset page of the vault's ciphertext chunk hashes, capped at 512, and is informational only: no continuation endpoint ships, sonextChunkCursornames where continuation would start and a client must not treat the page as a complete inventory. URLs come fromPOST /sync/attachments/presign-batch, which keeps this response bounded no matter how attachment-heavy the vault is.packs— always an empty array, and now permanently so. The route returnspacks: []unconditionally. It was reserved so the pack pipeline could plug in without a protocol change, but the pipeline did not use it: the client discovers packs throughGET /sync/packsinstead, and reads this field only to log its length. It is dead weight inBootstrapOpenResponseSchema— it cannot be removed without a contract change, so it stays, but nothing should be built on it.
Presigned R2 Transfers
R2 speaks the S3 protocol, so an object can be handed to a client as a plain SigV4 query-string presign with region auto, service s3 and payload hash UNSIGNED-PAYLOAD. That turns a chunk or pack transfer from client → Worker → R2 into client → R2, taking the Worker out of the byte path entirely.
The presigner is hand-rolled (services/r2-presign.ts) rather than pulled from aws4fetch or the AWS SDK: it is about a hundred lines on Web Crypto, there is no dependency to audit for one HMAC chain, and the signature path is pinned byte-for-byte against AWS's published known-answer vector. It is deliberately two layers — presignS3Url is pure protocol with no policy, which is what lets the vector pin it, and presignR2Url is the deployment policy wrapper that derives the path-style address and clamps the TTL.
Where URLs are issued
| Site | Method | Scope |
|---|---|---|
POST /sync/attachments/upload/initiate | PUT | One URL per chunk hash the client declares (≤128) |
POST /sync/attachments/presign-batch | GET | One URL per chunk hash the caller owns (1–1024) |
GET /sync/packs | GET | One URL per pack on the page (≤50) |
Issuance has its own bucket, blob_presign, at 120/min — an order of magnitude under the chunk ceilings it serves, and deliberately not elevated by a bootstrap session. Presigning is cheap (one HMAC chain plus one indexed D1 read per hash) but each issued URL unlocks a direct transfer that bypasses the proxied buckets entirely, and one batch already arms up to a full manifest's chunks per TTL window.
What a URL is scoped to
One object, one method, one bucket, for five minutes. DEFAULT_PRESIGN_TTL_SECONDS and MAX_PRESIGN_TTL_SECONDS are both 300, and presignR2Url clamps whatever it is handed into [1, 300]. Only the host header is signed.
Scope enforcement is structural, not checked. Clients send chunk hashes; the R2 key comes back from a blob_chunks row selected by user_id AND vault_id, which is exactly the ownership check the proxied chunk GET performs. A foreign vault's or another user's hash is simply "not found" (404). Clients never submit key material, so a cross-vault scope escape has nowhere to enter.
assertPresignKeyInVault then runs immediately before signing, on the exact bytes about to enter the URL, requiring the key to start with <userId>/vaults/<vaultId>/. That is defence in depth: a presigned URL bypasses every other auth check the Worker performs, so the last thing before the signature is a prefix assertion.
A leaked URL exposes only end-to-end encrypted ciphertext. It is still a credential and is treated as one — see Presigned URLs are credentials.
Two properties of the upload direction follow from the Worker not seeing the bytes:
Content-Lengthis not covered by the signature, so an armed PUT URL accepts any number of bytes. At/completeevery direct chunk ishead-verified against R2 for existence and exact byte count before quota is credited, and an object larger than the whole session's ciphertext budget is reclaimed immediately.- An armed URL that is never registered leaves an invisible object. The expiry sweep walks
uploaded_chunksand the orphan sweep walksblob_chunks, and a presigned PUT appears in neither. Migration 0008 addsupload_sessions.presigned_chunks— the JSON array of hashes armed at initiate — and complete, abort and the expiry sweep each delete every armed hash that has no liveblob_chunksrow (services/presigned-chunk-reclaim.ts). Theblob_chunkscheck is load-bearing: a client may legitimately declare the hash of a chunk that already exists, and deleting that object would destroy another attachment's data. The column is nullable and deliberately not backfilled — NULL means no URLs were ever armed, which is what every row written by the old server, by a proxied-path client, or on a deployment without the secrets holds.
Configuration and the graceful fallback
| Binding | Value |
|---|---|
R2_ACCESS_KEY_ID | R2 API token id |
R2_SECRET_ACCESS_KEY | Its secret |
R2_S3_ENDPOINT | https://<account-id>.r2.cloudflarestorage.com |
R2_S3_BUCKET | The bucket bound as STORAGE in wrangler.toml |
None of the four is declared in wrangler.toml — there is no [vars], [env.staging.vars] or [env.production.vars] entry for any of them. They appear only in .dev.vars.example and in the Bindings type in src/types.ts. All four are therefore supplied per environment out of band, and the usual var-versus-secret split does not apply to any of them; the table above deliberately does not claim one. Scoping the API token narrowly — one bucket, Object Read & Write — is a sound deployment practice and nothing in the code checks it, so it is a recommendation, not an invariant the Worker can enforce.
resolveR2PresignConfig returns null unless all four are present, the endpoint parses as a URL, its protocol is https:, and it has a host. A trailing slash on the endpoint is normalised away, because it would double up in the canonical path and break the signature.
A null config is not an error anywhere. presign-batch answers a typed STORAGE_PRESIGN_UNAVAILABLE (501); upload/initiate simply omits chunkUrls and the session stays valid on the proxied path; GET /sync/packs omits url and expiresAt. Clients read all three as "use the proxied path". Old servers answer 404 to the presign route, which clients treat identically.
R2_S3_BUCKET must name the same bucket the STORAGE binding points at. Nothing in the code cross-checks this — the Worker cannot read a binding's bucket name — so a mismatch fails at /complete instead: the presigned PUTs land in the other bucket, the Worker's head against STORAGE finds nothing, and every affected chunk is rejected with UPLOAD_INCOMPLETE (400). That is fail-closed by construction, since quota is only ever credited for storage the Worker has verified exists, but it is also silent until an upload completes. Verify the pair on every environment before enabling the secrets.
Pack Discovery
Lists the compaction packs available for the caller's vault. A pack is an immutable byte-concat of already-encrypted blobs plus a trailing index block — see Vault Packs for the format, the compaction pipeline and the client's apply path. Everything here is additive: old clients never call it, and the item-granular endpoints remain the source of truth.
Auth and gating. The route sits on the sync router after paidSyncMiddleware, so it carries the same stack as /sync/pull: authenticated, client-gated, paid-gated, and scoped to the X-Memry-Vault-Id header. Every query filters on user_id AND vault_id; there is no cross-vault read path. It is registered at both /sync/packs and /sync/records/packs, mirroring the other record routes. Its bucket is sync_packs at 60/min, not elevated by a bootstrap session — a bootstrap lists packs a handful of times, not continuously.
Pagination is keyset on (max_cursor DESC, id DESC), so packs arrive newest-first. limit defaults to 20 and is clamped to MAX_PACK_PAGE_LIMIT = 50. The cursor token is "{max_cursor}:{id}" and is validated at the route against ^\d+:\S+$. The composite token is what keeps tie-grouped rows — two ranges sharing a max_cursor — stable across pages: a bare max_cursor < ? filter could skip a row or serve it twice as pages churn. nextCursor is present only when another page exists.
Response is { packs: PackSummary[], serverTime, nextCursor? }. Each summary carries:
| Field | Meaning |
|---|---|
id | pack_index row id; also the second half of the page cursor |
itemKind | record | crdt_snapshot | crdt_update |
packKey | R2 object key |
minCursor/maxCursor | Range bounds on the kind's ordering axis |
itemCount | Entries written into the pack (holes are not counted) |
byteSize | Payload-region bytes only — header, index block and footer are excluded, so it is not a file length and must not be used to Range-request |
createdAt | Epoch seconds |
url / expiresAt | Presigned GET and its expiry, both present or both absent |
A row whose item_kind fails schema validation is dropped from the page rather than served: an unknown kind must not silently reach a client that switches on it.
An absent url means the deployment cannot presign, which is the same graceful degradation every other presign site takes. There is no proxied pack GET to fall back to — the client filters those packs out of its listing entirely and bootstraps through the item-granular endpoints. An expiresAt already in the past (a page held too long) has the same effect; the client applies a 30-second clock-skew margin and re-lists to mint fresh signatures when a queued pack ages out mid-run.
Deploy Prerequisites and Ordering
This epic adds bindings, queues and migrations that are individually optional, plus one ordering constraint that is not.
Per-environment Worker secrets — R2_ACCESS_KEY_ID, R2_SECRET_ACCESS_KEY. Absent (or partially set) means no presigned transfers anywhere; everything stays on the proxied paths.
Per-environment vars — R2_S3_ENDPOINT, R2_S3_BUCKET. Neither is declared in wrangler.toml, so both are set out of band per environment. R2_S3_BUCKET must match the STORAGE binding's bucket; see the fail-closed note above.
Optional secret — BOOTSTRAP_SESSION_HMAC_KEY. Absent means typed 501s and zero elevation.
Queues — memry-pack-compaction-staging and memry-pack-compaction-production must exist before deploying those environments. Queue bindings are not inherited from the top-level wrangler.toml block, so each environment wires its own producer and consumer; a missing producer makes every enqueue a silent no-op and a missing consumer leaves messages unconsumed. The same is true of the RATE_LIMITER Durable Object binding, which is fail-closed — an absent binding 500s every rate-limited request in that environment. wrangler.test.ts asserts all of this against the TOML.
Migrations apply on deploy. There are two independent files numbered 0007 — 0007_pack_index and 0007_bootstrap_sessions — plus 0008_upload_sessions_presigned_chunks. All three are additive:
| File | Adds |
|---|---|
0007_pack_index | Tables pack_index, pack_watermarks; indexes idx_pack_user_vault_cursor, idx_pack_user_vault_min, and idx_crdt_snapshots_created |
0007_bootstrap_sessions | Table bootstrap_sessions; indexes idx_bootstrap_sessions_user_expires, idx_bootstrap_sessions_vault |
0008_upload_sessions_presigned_chunks | Nullable column upload_sessions.presigned_chunks, deliberately not backfilled |
Three new empty tables, five new indexes, one nullable column. Only idx_crdt_snapshots_created touches a pre-existing table — it indexes crdt_snapshots(user_id, vault_id, created_at) so pack selection can scan snapshots above the watermark. A server running the previous code against the migrated schema behaves exactly as it did.
Telemetry event schema first, then the desktop build
The sync-server carrying the new telemetry event schema must be deployed before any desktop build that emits sync_bootstrap.
The reason is batch-level, not event-level: telemetry events are validated as a batch, and one unrecognised event name rejects the entire batch with a 400. A desktop build ahead of the server therefore does not just lose its bootstrap metrics — it loses every other event that happened to travel with them. Deploying the server first costs nothing, since an event name nothing emits is inert.
The sync_bootstrap event itself is consent-gated through the ordinary telemetry client, fires at most once per action per bootstrap (backed by a per-minute floor so even a begin/complete loop cannot spam), and carries three actions — interactive, full_text, throughput. It ships coarse buckets only (note_bucket, size_bucket) and never content, titles, paths, ids or keys.
Realtime Socket Auth
The change-notification WebSocket (/sync/ws) authenticates once at handshake with a Bearer access token. The server pins that token's expiry to the connection and sweeps every 60s, closing any socket whose token has expired (WS_TOKEN_EXPIRED, close code 4003).
Because access tokens are short-lived, the client renews the connection in place rather than riding each token to expiry. Whenever the token manager refreshes — the same cycle that serves HTTP requests, and always well before expiry — the client sends the fresh token over the open socket:
| Direction | Message | Meaning |
|---|---|---|
| client → server | { type: 'auth', payload: { token } } | Renew this connection with a fresh token |
| server → client | { type: 'auth_ok', payload: { exp } } | Accepted; the connection now expires at exp |
The server verifies the token and requires it to belong to the same device before extending the expiry. Renewal is best-effort: a rejected or unanswered auth leaves the original expiry in place, so the socket closes at expiry and the client reconnects with a fresh token as it otherwise would.
The renewal hook belongs to the running sync runtime, not to the token manager: the runtime installs it at start and detaches it at stop. A refresh that lands after a vault switch or sign-out therefore renews nothing instead of reaching into a torn-down socket and CRDT queue, and the runtime that replaces it installs its own hook.
The message contract, and the mobile client
The socket's message names, the keepalive string, the close codes and a parser live in packages/contracts/src/sync-socket.ts. Desktop parses every frame with its parseSyncSocketFrame; mobile parses against the same module. An unrecognised type parses successfully and is then ignored rather than failing the frame, so a server that starts sending a new message cannot break a client that shipped before it. Desktop drops an ignored frame with a debug log; only a frame that is not a { type, payload? } envelope at all raises the socket's error event.
The parser narrows calendar_changes_available (sourceId), linking_request (sessionId, newDeviceName, newDevicePlatform) and linking_approved (sessionId) alongside the older types. Unknown payload keys are stripped, and a frame missing a required field is ignored rather than forwarded, so a malformed linking frame no longer reaches the renderer.
Mobile is a second implementation rather than a port, because React Native's WebSocket is not the same object as ws. It has no terminate(), no ping/pong events and no unexpected-response, so a rejected handshake surfaces as a bare error and a synthetic 1006 close with the HTTP status nowhere in reach. The mobile client therefore probes the same URL over plain HTTP after a connect that never opened, and reads the real status and error code from there. Headers are the only auth channel; RN's third constructor argument carries them, and X-App-Version goes on the wire without its +build suffix, because the server's version comparison parses 2+318 as NaN and would pass the gate by accident. X-Memry-Vault-Id is equally required in practice: the Durable Object filters every broadcast by the socket's attached vault, so a socket without it connects and then hears nothing.
Mobile does not pin certificates (see below) and relies on the OS trust store. It connects on the foreground and online edges and closes the socket deliberately when the app backgrounds, so a close the OS delivers while suspending the process cannot arm the reconnect backoff and spend the handshake budget, which is 15 per 60 seconds keyed by user and shared across all their devices.
Certificate pinning on the socket
A wss:// socket connects through a pinned https.Agent. One agent is shared for the process rather than rebuilt per reconnect: the agent holds nothing between connects (keepAlive is off, and the WebSocket upgrade detaches its socket), so a fresh one per reconnect only produced garbage.
Sharing does not freeze the pin. The check lives in the agent's checkServerIdentity, which resolves the connecting hostname's pins from certificate-pins.ts on every TLS handshake — an updated pin table applies to the next handshake with no restart and no cache flush. Only the two decisions made when the agent is constructed are cached with it: whether pinning is disabled (unpackaged dev/test builds) and whether the configured host still carries placeholder pins. Both form the cache key, so a change in either destroys the cached agent and builds a new one.
A pin mismatch is terminal for the session: the manager latches certificate_pin_failed, stops reconnecting, and requires an app restart.
Vault-Key Verification
Before syncing — and whenever an entire pull page fails to decrypt — the client verifies its local master key against the account's key verifier (local cache first, GET /auth/key-verifier as fallback). A confirmed mismatch stops the pull cycle without quarantining items or marking them corrupt, escalates once into the recovery flow, and signs the install out so sign-in + recovery phrase can restore the correct key. See Vault-Key Mismatch Detection.
Error Modes
| Failure | Behavior |
|---|---|
| Offline | Outbox queues; retry with backoff |
| Server unreachable | Machine still has a link, so requests are retried with exponential backoff, not instantly |
| Auth expired (401) | Refresh the access token and retry the request once; only a failed refresh prompts sign-in |
| Refresh rejected | Stop refreshing entirely (see below); prompt the user to sign in again |
| Payment required | Sync stays local-only until a paid plan is active |
| Client below floor (426) | Read-only mode; outbox parked; resume after update |
| Platform writes off (403) | Read-only mode; outbox parked; resume when the switch is flipped back |
| Quota exceeded | Surfaces in Settings → Vault |
| Socket token expiry | In-place renewal over the open socket; a rejected renewal falls back to close + reconnect |
| Server unavailable | Exponential backoff; status indicator turns yellow |
| Blob hash mismatch | Reject the item; log; alert health view |
| Vault-key mismatch | Stop pulling without branding items; prompt recovery; sign out to restore the correct key |
| Bootstrap unavailable (404/501/409/429) | No elevated window; steady-state pacing, i.e. pre-#1837 behavior |
| Bootstrap renewal refused | Close locally; pacing reverts on the next chunk |
| Presign unavailable (501) | Transfers stay on the proxied blob paths |
| Pack unusable at any stage | That pack or entry falls back to its item-granular GET; cursor untouched |
Rejected Refresh Tokens
A 401 on /auth/refresh means the refresh token itself is dead, so no retry can succeed. Because every part of the app asks for a valid access token on demand — sync passes, websocket reconnects, CRDT pushes, attachment transfers, calendar sync, billing checks — an unlatched failure would let each of them re-enter the refresh path forever.
A rejection therefore latches. The first two rejections open a backoff window (1 minute, then 5) during which no refresh request reaches the network at all; that spacing exists only so a transient server-side 401 can recover before the session is written off. The third rejection is terminal: the client stops refreshing for good and prompts the user to sign in again.
Signing out remains an explicit user action. The session is already dead on the server, but local key material is never cleared on the strength of an HTTP status alone. Signing in again clears the latch and sync resumes.
Encryption Stays End-to-End
The server never sees plaintext. See Cryptography for the key hierarchy.
Crypto worker and main-thread fallback
Push encryption and pull decryption run in a worker thread so a large batch does not block the main process. The worker is an optimisation, never a dependency: whenever it is unavailable the same batch is encrypted or decrypted on the main thread instead, and sync continues at reduced speed rather than failing.
"Unavailable" covers both a worker that never started and a running worker that rejects a request — a request timeout, the worker crashing or exiting mid-batch, or a message kind the worker build does not implement, which is what a partially updated install looks like. The batch that was in flight when any of those happen degrades to the main thread with the rest; it is not lost.
Degrading cannot mask a bad payload. The worker reports per-item crypto outcomes in its reply — a failed decrypt or a signature mismatch comes back as a per-item failure, not as a rejected batch — so a rejection only ever means the worker itself was unreachable. The main-thread path then runs the identical encryption and signature verification over the same inputs, so an item that genuinely fails crypto still fails; it just fails on the main thread. Push payloads are resolved once and shared by both paths, so the fallback encrypts exactly what the worker was handed.
A worker that crashes takes itself out of the rotation, because the exit leaves nothing to send to. A worker that is alive but silent does not: it looks healthy, so every batch would ask it again and wait out the 60-second request timeout before degrading. The bridge therefore counts consecutive failed requests and stops offering itself after three, from which point batches go straight to the main thread with no round trip. The penalty for a silent worker is bounded at three timeouts for the whole session rather than one per batch.
Three is deliberate on both sides. One failure is noise — a single timeout under load should not cost the session its worker — so a successful batch resets the count and only consecutive failures latch. Waiting longer is expensive, because the penalty is paid in whole minutes. The thread is left alive rather than terminated, and restarting the sync runtime gives the bridge a fresh worker and a clean count.
In-flight requests are bounded too. Each one carries the batch's items and key material until it is answered, so the bridge refuses to hold more than 1,000 at once; past that point requests are rejected at the door and fall to the main thread, which is the same degradation a wedged worker already triggers. A sweep runs while requests are outstanding and collects any request that outlived its own timeout, and it stops itself as soon as nothing is pending, so an idle bridge costs nothing.
A healthy worker that has had no request for two minutes is stopped to give back its V8 isolate and libsodium heap. The bridge still reports itself running while the thread is down for being idle, so crypto keeps routing to it, and the next request spawns a fresh thread; concurrent requests share that spawn. A request that arrives after the idle stop began but before the old thread exited waits for the fresh thread instead of being posted to the exiting one, so it cannot count toward the failure latch. An explicit stop stays stopped, a failed respawn falls back to main-thread crypto, and a latched-off worker is never idle-stopped. The first batch after an idle stop pays the thread and libsodium startup, which runs off the UI path.
Shutting the bridge down asks the worker to exit and waits three seconds before terminating it. A worker that misses that window is fully detached first, so the exit that terminating eventually produces cannot land on a bridge that has already been restarted and cancel the new worker's in-flight batches.
The bridge only listens to the thread it is currently routing to. A message, error, or exit that arrives from a thread it has already walked away from is ignored, and that thread is disconnected outright. Without this, the late exit of a terminated worker would take the live worker out of the rotation — a bridge reporting no worker at all while a healthy thread sat idle, leaving every batch for the rest of the session on the main thread.
Starting and stopping the bridge are serialised against each other. A start requested while a shutdown is still running waits for that shutdown to finish and then spawns a fresh thread, instead of mistaking the thread on its way out for a running one and returning with nothing behind it. For the same reason a thread that has been asked to exit no longer counts as running, so batches raised during the shutdown window go straight to the main thread rather than waiting out a request timeout against a worker that is leaving. This is what keeps a vault switch or a sync restart from ending up on main-thread crypto for the rest of the session.