QxCloudSync Protocol

QxCloudSync Protocol

Device-to-device sync over a blind relay: full-mesh pairing with no transitive trust, hybrid post-quantum handshake, epoch keys, self-healing routes, AES-256-GCM envelopes, and deepMerge rules.

QxCloudSync Protocol

#

QxCloudSync synchronizes rooms, messages, parameters, and friend-room keys between two sessions of the same user_id in deepMerge mode: every object carries { updatedAt, by }, rooms and messages carry ids, and merges are union plus last-writer-wins.

Blind, amnesic relay

The server stores nothing, reads nothing, and logs nothing. It checks the session, the rate limit, and the opaque byte cap, then fans the ciphertext out to sibling sessions. All authentication and confidentiality are enforced client-side only.

Transport

Client to server (opcode 60):

{ "op": 60, "d": { "toClientId": "<48 chars | empty = broadcast>",
  "encrypted": { "...opaque..." }, "requestId?": "<≤128 chars>" } }

Server to siblings (opcode 61, forwarded as-is):

{ "op": 61, "d": { "fromClientId": "...",
  "toClientId": "...", "encrypted": { "...opaque..." } } }

Server to sender, only when requestId is present (relay ack with delivery report):

{ "op": 60, "d": { "ok": true, "delivered": 2, "dropped": 0,
  "peers": ["<clientId>", ...], "peerCount": 3, "requestId": "..." } }

delivered counts the sibling queues that accepted the frame, dropped those whose queue was full (512) or gone — never silent anymore. peers lists the clientIds that received it, peerCount the total sibling sessions for this user_id. The op 61 wire format is unchanged, so older clients keep decrypting and simply ignore the extra ack fields.

Presence directory, client to server (opcode 62) — same-user_id siblings only, sorted by clientId:

{ "op": 62, "d": { "requestId?" } } →
{ "op": 62, "d": { "ok": true, "self": "<own clientId>",
  "peers": [{ "clientId": "...", "platform": "web|mobile|desktop" }] } }

Presence events, server to siblings (opcode 63, best-effort, ignored by older clients):

{ "op": 63, "d": { "event": "join|update|leave",
  "clientId": "...", "platform": "..." } }

The handler is relay_cloud_sync_op (websocket/protocol.rs), modelled on relay_room_signal (op 55 → 56):

  • Rate limit cloud_sync:session:<sid> 120 per 10 seconds (raised from 30: a full mesh needs N(N−1)/2 handshakes plus chunked snapshot floods on join, which choked meshes of 3–4+ devices mid-handshake so they never converged). Directory sync_peers:session:<sid> 15 per 10 seconds. The global WebSocket guard (1200 per 60 seconds) still caps sustained abuse.
  • d.encrypted must be a JSON object with a serialized length of at most 64 KiB (MAX_CLOUD_SYNC_BYTES), else "Payload too large".
  • The sender must be identified (user_id plus username non-empty; session revalidation is already enforced by dispatch) and carry a non-empty clientId, else "Missing clientId". Without a return address siblings could only answer by broadcast, turning every handshake into a mesh-wide storm at 3+ devices.
  • Fan-out targets players with the same user_id, a different session_id, and an empty toClientId or a matching client_id. Recipients are collected under the read lock, then try_send after the lock is dropped; full or closed queues increment dropped instead of vanishing silently (same 512-message queue as op 55/111).
  • No database write, no room_messages push, no dead-drop, no history. Invisible devices are included, unlike room broadcasts.
  • Stale routing is now visible: a toClientId that no longer maps to a connected session yields delivered: 0 while peerCount shows the live mesh — the sender falls back to broadcast and refreshes via op 62 immediately instead of stalling. A broadcast with peerCount: 0 is the normal single-device case and needs no retry.

Mesh healing contract (3+ devices)

  • Unicast frame with delivered == 0 means a stale route: resend once with empty toClientId, then re-learn the mesh with op 62. Do not spin — dropped > 0 means congested queues, so retry that leg with backoff.
  • On op 63 join from an unknown clientId, open a handshake to it; on leave, mark that leg's peerWs stale at once instead of pushing into the void; on update, refresh the Settings → Sync icon.
  • On boot (or after reconnect), fetch op 62: if the mesh is non-empty and no session exists for a listed peer, handshake it directly — do not wait for its hello. This keeps 4–8-device meshes fully interconnected without the manual Pair button.

Trust root and key schedule (client-only)

The masterSecret comes from the 12 recovery words (PBKDF2-SHA256 100k, salt qxphantom:master, then HKDF qxp-master) and is never transmitted.

KeyDerivationRole
syncRootHKDF(master, "", "qxcloudsync:root:v1")Sync trust root, RAM-only.
syncAuthHKDF(syncRoot, "", "qxcloudsync:auth:v1")HMAC key authenticating hellos.
syncMasterHKDF(ecdh || ss1 || ss2 || syncRoot, transcript, "qxcloudsync:master:v1")Session master, RAM-only. Ephemeral private keys are wiped after use.
syncEpochKey_kHKDF(syncMaster, BE64(k), "qxcloudsync:epoch:v1")Data key, 7-day TTL, auto-rotation at TTL − 10% via a signed rekey hello (no new ECDH). Old epoch keys are dropped.
wrapKeyHKDF(epochKey, "", "qxcloudsync:roomkey-wrap:v1")Wraps roomKey with AES-GCM; the room key never travels raw.

Handshake, all inside d.encrypted and signed:

  • hello (ephemeral P-256 plus ephemeral ML-KEM-768 public key) → accept (plus ML-KEM ciphertext to the initiator) → confirm (plus ML-KEM ciphertext to the responder).
  • Each hello carries auth = HMAC(syncAuth, canonical) plus a hybrid signature: device ECDSA P-256 and SLH-DSA-SHA2-128f (FIPS 205, 17088-byte signatures, crypto/slhdsa.ts). Both must verify (fail closed); a wrong HMAC or a missing/invalid PQ signature is silently dropped. Each device holds a long-term SLH-DSA identity keypair (qxcloudsync-device-v1).
  • Post-quantum posture: KEX = ECDH + 2× ML-KEM-768 (FIPS 203), safe via the ML-KEM component; identity = ECDSA + SLH-DSA (FIPS 205); session data = AES-256-GCM (PQ-safe symmetric, Grover halves 256-bit to a comfortable 128-bit). PQ signatures are deliberately handshake-only: a 17 KiB signature on every data part would eat the 64 KiB relay budget and CPU per push, while AES-GCM already gives PQ authenticity per part.

Sessions are additionally persisted client-side, AES-GCM encrypted under HKDF(syncRoot, "", "qxcloudsync:persist:v1"), so a browser restart does not force a re-pair (the server still stores nothing). The topology is a full mesh of pairwise sessions with no transitive trust and no forwarding: a device paired with two others does not make those two know each other, and it never relays between them — every pair needs its own direct handshake (N devices = N(N−1)/2 handshakes). With sync enabled and the 12 words present, a client holding zero sessions broadcasts a signed hello on boot (jittered 2–6 s); any holder of the same words verifies the HMAC and answers. A client that already holds sessions also handshakes proactively: any peer learned via the op 62 directory or an op 63 join event without a live session gets a direct hello — the manual Pair button and re-handshakes (see below) remain as fallback. Duplicate hellos are ignored while the existing session shows recent verified inbound traffic, but answered when that session is stale (reinstall / divergence recovery) — the completed handshake then replaces the old session and its secrets are wiped. Hello is broadcast, accept/confirm/data are unicast via toClientId.

Hellos carry a signed normalized platform (mobile | web | desktop) shown in Settings → Sync with one icon per type. Data envelopes are AES-256-GCM under the epoch key (AAD syncId:epoch:n:from:to) plus a device ECDSA signature, with anti-replay on (syncId, epoch, n).

Mesh self-healing (client-only)

The peer routing address (peerWs, the peer's WebSocket client_id used for unicast) goes stale on every reconnect, tab reload, or phone wake. The relay now reports this instantly (delivered: 0 on the op 60 ack, plus op 63 leave/join events and the op 62 directory — see Mesh healing contract above); the mechanisms below remain as the safety net for older servers and for frames sent without requestId. A rekey lost in a stale route still splits epochs (both sides then drop everything) — hence the epoch healing below.

  • Each session tracks lastInboundAt (last verified inbound traffic only — the displayed lastSeen is also bumped on send, so it can never detect a one-way black hole).
  • Past 3 minutes without inbound, pushes, rekeys and revokes fall back to broadcast (empty toClientId, fanned out to all same-user_id sessions). Reception refreshes peerWs via fromClientId and the ack heals the way back — one round trip restores the leg, no user action.
  • Epoch healing on receive, only for envelopes addressed to us: a missed rekey is absorbed by a bounded local fast-forward (≤ 10 epochs; the subsequent envelope open authenticates, so only the master holder can trigger adoption); a behind peer gets the current signed rekey re-sent; an unknown syncId triggers a fresh handshake. Re-handshakes and rekey resends are rate-limited to one per minute per peer, and the epoch is never advanced without a successful rekey send.

deepMerge rules

  • Version vectors { deviceId: counter } select deltas; merge is union plus last-writer-wins (tie-break: lexicographically greater by wins). Ratchets merge by max.
  • Trusted sender keys merge by union; a divergent JWK keeps the local one and raises an error (no silent overwrite).
  • Room-key conflict policy: refuse plus manual choice — the local key is kept, the conflict is surfaced in Settings → Sync, and the user picks Keep local / Request remote re-push.
  • Large snapshots are split into valid sub-snapshots (rooms/params chunk plus message batches of ~100 per room) so each relay frame stays ≤ 64 KiB. Full history lives in client IndexedDB (qxcloudsync-v1), beyond the 500-per-room localStorage cap. Sync pauses under client-lock, RAM-only OPSEC, or decoy.
  • Room deletion (op 57/58) travels as a deleted tombstone collection (30-day TTL): the deleted room is dropped locally (lists, messages, keys, pins, IndexedDB) and never re-imported from stale snapshots while the tombstone lives. Pins (pinnedRooms, ≤ 5) sync last-writer-wins; room members are attached for newly imported rooms only, so live rosters are never clobbered.

Event propagation

Snapshots are full-state and idempotent, but they are not only periodic: persist() itself notifies subscribers (internal mutation calls included), coalesced into one push 2.5 s after the last change. Out-of-band stores (custom theme, locale) are observed with synchronous watchers. Applying a remote snapshot never re-notifies (internal guard), so there is no echo loop. The 90 s timer remains as a safety net.

Security properties (audited)

  • Quantum-proof (hybrid): KEX = ECDH P-256 + 2× ML-KEM-768 (FIPS 203), safe via the ML-KEM component; identity = ECDSA P-256 + SLH-DSA-SHA2-128f (FIPS 205), both required; session data = AES-256-GCM (≈128-bit PQ margin).
  • Transcript integrity: the KDF transcript is hashed over canonical JSON (sorted keys), never JSON.stringify — the server re-serializes in sorted order, so a naive hash would diverge per side.
  • Anti-replay: per-peer (syncId, epoch) window with monotonic high-water (tolerance 5000 for reordering) plus a bounded dedup set (6000 entries, pruned). Survives restarts via the persisted sendN; last-writer-wins merge makes residual replays harmless.
  • At rest: without client lock, session blobs are AES-GCM under a syncRoot-derived key (device boundary, same as the stored recovery words). With client lock active, sessions and the SLH-DSA identity rest only under AES-GCM envelopes keyed by the lock key; locking wipes all RAM secrets (masters, epoch keys, syncRoot, SLH cache) and never downgrades envelopes.
  • Residual (accepted): hello replay is a bounded nuisance (server rate 120/10 s, 5 recent pendings max, 120 s sweep); the server observes timing, frame counts and approximate sizes (no padding to fixed buckets, unlike PHANTOM); PBKDF2-100k for the master follows the pre-existing PHANTOM parameters.