← All posts

CRDT Lab and the hard part of offline collaborative editing

CRDT Lab turns offline collaborative editing into a small, inspectable systems problem. It shows what convergence means on three replicas, how to verify it, and where replicated text models often fail.

Distributed editing fails in quiet ways. A shared document looks simple until you add latency, dropped packets, and local edits made without a network. Then the hard part appears. Each replica must accept writes on its own timeline and still converge on the same state later.

Practitioners run into this problem in editors, note apps, whiteboards, and any workflow where people expect local-first behavior. The challenge is not syncing bytes. It is preserving intent under reordering, duplication, and delay. CRDT Lab puts this problem in a small, inspectable form. You edit offline on three replicas, reconnect them, and check whether they still agree.

This matters because many systems look correct in a happy-path demo. They break when two users touch the same region, when one device reconnects late, or when an operation arrives twice. A compact lab is useful because you can force those cases on demand. You do not need a production stack to learn where convergence logic holds and where it leaks.

What problem this demo isolates

A replicated editor has two jobs. It must keep working while disconnected. It must merge later without central arbitration of every keystroke. Those jobs pull against each other.

If you assign one server as the source of truth, offline edits become awkward. If you accept local writes first, you need a merge model with strong convergence properties. CRDTs exist for this setting. The core promise is narrow and technical. Replicas that apply the same set of operations, even in different orders, should converge on the same result.

That promise is easy to repeat and hard to validate. You need to inspect the edge cases.

  • Two replicas insert at nearby positions while disconnected.
  • One replica deletes content another replica still references.
  • Messages arrive out of order.
  • The same operation is delivered more than once.
  • A replica lags far behind, then catches up.

A useful lab does not hide these cases behind polished collaboration UI. It exposes them. Three replicas matter here because pairwise agreement is not enough. Many merge bugs stay hidden with two participants and appear only when a third path of causality enters the system.

What to inspect in a three-replica CRDT setup

When you use the demo, treat it as a state machine, not a text box. Make edits on one replica while the others stay disconnected. Then create conflicting structure elsewhere. Reconnect in different orders. Your goal is to observe whether the merge rules depend on topology or timing.

Start with simple checks.

  • Type distinct text on each replica while all three are offline.
  • Reconnect one pair first, then the third.
  • Reset and reconnect in a different sequence.
  • Compare final state across all three replicas.

If a CRDT is doing its job, the order of delivery should not change the eventual document. Intermediate views can differ. End state should converge.

Then inspect positional behavior. Text data structures often fail around indexing assumptions. Local cursor positions feel linear to users. Distributed inserts are not. Under the hood, each element usually needs an identity with ordering rules that survive concurrent inserts. You want to see whether nearby concurrent edits settle into a deterministic order.

Next, inspect delete behavior. Deletes are where many readers underestimate complexity. A delete often targets an identity, not a raw position, because positions drift as inserts land. If one replica removes content while another edits around it, the system needs a stable way to interpret both operations later.

The practical signal is simple. After enough reconnects, all three replicas should render the same content. If they do not, you likely found a flaw in causal tracking, operation identity, deduplication, or ordering.

How you would verify convergence claims

A good verification approach is adversarial and repeatable. Do not trust one successful merge. Try to falsify convergence.

Use a small matrix of scenarios.

  1. Concurrent inserts. Put each replica offline. Insert different characters at the same logical position. Reconnect in multiple sequences.
  2. Insert versus delete. Delete a character on one replica while another inserts beside it. Reconnect and compare.
  3. Duplicate delivery. If the lab exposes message replay or repeated sync, check whether applying the same operation twice changes state.
  4. Long partition. Let one replica diverge with several edits while two others sync among themselves. Reconnect the isolated one last.
  5. Mixed fan-in. Sync A with B, then B with C, then A with C. Reset and change the order.

You are looking for two properties.

  • Convergence. Same final state after all operations propagate.
  • Idempotence and commutativity at the operation set level. Replays and reorderings should not create drift.

You do not need internal source access to learn a lot. Visible behavior exposes core guarantees. If the final text differs across replicas after all have synced, the model failed its main contract. If the final text matches but only for one sync order, you have a timing-sensitive implementation issue.

For a stronger check, use recognizable tokens instead of free typing. Insert markers such as A1, B1, C1 from different replicas. That makes it easier to track whether one operation vanished, duplicated, or moved unexpectedly after merge.

Where systems like this commonly go wrong

The most common mistake is confusing last-write-wins conflict handling with collaborative text merging. Last-write-wins is easy to explain and poor at preserving concurrent user intent in rich text. It collapses races by timestamp or version and drops structure.

Another failure point is identity design. In distributed text, elements need stable identifiers. If the implementation leans too hard on array indexes, concurrent inserts and deletes become ambiguous after remote operations shift positions. Bugs then appear as lost edits, duplicate characters, or different final orders on different peers.

Causality tracking is another weak spot. A replica needs enough metadata to know what it has seen and what another peer is missing. If the sync logic skips this, late-arriving operations can be misapplied or ignored. In three-replica topologies, this shows up when A and B agree, B and C agree, but A and C still diverge after relaying through B.

Deduplication also matters. Real networks replay. Retries happen. Peers reconnect after partial sync. An operation-based design needs a durable notion of operation identity so repeated delivery does not mutate state twice.

Tombstones and garbage collection are another tradeoff. Some CRDTs retain deleted-item metadata to preserve ordering and causality. Drop it too early and a lagging replica cannot merge correctly. Keep it forever and state growth becomes a concern. A lab like this does not need to solve storage policy in full, but it helps you see why deletion in distributed text is more than removing a character from an array.

What signal to read from the results

The key signal is not whether merge output matches your intuition in every case. Distributed order among concurrent edits is often deterministic but arbitrary. The stronger signal is whether all replicas land on the same result from the same operation set.

That distinction matters in product work. Users care about predictability and preservation of edits. Engineers need stronger invariants underneath.

Read the results in layers.

  • Agreement across replicas. The first invariant.
  • Stability across sync order. A sign the implementation is not topology-sensitive.
  • No duplicated or missing edits. A sign operation IDs and deduplication hold.
  • Reasonable placement of concurrent inserts. A sign ordering rules are consistent.

If a system passes agreement but yields surprising order, you have a product question about presentation and intent resolution. If it fails agreement, you have a correctness problem.

Three replicas are useful because they expose transitive sync bugs. A pair of peers often masks issues. Add a third, and hidden assumptions about who saw what start to break. That is why this demo is more than a toy. It compresses a hard distributed-systems problem into something you can probe by hand.

What to watch next

If you explore this lab, watch how it behaves under repeated reconnects and asymmetric histories. Those are the conditions where convergence models earn their keep. After agreement, the next engineering questions are storage growth, cursor behavior, richer document structures, and whether the same merge rules still hold outside plain text.

A small demo will not answer every production concern. It does show the core requirement with unusual clarity. Offline editing is easy to promise. Agreement after independent work is the part worth testing.