How real-time collaborative editing works

Live cursors look like the hard part and are the easy part. The real problem is what happens when two people change the same thing at once, and there are only a few honest answers to that.

Collaborative editing feels like one feature and is actually three, stacked, in increasing order of difficulty: showing who else is here, sending changes as they happen, and deciding what the document means when two changes conflict. Teams routinely build the first two, discover the third late, and find that it was the whole problem.

Presence is the easy layer

Presence is who is connected, where their cursor is, what they have selected. It looks impressive and it is the simplest part, because presence data is disposable. If you drop a cursor update, the next one arrives forty milliseconds later and nobody notices. There is no history to reconcile and nothing to be permanently wrong about.

The only real design decision is not persisting it. Presence belongs in memory with a short expiry, because a cursor position has no meaning once its owner has gone. The characteristic bug in this layer is ghost users: someone's connection dropped without a clean disconnect and their avatar sits in the header for an hour. Heartbeats with a timeout solve it, and the timeout wants to be short.

Transport is a solved problem with one trap

Getting changes from one client to others is well-served: a websocket connection, or a hosted service that manages them for you. Either is fine, and the choice is mostly about whether you want to operate reconnection logic yourself.

The trap is that the network will interrupt you, and the interruption is not rare. Laptops sleep, tunnels drop, phones move between networks. Every one of those produces a client that was disconnected for a while and now needs to rejoin, and 'what happens to the edits made during that gap' is a question the transport layer does not answer for you.

Every collaborative system is really a system for handling reconnection. The steady state is the easy case.

Conflict is the actual problem

Two people edit the same field within the same second. Both changes reach the server. What is the document now?

There are only a few honest answers, and the industry has converged on a small set.

Last write wins is the simplest: keep whichever change arrived later and discard the other. It is trivial to implement and it silently loses work, which is acceptable for a status field and unacceptable for a paragraph somebody was writing.

Locking avoids conflict rather than resolving it. One person holds a field, everyone else sees it read-only. Easy to reason about, and it turns collaboration into queueing, which users experience as the tool getting in their way.

Operational transformation sends operations rather than states, and transforms incoming operations against ones already applied so intent survives. 'Insert at position 5' becomes 'insert at position 7' if somebody else inserted two characters earlier in the document. It works well and it needs a central server to establish a canonical order, and the transformation functions are famously subtle to get exactly right.

Conflict-free replicated data types take a different route: design the data structure so that concurrent operations commute. Apply the same set of changes in any order on any replica and you land on the same state, no central authority required. This is why they underpin offline-capable and peer-to-peer collaboration.

The tradeoff nobody mentions

Convergence guarantees that everyone ends up with the same document. It does not guarantee the document is what anyone wanted.

Two people concurrently rewriting the same sentence will converge on a merge containing both, deterministically and identically for everyone, and that merge can be nonsense. The mathematics is satisfied. The writer is not. This is why serious editors layer intent-preserving behaviour on top of a converging structure rather than treating convergence as the finish line.

There is also a cost that surprises people: metadata. To know that concurrent operations commute, the structure keeps identity and causality information per element, and that bookkeeping can dwarf the content it describes for a long-lived document. Garbage collecting it without breaking the guarantees is one of the genuinely hard parts, not a footnote.

How to choose

That last option is underrated. A great deal of engineering effort goes into resolving conflicts silently in situations where surfacing the disagreement would have been both easier and more honest.

Automatic merging is a user-experience decision disguised as an algorithm choice.

The order to build it in

Decide the conflict model first, before any transport code, because it constrains everything above it and it is the layer you cannot retrofit. Then build the transport and take reconnection seriously from the beginning. Add presence last: it is the most visible layer and the least structural, and building it first tends to produce a demo that convinces everyone the hard part is done.