Screen sharing sends a picture of your page
I have a problem with screen sharing.
Not a quality problem. A category problem.
When I present from a browser, I have a real web page. Real text, real links, real elements. The screen share takes all of that, flattens it into video, and ships the pixels. The audience gets a picture of my page.
On a laptop in a video call, that is fine. In a room, it is not. The phone in the back row gets a blurry rectangle. Venue wifi turns it into a blurry rectangle that freezes. Someone who joins five minutes late gets whatever frame happens to be on screen, with no way to look back.
The page was already structured. We throw the structure away and then spend bandwidth trying to recover what we threw away.
Why not just send everyone the slides link?
That works if your talk lives inside one deck tool, and some of them already do live following well. It does not work when the thing you are presenting is a GitHub diff, a Grafana dashboard, an Excalidraw board, or the actual app you are demoing. Those are just web pages. There is no link that puts the audience on the same page, in the same state, at the same moment.
So I started building one.
I shelved this idea once#
This is not my first attempt.
Years ago I built a prototype that made one browser follow another. Cursor movement, clicks, typing, all relayed with DOM selectors. It worked in the demo. It fell apart for two reasons.
First, selectors are fragile. div.card:nth-child(3) only means something
against the exact DOM it came from. If the receiving page differs at all, the
click lands on the wrong element or on nothing.
Second, it was a security hole. Replaying my clicks against your browser means my input is driving your logged-in session. That is not a sync feature. That is session hijacking with extra steps.
I shelved it. Both problems stayed in my head for years.
The word that fixed it#
Linguists have a word for “this”, “here” and “that one”: deictic expressions. They only mean something relative to the person saying them. “Click this” is useless unless you know where the speaker is standing.
That was the whole bug. My prototype transplanted “this” from one browser into another. It should have resolved “this” against the receiving side.
Both failures are the same failure. A selector is a “this” that only makes sense in my DOM. A replayed click is a “this” that only makes sense in my session. Stop transplanting, start resolving, and both go away.
So the engine I am building is called deictic, and resolving is its whole
job.
Two modes, picked per region#
Anchor mode is for markup I control. Elements carry a stable data-pt-id.
Instead of “click at (412, 380)”, the source sends intents like “go to this
slide” or “point at this element”. The receiver looks the id up in its own DOM
and shows the equivalent. No coordinates, no selectors. It is cheap on
bandwidth.
Mirror mode is for pages I do not control. DOM mutations are captured as diffs, rrweb-style, and rebuilt on the receiving side as a read-only replica.
Why not just use Mirror mode for everything?
Because it costs more bandwidth, and fuzzy page matching is exactly what anchors let you skip. You use anchors when they exist and fall back when they do not. The fallback chain is anchor, then mirror, then a coarse point-and-highlight. It never fails silently.
Mirror mode is also what kills the old security hole. The audience is looking at a description of my page. There is no logged-in session underneath for my input to drive. The hijack is not blocked, it cannot exist.
The DOM is the state#
This was the decision that shaped everything else.
deictic keeps no session model, no event buffer, no shadow document.
snapshot() reads the live DOM into an object. apply() writes into the live
DOM. That is it.
The payoff is catch-up. Someone joins late? Apply the latest snapshot, then the deltas after it. Someone’s wifi dropped for thirty seconds? The same two steps. Late join and reconnect are the same operation, and there is no history to replay, so nobody watches a fast-forward animation of what they missed. They just land on now.
Recovery is always a resync, never a patch-forward. There is no patch-forward function to call, so nobody can call it.
The server never reads a frame#
A frame is a readable envelope around an opaque payload. The envelope holds
the session id, sequence number, frame type and protocol version, in protobuf.
The payload is bytes that only deictic can encode or decode.
The Go server reads envelopes, assigns sequence numbers, fans out and stores blobs. It never decodes a payload. It does not import a single line of engine code.
Why Go for the server and TypeScript for the engine? You write Rust.
The engine is DOM-bound. It runs MutationObserver callbacks, attribute reads and node insertions on every frame. In Rust with WASM, every one of those would cross the bindgen boundary. The server is websocket fan-out with almost no compute, and goroutines fit that directly. If fan-out tail latency ever becomes the problem, the server is small enough to rewrite.
The protocol version is in every frame from the first commit. A mismatch means reject and reload, never a best-effort decode. A tab left open across a deploy is the normal case, not the edge case.
One test to rule them all#
The whole suite hangs off one property:
For any sequence of source operations, after all frames are applied, the target DOM equals the source DOM.
Operations are randomly generated, not hand-written. When a 200-step sequence fails, the framework shrinks it to the three steps that matter. Then I drop random frames, force a resync, and assert convergence again, because bad networks are the whole premise.
Real pages are the hard part#
Toy markup always works. Real pages are where this gets honest.
Excalidraw’s canvas comes through. Grafana Play works. GitHub diffs work.
Iframes do not. Fonts and styles break intermittently. Nested element scroll does not always apply on the receiving side.
If you have fought iframe capture in MutationObserver land, I genuinely want to hear how you did it.
What I got wrong#
I showed early versions to friends and colleagues. Their feedback was “I don’t see how this is different from screen share”. They were right. Nothing was broken, they just never feel the problem. Watching a call on a laptop, pixels are fine.
The people who feel it are the ones presenting to a room: teachers, conference speakers, anyone doing a live demo on bad wifi. I should have started there.
Where this is going#
I am building this as Present Tense. It is in early beta, rough edges included.
If you present to real rooms and this problem sounds familiar, write to hi@present-tense.io. I would rather hear from ten people who feel it than a hundred who do not.
If you have feedback regarding this blog post, click on an issue on GitHub
That's it in this post. In case I don't see ya, good afternoon, good evening and good night.