← Teardowns

How Is Miro Built? Real-Time Whiteboard Infra and Its Actual Moat

A deep, honest look at how Miro's real-time whiteboard actually works under the hood — CRDTs, sync infra, and whether its moat is tech or something else entirely.

What Miro actually is, under the marketing

Strip away the templates and icebreaker stickers and Miro is a distributed document editor where the document happens to be an infinite 2D canvas instead of a page of text. That sounds simple until you remember what it needs to do: let 50 people drag shapes, type text, draw freehand lines, paste images, and run a Mural-style brainstorm simultaneously, with sub-100ms perceived latency, on a canvas that can scale to tens of thousands of objects, viewed on everything from a laptop to a touch display in a conference room.

The product is really three systems wearing one UI: a real-time sync engine (who sees what, when), a canvas rendering engine (how thousands of objects stay smooth at 60fps while you zoom and pan), and a collaboration/workflow layer (comments, voting, video huddles, permissions, templates) that turns a whiteboard into a place teams actually keep coming back to. The first two are hard engineering. The third is where the moat actually lives.

The sync engine: why whiteboards are harder than Google Docs

Google Docs solved real-time collaborative text years ago using operational transformation, later moving toward CRDT-flavored approaches. Text is a linear sequence — inserts and deletes at an index. A whiteboard is a scene graph: shapes have position, size, z-order, rotation, style, and nested groups, and two users can simultaneously drag the same object in opposite directions, resize a frame that contains 40 child objects, or undo an action that a third person has already built on top of. Conflict resolution on a 2D canvas with hierarchical objects is a materially harder problem than resolving two cursors inserting characters into the same paragraph.

Miro's approach is built around operation-based synchronization with last-write-wins and field-level merging rather than naive whole-object locking — when two people move the same shape, you don't want the whole object to 'lock,' you want both edits to resolve into something coherent without a jarring snap-back. Under the hood this looks like versioned operations per object property, broadcast over WebSockets through a pub/sub backbone, with the server acting as the arbiter of ordering so every client converges on the same state even if messages arrive out of order or a client reconnects after dropping offline for ten seconds on a flaky conference-room wifi.

This is the part that's genuinely hard to clone in a weekend with an AI coding assistant. You can get a toy real-time canvas working with Yjs or Liveblocks and a canvas library in an afternoon — and for 5 users and 200 shapes it'll feel identical to Miro. But the failure modes only show up at scale: operation storms when 40 people paste large images simultaneously, server-side conflict arbitration under packet loss, selective object sync so a board with 50,000 shapes doesn't ship the whole scene graph to every client on load. That's the gap between a demo and infrastructure that enterprises will run their quarterly planning on.

Rendering an infinite canvas without melting the browser

The second hard problem is purely client-side: rendering. A naive DOM-based whiteboard dies around a few hundred elements. Miro renders to canvas/WebGL with aggressive viewport culling — only objects actually visible (plus a buffer) get drawn or even kept hydrated in memory — combined with level-of-detail tricks so a zoomed-out board shows simplified shapes instead of full vector paths and text layout.

Then there's the object model itself: frames, groups, connectors that need to re-route when you move an endpoint, sticky notes that auto-resize to fit text, tables, mind maps. Each object type has its own hit-testing, selection, and serialization logic, and all of it has to stay perfectly in sync with the collaborative state engine described above. This is the unglamorous 80% of the codebase — not the sexy real-time sync, but thousands of small interaction details (snap-to-grid, multi-select marquee across a rotated viewport, keyboard shortcuts that don't fight browser defaults) that took years of user feedback to get right. An AI tool can scaffold a canvas and a few shape types quickly. It cannot shortcut years of UX refinement encoded in edge-case handling.

Technical cloneability: higher than Miro would like you to think

Be honest about this one: the core real-time whiteboard primitive is not proprietary magic anymore. Libraries like Yjs, Automerge, and Liveblocks have commoditized CRDT-based sync; tldraw has open-sourced a genuinely good infinite-canvas rendering engine; Excalidraw proved a scrappy team can ship a credible whiteboard in months. If you're scoring 'could this be rebuilt by a strong team with modern AI tooling in a few months,' the answer for a Miro-like MVP is yes — basic shapes, sticky notes, real-time cursors, and simple sync are achievable fast.

Where cloneability drops off is at the tail: enterprise-grade conflict resolution at high concurrency, offline/reconnect handling that doesn't corrupt state, accessibility and internationalization across dozens of locales, SOC 2 / HIPAA-grade security controls, and performance on boards with tens of thousands of objects used by Fortune 500 design and engineering orgs. That tail is 90% of the real engineering effort and 10% of what a demo video shows you. So: technically cloneable at the surface, genuinely expensive to clone at production depth.

Where the real moat is: not the canvas, the workflow lock-in

Miro's defensibility isn't the WebSocket layer — it's that it became the shared visual memory of hundreds of thousands of teams. Every retro, every roadmap, every user journey map, every workshop template lives on boards that are cross-linked, commented on, exported into decks, and referenced in Jira tickets via the Miro-Atlassian integration. That's a data gravity moat: the switching cost isn't 'learn a new tool,' it's 'migrate two years of institutional diagramming history and retrain 10,000 employees on new keyboard shortcuts.'

On top of that sits an integration and app-marketplace moat — Miro ties into Jira, Asana, Slack, Microsoft Teams, Confluence, Salesforce, and ships an app framework so vendors build inside Miro rather than around it. Every integration is a small but real piece of distribution and retention that a clone starts at zero on. And then there's the enterprise sales layer: SSO, SCIM provisioning, admin governance, audit logs, procurement relationships with IT departments — boring stuff, but it's exactly the stuff that makes a CTO renew instead of switching to a cheaper, technically-equivalent competitor.

This is the classic oneprompt split: the canvas and sync engine score medium-to-high on cloneability (a capable team with AI tooling can get 70% of the way there fast), but the business moat scores high — not because the tech is unreproducible, but because the value accrued in the data, the integrations, and the enterprise relationships, not in the rendering loop.

What a 'Miro clone' prompt actually gets you

If you ask an AI coding tool to build 'a Miro clone,' you'll get a working infinite canvas with shapes, text, and real-time cursors surprisingly fast — that part of the stack has become a solved, well-documented problem. What you won't get in that first prompt: robust operational-transform-style conflict resolution under real network conditions, a performant renderer at 10,000+ objects, an integration ecosystem, or any of the enterprise trust infrastructure that actually wins contracts.

That's the useful lens for evaluating any 'can this be rebuilt' question: separate the demoable core loop (fast to clone, often open-sourced already) from the compounding moat (data, integrations, workflow habit, enterprise trust — slow to clone, if cloneable at all). Miro's canvas is a commodity today. Miro's position as the whiteboard every cross-functional team already has muscle memory for is not, and that's the part worth actually worrying about if you're a competitor — or betting on if you're an investor.

Want the same teardown for any site?

Analyze a site →