Lesson 36 of 36
Worked Scenario: Design an Infinite Canvas (Figma / Miro / Excalidraw)
A full worked answer to a Senior-FE prompt that most candidates underprepare for — a pannable, zoomable canvas with thousands of shapes, hit-testing, selection, and multi-user editing done honestly.
A Figma/Miro/Excalidraw-shape infinite canvas is a prompt that most Senior frontend candidates underprepare for — it's an unusual mix of rendering, input, and real-time concerns, none of which the standard React patterns cover cleanly. The honest answer is pragmatic, not academic.
Clarifying requirements first
Before proposing anything, the questions worth asking out loud: How many shapes per board at the extreme — hundreds, thousands, tens of thousands? Zoom range — 0.1x to 10x, or something more aggressive? Shape complexity — rectangles and arrows, or rich components with text layout? Multi-user required? Mobile support? For this answer, assume: boards of up to ~10 000 shapes, zoom 0.05x to 20x, mixed shapes including text and bitmap images, multi-user real-time required, desktop-first with usable touch panning on mobile.
Rendering primitive: a single canvas, viewport-culled
A board of 10 000 shapes cannot be a DOM tree of 10 000 nodes. The honest primitive is a single canvas (2D for simplicity, WebGL once the shape count or per-shape complexity demands it), and every frame does:
- Step 1
- Step 2
- Step 3
- Step 4
The spatial index is what keeps culling O(log n) rather than O(n). A naive "iterate every shape, check if visible" works up to ~1 000 shapes and degrades predictably beyond.
The transform: one source of truth for pan + zoom
world coords
What each shape stores. Independent of viewport.
transform = translate(pan) · scale(zoom)
A single source-of-truth matrix representing the current view. All screen→world and world→screen math goes through it.
screen coords
What the canvas renders in. Derived from world coords and the transform.
Keeping one transform (rather than independently tracked pan and zoom scattered across components) means hit-testing, drawing, and cursor display all share a single truth. The common bug when the transform is not centralised: the cursor and the drawn shape disagree about where "here" is after a zoom change.
Hit-testing: a separate concern from rendering
A click at (screenX, screenY) needs to resolve to "which shape, if any".
The naive approach is to iterate every shape in reverse Z-order and check
containment — fine up to ~1 000 shapes, slow beyond.
The scalable approach is to reuse the same spatial index that powers culling: query the index for shapes whose bounding box contains the click point (world-coordinates), then do precise hit-tests against that short list. For irregularly-shaped elements (text with its baseline, arrows with their curves), the precise hit-test can be expensive — doing it on a short list of candidates rather than every shape is what makes it tractable.
An alternative on 2D canvas: an off-screen "hit canvas" where each shape is drawn in a unique colour-id; a one-pixel read at the click point tells you the shape id instantly. The cost is drawing everything twice per frame — a trade worth considering at low shape counts, usually not worth it at high ones.
Dragging and multi-user editing
A drag is local-first on the dragger's screen (zero network latency for your own hand) and progressive to peers:
- Step 1
- Step 2
- Step 3
- Step 4
- Step 5
The anti-patterns are familiar from the collaborative-wiki scenario: locking a shape on drag-start kills the live feel; drag-end-only sync makes peers see shapes teleport; unthrottled per-move WebSocket writes flood the channel.
Zoom-out at scale: tiled rasters
At extreme zoom-out, a 10 000-shape board produces a few hundred visible pixels — drawing every shape just to produce the overview is wasted work. At a specific zoom threshold (and below), switch to a different render path: pre-rendered raster tiles of the board at that zoom level, cached and invalidated per-tile on edits that overlap.
Tile cache
Keyed by (zoom level bucket, tile x, tile y). An LRU with sensible memory bounds.
Render path
Above the zoom threshold, render shapes directly from the spatial index. Below, composite tiles.
Invalidation
On edit, mark overlapping tiles dirty. Dirty tiles re-render on next viewport intersection, not eagerly.
This is the map-tile pattern applied to a vector canvas. The honest trade: first-visit zoom-out is a bit slower (tiles need generating) but pan/zoom in that range becomes essentially free.
Memory and off-main-thread work
A big board plus a lot of history (undo stack, cursor streams) can grow memory in unobvious ways. Two disciplines:
- Shapes are plain objects, not React components. Nothing in the rendering path should produce React fibers.
- Expensive work (text layout for a complex label, image decoding, PDF export) runs in a Web Worker. The main thread owns input and canvas drawing; everything else is a worker job.
What's explicitly out of scope, and why
Not solved in this answer: shape rendering engines (text layout, bézier curves) in detail — these are mostly a known-work problem solved by canvas APIs or small libraries; collaborative text editing within a shape (a separate CRDT problem — see the wiki scenario); vector-image file formats (SVG import/export); the server infrastructure behind the real-time channel. Naming these as scoped out is a strength.
What to remember
- One canvas, viewport-culled via a spatial index (quadtree/R-tree). The DOM-per-shape approach collapses at hundreds of shapes.
- A single transform matrix is the source of truth for pan + zoom. Hit-testing, drawing, and the cursor all derive from it.
- Hit-testing reuses the spatial index: find candidates fast, do the precise test on the short list. An off-screen hit canvas is a valid alternative at smaller scale.
- Drag is local-first on the dragger (zero network latency), throttled with interpolation to peers, and reconciled on commit via a CRDT or OT — not locked.
- At extreme zoom-out, switch to tiled rasters. The map-tile pattern applied to a vector canvas.
- Everything expensive goes in a Web Worker. The main thread owns input and canvas drawing.
Check yourself
3 questions · pass 3/3 to finish the course
1.A canvas holds ~5 000 shapes across an effectively infinite coordinate space. Rendering every shape to the DOM every frame is a non-starter. What's the right rendering primitive, and what's the first optimisation on top of it?
2.A user drags a shape across the canvas. Two others are viewing the same board in real time. What's the honest contract for every piece of this interaction?
3.On a zoomed-out view (0.1x), a 10 000-shape board is intelligible even when individual shapes are smaller than a pixel. What's the right rendering optimisation for this zoom range?
3 left to answer
Discussion
Sign in to postNo comments yet. Be the first to say something.