AniUI Academy

Worked Scenario: Design a Browser Code Editor (Monaco / VS Code Web)

A full worked answer to the Microsoft-style browser-editor prompt — tokenization that keeps up with typing, large-file performance, extensions without a plugin exploit surface, and the honest limits of a web runtime.

16 min read

A browser code editor — Monaco (the engine behind VS Code web, GitHub code view, StackBlitz, CodeSandbox) — is one of the Microsoft-style prompts that sorts out whether a candidate has thought about performance under input rather than just performance in general. Typing is sacred: every feature in an editor is in a budget race with the main thread's render loop for the next frame.

Clarifying requirements first

Before proposing anything, the questions worth asking out loud: Largest single file the editor must handle — tens of thousands of lines, hundreds? Languages supported — the top five, twenty, extensible? Extensions required, and if so, by whom (first-party only, or user-installable)? LSP (language-server) support — in-process or over a wire? Collaborative editing in-scope, or single-user? For this answer, assume: files up to ~100 000 lines in the worst case; ~20 languages supported first-party plus a user- extension model for the rest; a wire-protocol LSP (the server runs somewhere, the editor talks to it); single-user editing (collaborative is a separate design — see the wiki scenario).

Rendering strategy: pure CSR, mounted after shell

The editor is an application, not a document. SSR is irrelevant here — nothing is crawlable, every byte of the editor is interactive. The host page can SSR a shell (file tree, header), and the editor itself mounts client-side after that shell paints. Code-splitting the editor is mandatory: no page should pay for Monaco's megabytes until a file is being opened.

The document model: lines, pieces, and no React

Lines

The file as an array of lines indexed by number — the primary addressable unit for most editor operations.

Piece table / rope

Edits as insertions and deletions over the original buffer — O(1) local edits even on huge files, O(n) only when reading back the whole document.

Decorations

Overlays on the document — syntax tokens, error squiggles, diagnostics, selections — stored separately from the text so they can be re-applied without reflowing the buffer.

The surface that most candidates get wrong: there is no React component tree per line in a serious editor. Monaco renders characters to a canvas-like layer or a tightly-managed set of absolutely-positioned divs — the viewport holds dozens of rendered rows, not every line in the file. React re-rendering on every keystroke across a 50 000-line file is a non-starter; the editor owns its render loop and only mutates what changed.

Tokenization: main thread for the viewport, Worker for the rest

A 50 000-line file tokenized synchronously on open freezes the editor for over a second. The right split:

  1. Step 1

  2. Step 2

  3. Step 3

  4. Step 4

The user sees a readable, highlighted viewport the moment the file opens. Scrolling into un-tokenized regions shows plain text for a few hundred milliseconds until the Worker catches up — honest and acceptable, where a one-second freeze on open is not.

Prioritising per-keystroke work by latency budget

Several features all react to every keystroke: syntax tokenization, bracket matching, autocomplete, error diagnostics, formatting hints. Running them all synchronously on every tap collapses input rate on a slow machine. The honest model is to tier them by budget:

Immediate (sub-16ms)

Viewport-visible syntax tokens, cursor position, bracket match. Must land in the same frame as the input.

Short debounce (~50ms)

Autocomplete: feels responsive without firing on every single character of a fast burst.

Long debounce + preemptible (~300ms, in Worker)

Document-wide validation, formatting hints, slow LSP requests. Cancelled by the next keystroke.

Blanket-debouncing everything is the common failure — it makes autocomplete feel sluggish and the highlight visibly lag the cursor. Different work has different budgets.

LSP over a wire: honest about round-trip costs

If the language server runs server-side, every autocomplete, hover, and go-to-definition is a network round-trip. Design the UI around that cost rather than pretending it's local:

  • Autocomplete shows an immediate heuristic result (local tokens, previously- seen identifiers) while the LSP request is in flight; the real result replaces it when it lands.
  • Hover info is a two-stage render: an instant local tooltip ("variable", from the token type) swapped for the LSP's richer tooltip when it returns.
  • Cancel in-flight requests on the next edit; a reply arriving two seconds late to a cursor that has moved is noise.

Extensions: sandboxed, with a narrow structured API

In-page extensions are a security story, not a performance one. A buggy or malicious extension loaded as a <script> has access to every cookie, every auth token, every API the host hits, and the whole DOM. The right architectural response is a hard sandbox: each extension runs in its own Web Worker (or a cross-origin iframe if DOM access is explicitly needed), and the host exposes a narrow structured API over postMessage:

  • getDocument() returns a snapshot
  • insertAtRange(start, end, text) applies an edit
  • registerCommand(id, label) registers a command the host can invoke
  • showNotification(message, severity) surfaces to the user

VS Code desktop runs extensions in a separate process for exactly this reason; the browser equivalent is a Worker or iframe boundary. The sandbox is what makes the trust model honest — not a performance tweak.

Large files: the honest limits of a browser

A 100 000-line file is handleable; a 10 000 000-line file is not, no matter how clever the editor. Browser memory is finite, text layout at that scale is expensive even in Monaco, and the honest answer to "can we open a 2GB log in the editor?" is "no, you want a dedicated log viewer" (see the CI scenario). Shipping an upper bound (e.g., refuse to open files above a documented size, with a clear error message) is better than crashing the tab or appearing to load forever.

What's explicitly out of scope, and why

Not solved in this answer: collaborative editing (its own prompt — see the wiki case study for the CRDT/OT reasoning); language-server implementations themselves (a backend concern); the test-runner integration (its own feature with its own data model); vim/emacs keybinding emulation (a feature set, not a system-design question). Naming these as deliberately out of scope is a strength.

What to remember

  • The editor owns its render loop. There is no React component per line in a serious editor at scale; the viewport holds dozens of rendered rows, mutated directly on edit.
  • Tokenization is a split job: main thread for the viewport on open and per-keystroke; Worker for the rest of the file, preempted by typing.
  • Per-keystroke work is tiered by latency budget — the viewport's highlight is immediate, autocomplete short-debounced, document validation long-debounced and preemptible — not blanket-debounced.
  • LSP over a wire is designed around round-trip latency: heuristic local results immediately, real results when they arrive, in-flight requests cancelled on the next edit.
  • Extensions run in a sandbox (Worker or iframe) with a narrow structured API. In-page extensions are a supply-chain compromise waiting to happen, not a feature.
  • Browsers have finite memory. An upper file-size limit, visibly surfaced, is better than appearing to load forever.

Check yourself

3 questions · pass 3/3 to unlock Worked Scenario: Design a Chat With Threads (Teams / Slack)

up to 50
  1. 1.A 50 000-line file is opened in the editor. Highlighting the whole file synchronously on open blocks the main thread for 1.5s. What's the right split between the main thread and workers, and why?

  2. 2.The editor supports extensions. A naive design loads each extension as a script in the main page. What's the actual threat here, and what's the right architectural response?

  3. 3.On every keystroke, the editor runs syntax validation, formatting hints, and autocomplete. On a slow machine these cumulatively drop input to 20fps. What's the right prioritisation model?

3 left to answer

Discussion

Sign in to post
Sign in to join the conversation.Sign in

No comments yet. Be the first to say something.