Lesson 29 of 36
Worked Scenario: Design a Code Review Diff Viewer (GitLab / GitHub)
A full worked answer to the GitLab-style diff-viewer prompt — side-by-side vs unified rendering, virtualization across 10 000-line files, syntax highlighting without blocking paint, and inline comments anchored across edits.
A code-review diff viewer — GitLab, GitHub, Phabricator, Reviewable — is the kind of prompt that reveals whether a candidate has actually built one, or only used one. The naive "render each file's diff into the page" approach collapses the second you give it a real MR, and nearly every real-world fix sits in territory this course has already covered: virtualization, Web Workers, content-addressed anchors.
Clarifying requirements first
Before proposing anything, the questions worth asking out loud: Average and worst-case MR size — tens of files, hundreds? Largest single file to render? Side-by-side only, or unified too? Inline comments required (anchored to specific lines)? Syntax highlighting for every supported language, or just the top handful? Mobile first-class or secondary? For this answer, assume: up to a few hundred files per MR with a worst case of ~12 000 lines in a single file; both side-by-side and unified views; inline comments anchored across force-pushes; highlighting for the top ~20 languages; desktop-first.
Rendering strategy: SSR shell, CSR diff
Behind a login, so SSR for the shell (header, file list, MR metadata) with cookies attached gives a visually-complete first paint. The diffs themselves are a CSR concern — they are heavy, interactive (expand context, add comment, mark viewed), and nothing about them benefits from being in the first HTML payload. Streaming SSR for the file list means the shell doesn't wait for the full file manifest to resolve before shipping.
The DOM is the actual bottleneck
A 240-file MR, side-by-side, with highlighting and comment affordances per row can produce more than 100 000 DOM nodes if everything mounts eagerly. That's where the viewer gets slow — paint, layout, scroll, and memory all scale with mounted node count, not with the diff's conceptual size.
The structural fix is two levels of virtualization:
- File-level: each file's diff block starts as a short placeholder with
the filename, status, and line counts. When the placeholder scrolls near
the viewport (
IntersectionObserver, generous rootMargin so there is no flash on scroll), the diff mounts. When it scrolls back out by enough, it unmounts and leaves the placeholder behind. - Line-level: within a mounted diff, only the window of visible rows (plus overscan) is in the DOM. Fixed-row-height windowing is enough — diff rows are uniform-height.
This bounds the DOM to a few hundred nodes regardless of how large the MR is. The 100 000-node case stops existing.
Syntax highlighting: visible on main thread, rest in a Worker
A synchronous tokenizer on a 12 000-line file blocks the main thread for most of a second — enough to feel like the whole page froze. The honest answer is to split the work:
- Step 1
- Step 2
- Step 3
The honest trade-off: highlighting trails scroll slightly on first visit to a big file. In exchange, the file is readable immediately and the viewer stays responsive. A session-level cache keyed by file content hash means the second visit is free — no re-work.
Server-rendering the highlighted HTML is tempting but trades the browser's CPU for a serious server CPU bill, and still ships a huge DOM. The Worker approach keeps the cost where the viewer is.
Inline comments: anchor to content, not line number
The specific failure mode: a reviewer comments on line 203. The author pushes a new commit that adds eight lines near the top. Line 203 is now line 211, and the old line 203 is a completely different piece of code. A viewer that anchors by raw line number is now showing the comment against the wrong line.
The right primitive is a content-plus-context anchor:
Content hash
Hash of the exact line the comment is on. Follows the line through re-numbering.
Context fingerprint
Hashes of the 2-3 lines above and below. Disambiguates identical lines elsewhere in the file.
Fallback: outdated marker
If neither matches after a force-push, the comment is marked outdated against the old revision and remains visible in the review history.
Comments follow the real line through re-numbering; they become "outdated" only when the content they anchor to actually changes. The outdated state is visible rather than silent — the reviewer sees "this comment was on a line that no longer exists" and can decide whether to re-post against the new code. GitHub and GitLab both do a version of this; the differences are largely in how aggressive the fingerprint matching is.
Side-by-side vs unified: a URL-driven view, not a mode state
The two views differ only in how the rows are laid out — side-by-side is a two-column grid of (old row | new row), unified is one column of alternating removed/added/context rows. The underlying diff state is the same, so the toggle lives in the URL (search param) rather than component state. Shareable, back-button navigable, and the user's preference is their URL rather than a hidden localStorage flag the next person won't understand.
What's explicitly out of scope, and why
Not solved in this answer: the diff algorithm itself (a backend concern — the viewer renders what the server computes); binary file diffs (a separate renderer that doesn't share the text-diff code path); the review state machine (approvals, required reviewers — a backend permissions model that the frontend reflects); conflict resolution UI (merging is a different feature). Scoping these out is a strength.
What to remember
- Two levels of virtualization — file-level
IntersectionObserver-driven mount, line-level windowing within an open diff — bounds the DOM independent of MR size. - Syntax highlighting is a split job: the main thread highlights the visible range, a Web Worker tokenizes the rest and posts results back. Highlighting trailing the scroll is acceptable; the file freezing is not.
- Inline comments anchor to a content hash plus a small context fingerprint, not a raw line number. Comments follow the line through re-numbering and are visibly marked outdated when the line content actually changes.
- Side-by-side vs unified is a URL-driven view, not a hidden mode — same rule as filters on a list.
- The diff algorithm, binary diffs, and merge resolution are deliberately not in this answer; naming what you're not solving is itself a strong move.
Check yourself
3 questions · pass 3/3 to unlock Worked Scenario: Design a CI/CD Pipeline Dashboard (GitLab / GitHub Actions)
1.A merge request has 240 changed files and the largest changed file is ~12 000 lines. A naive implementation renders every file's diff on page load and the browser becomes unresponsive. What's the right structural fix?
2.A reviewer adds an inline comment on line 203 of
app/foo.ts. The author force-pushes a new commit that changes line 203 to line 198 (lines added above). What should the diff viewer do with that pending comment?3.Syntax-highlighting a 12 000-line file blocks the main thread for ~900ms with a naive synchronous tokenizer. The design calls for the diff to feel responsive. Which approach honestly fixes this, and what's the trade-off?
3 left to answer
Discussion
Sign in to postNo comments yet. Be the first to say something.