Lesson 30 of 36
Worked Scenario: Design a CI/CD Pipeline Dashboard (GitLab / GitHub Actions)
A full worked answer to the GitLab-style pipeline-UI prompt — a stage graph that keeps up with live status, a log viewer that handles 50MB builds, and a history table that stays usable across thousands of runs.
A pipeline dashboard — GitLab CI, GitHub Actions, Buildkite, CircleCI — is one of those prompts where the hard problems are hiding in plain sight. The stage graph looks small until you need it to update live without reflowing. The log viewer looks trivial until the job writes 40 megabytes. The history table looks ordinary until it's got five thousand rows and the product wants them filterable.
Clarifying requirements first
Before proposing anything, the questions worth asking out loud: Rough worst case for a single job's log output — megabytes, tens of megabytes? Longest pipeline — number of stages, number of jobs per stage? How much history is shown — last 100 builds, last 10 000? Do users need to re-run or cancel from the UI, or is this read-only? Mobile first-class or secondary? For this answer, assume: up to ~50MB of log output per job in the worst case; up to ~80 nodes across 12 stages for the biggest pipeline; the history table may hold tens of thousands of rows for a mature project; re-run/cancel are first-class; desktop-first with a usable mobile view.
Rendering strategy: SSR shell, CSR everything-that-changes
The dashboard is behind a login. SSR for the shell — project header, pipeline summary, the current run's top-level status — makes the first paint feel instant on cold navigation. The three live surfaces (stage graph, log viewer, history table) are CSR, driven by a WebSocket (or server-sent events) that pushes status transitions. Streaming SSR for the shell is a cheap bonus: the project header ships before the current run's details resolve.
The stage graph: layered DAG layout, not nested flex
A pipeline's stage graph is a directed acyclic graph with cross-stage
dependencies — a job in stage 4 can depend on a specific job in stage 2.
Trying to express that in nested CSS flex produces a layout that re-flows
awkwardly every time a node's size changes with its status (running has a
spinner, failed has an error chip, success is neither).
- Step 1
- Step 2
- Step 3
- Step 4
The honest cost: the layout algorithm runs on load (and on structural changes to the graph, which are rare — mostly the pipeline shape is fixed once started). Every status tick after that is a paint-only update. For 80 nodes this is a non-issue; the alternative (CSS flex re-lay-out on every status change) stutters visibly.
Canvas is tempting at larger graph sizes but the trade-off is severe: hit-testing, accessibility, and selection all get harder, and 80 nodes isn't where DOM breaks down. Keep SVG+DOM until the node count genuinely forces canvas.
The log viewer: virtualized text with a bounded live tail
Logs stream over the WebSocket as the job runs. The viewer has three operating modes it needs to switch between cleanly:
- Tailing (default while running): new lines append and the view auto-scrolls to the bottom. Auto-scroll turns off the moment the user scrolls up — a critical UX detail that this course has not covered elsewhere, and the one most log viewers get wrong.
- Scrolled back (reading history): no auto-scroll. A floating "jump to latest" button is the way back to tailing mode.
- Finished (job done): a terminal state where the full log is static and can be virtualized plainly.
Memory discipline: the log lives in a plain JS array in state, capped at some reasonable in-memory size (say 20 000 lines). Older lines are offloaded; a "load earlier" affordance fetches them on demand from the backend. The virtualized list only renders the visible range into the DOM, so a 40-minute job does not grow a 2GB tab. Fixed row height — log lines are uniform — so the simpler windowing works.
ANSI colour and search without blocking
Log lines carry ANSI escape codes. Parse-on-render is cheap for the visible window (dozens of lines, not tens of thousands); don't precompute every line's HTML upfront. Full-log search is a Web Worker job: the raw string array is transferable, the worker scans and posts back hit ranges, the viewer renders a highlighted overlay on the matching rows. Doing it on the main thread blocks typing in the search box, which is exactly the feedback users need.
The run-history table: server-filtered, URL-stated, paginated-on-scroll
URL search params
Current filter (branch, status, committer, date range) — the single source of truth. Shareable, back-button navigable.
Server query
The URL params map directly to a server query. Pagination is cursor-based: give me the next page after row X.
Client cache
Keyed by (filter signature, cursor). A back-navigate restores the previous page from cache rather than re-fetching.
In-memory client filtering is tempting for the small-dataset case and wrong by the time the project has 10 000 builds — "fetch everything first" means each visit downloads megabytes of mostly-irrelevant history. Server-side filtering with URL-encoded params is the right primitive from day one; the cost is one extra network round-trip on filter change, which is what the user already expected.
Real-time updates on a running pipeline
A WebSocket channel scoped to the pipeline id streams events: node status
transitions, log line batches, pipeline-level status changes. Apply them
idempotently against the normalized pipeline state — a node's status is
written to its slot in nodes.byId, log lines append to the active job's
log array, pipeline-level status refreshes the header. On WebSocket
reconnect, re-fetch the current pipeline state rather than trusting a
server to replay missed events in order; same reasoning as the Kanban
scenario.
What's explicitly out of scope, and why
Not solved in this answer: the pipeline execution itself (a backend infrastructure problem); artifact browsing (a separate surface with its own storage model); the YAML pipeline editor (that's a code-editor problem, with its own prompt); the test-results view with per-test history (its own set of scaling concerns around tens of thousands of tests). Naming these as scoped-out shows the answer is bounded on purpose.
What to remember
- Compute the DAG layout once with a layered-graph algorithm, absolutely-position the nodes, render edges as SVG. Status updates then only re-paint, never re-layout.
- Treat the log as an append-mostly virtualized list with a bounded in-memory cap; "load earlier" fetches older chunks on demand. 40MB of logs is a specific solvable problem, not an open-ended one.
- Auto-scroll in the log viewer is user-controlled, not application-controlled: scrolling up turns tailing off, the "jump to latest" button turns it back on.
- The run-history table is server-filtered from day one, with the current filter in the URL — client-side filtering of 10 000 rows is a dead-end you don't need to walk down.
- Real-time events go against a normalized store idempotently; a WebSocket reconnect is a re-hydrate, not a replay.
Check yourself
3 questions · pass 3/3 to unlock Worked Scenario: Design a Browser Code Editor (Monaco / VS Code Web)
1.A long-running CI job streams ~40MB of log output over the course of an hour. The naive 'append each line to a scrolling panel' breaks the browser long before the job finishes. What's the right structural fix?
2.The stage graph for a complex pipeline is a DAG — 80 nodes across 12 stages, with cross-stage dependencies. Rendering it as nested CSS flex boxes produces a layout that re-flows awkwardly on every status update. What's the right rendering approach, and why?
3.The run-history table shows thousands of past builds. The product wants fast filtering by branch, status, and committer, with a shareable URL for every filter. Where does the filtering happen, and where does the state live?
3 left to answer
Discussion
Sign in to postNo comments yet. Be the first to say something.