Lesson 34 of 36
Worked Scenario: Design a Video Player (Netflix / YouTube)
A full worked answer to a Netflix/YouTube-shape streaming prompt — adaptive bitrate, custom controls, DRM boundaries, performance under player resize, and accessible captions.
A Netflix / YouTube-shape video player is a prompt that reveals whether a
candidate has thought about streaming as a system or treated it as a <video>
tag with buttons. The question is less about which library to use and more
about which concerns the frontend actually owns.
Clarifying requirements first
Before proposing anything, the questions worth asking out loud: VOD (video on demand) or live? Protected content (DRM required) or open? Low-latency a requirement (live sports) or standard streaming latency acceptable? What range of network conditions must the player handle — residential-only or including slow mobile? Multi-track audio + subtitles required? Mobile first-class? Analytics / QoE telemetry needed? For this answer, assume: VOD, DRM-protected (Widevine/FairPlay/PlayReady), standard streaming latency, network range from slow mobile to fibre, multiple audio tracks and subtitles, mobile first-class, full QoE telemetry.
Rendering strategy: pure CSR for the player, SSR for the page around it
A video player is an interactive widget — personalised, authenticated, not indexable. Pure CSR. The surrounding page (title, metadata, episode list, recommendations) can SSR for a fast first paint; the player itself mounts client-side once its container is in the viewport.
Streaming primitive: ABR over HLS or DASH
The naive approach — a single MP4 at a chosen resolution — breaks the moment the network dips. The honest primitive is adaptive bitrate (ABR) streaming over HLS or DASH:
- Step 1
- Step 2
- Step 3
- Step 4
For the frontend design round, you don't need to describe an ABR algorithm from first principles — the standard libraries (Shaka Player, hls.js, Video.js) implement well-known ones (BOLA, throughput-based). Your job is to know what the primitive is, what the client owns around it (quality picker, buffer-health indicator, stall telemetry), and when a custom algorithm is worth it (almost never for standard VOD).
DRM: the trust boundary
If content is protected, playback happens inside the browser's EME (Encrypted Media Extensions) layer. The frontend's contract:
License server (backend)
Issues short-lived decryption keys scoped to the user, device, and content.
EME in the browser
Handles the key exchange and the actual decryption of each segment — opaque to JS.
Your player code
Mediates: tells the player library to request a license when the manifest signals DRM, hands the response to EME, surfaces errors (unsupported device, expired license) to the user.
Trying to decrypt in JS yourself is a non-starter — the keys never leave EME. The frontend's job is orchestration, not cryptography.
Custom controls: not optional for a serious player
Native <video> controls are inconsistent across browsers, lack keyboard
polish, and can't surface quality or audio-track pickers. A custom control
bar is standard:
- Play/pause, scrub bar with buffer visualisation, time, volume, mute.
- Quality picker (manual override of the ABR choice, 'Auto' as default).
- Audio-track picker (dubs).
- Subtitle picker.
- Fullscreen (via the Fullscreen API) and picture-in-picture (via PiP API).
- Keyboard: Space play/pause, arrow keys seek, J/K/L (YouTube-shape) if desired, Esc exits fullscreen.
Each control is a button with a visible focus ring and an accessible label. A player that fails a keyboard user is not shippable.
Captions: WebVTT via <track>, styled for legibility
Captions are a <track kind="subtitles" src="…vtt" srclang="…"> element.
The browser handles timing; your code handles:
- A picker UI listing available tracks, invoking the native TextTrack API to switch.
- A styled cue container — a translucent scrim behind the text so captions stay legible against bright scenes, outline on the text for extra contrast.
- Respecting the OS/browser caption preferences (font size, colour) where available.
Burning captions into the video kills selectability (one burned language per file) and accessibility. Manual timeupdate-driven overlays drift and miss edge cases.
Seek, scrub, and buffer discipline
A seek into an unbuffered position triggers a silent stall unless the player intervenes:
- Step 1
- Step 2
- Step 3
- Step 4
Scrubbing (dragging the progress bar) is the continuous variant — the same cancel-and-refetch pattern applied continuously. Debounce the actual seek network request so a drag doesn't fire dozens of segment requests.
Performance: resize is the hidden gotcha
A <video> element re-decoded on every resize is a hidden performance
disaster. The container's dimensions change on fullscreen, picture-in-picture,
and simple layout shifts; the video element itself should hold a stable
pixel size and let CSS scale the output. GPU compositing (transform: scale)
beats CPU re-render.
Frame-rate matters too: a player should respect the OS's reduced-motion setting for any UI animation (progress-bar transitions, control-bar fade-in).
QoE telemetry: five numbers that matter
For a serious player, operational telemetry is non-optional. The five numbers to emit:
Startup time
Click 'play' to first frame rendered. The number that trades off against SEO.
Rebuffer ratio
Fraction of playback time spent buffering. The number ABR algorithms are optimised against.
Average bitrate delivered
Weighted average of segment renditions played. Trade with rebuffer ratio.
Error rate
Fraction of sessions ending in an unrecoverable error. Breaks DRM reliability or network robustness problems to the surface.
Playback failure to start
Sessions where play was pressed and first frame never rendered. The hardest failure to debug and the most important one to catch.
The frontend emits these; a backend pipeline aggregates them per rendition, device, region. Candidates who name these specifically land ahead.
What's explicitly out of scope, and why
Not solved in this answer: the manifest-generation pipeline (a backend encoding and packaging concern); the ABR algorithm's internals (library code, rarely worth re-implementing); social features (comments, likes); peer-to-peer delivery (its own architecture). Scoping out is a strength.
What to remember
- ABR over HLS/DASH is the primitive. The frontend owns the UI around it (quality picker, buffer indicator, stall telemetry) and the orchestration, not the algorithm.
- DRM lives inside EME; the frontend's job is to pass licenses between the server and the browser, never to touch keys.
- Custom controls are table stakes — native controls fail keyboard and feature-coverage expectations.
- Captions are WebVTT via
<track>, with a styled cue container and a respect for OS caption preferences. Burned-in captions are an anti-pattern. - Seek into unbuffered: show the buffering state, cancel pending segments, fetch a low-rendition segment for fast resume, raise quality once stable.
- QoE telemetry has five canonical numbers: startup, rebuffer ratio, average bitrate, error rate, failure-to-start. Candidates who name them are credible.
Check yourself
3 questions · pass 3/3 to unlock Worked Scenario: Design an Analytics Dashboard (Datadog / Amplitude)
1.A video player needs to handle variable network conditions without the user manually choosing a resolution. What's the right primitive, and what does the frontend actually own in it?
2.Captions must be togglable, selectable per language, and legible against bright backgrounds. What's the honest implementation?
3.A seek operation on a long video jumps to a point not yet in the buffer. What happens by default, and what should the player do explicitly?
3 left to answer
Discussion
Sign in to postNo comments yet. Be the first to say something.