---
title: "The 25-Hour Build Was a Queue"
description: "NOW TV's release pipeline took 25 hours and only about 30% of end-to-end runs passed, because nobody ran them. The problem was a queue for physical devices, not slow builds. Six months of measuring the wait and buying capacity where it was longest took a run to 1.5 hours and the pass rate past 95%."
author: "Filipe Brito Ferreira (Senior Front-End Platform Engineer)"
canonical: https://www.fbritoferreira.com/blog/the-25-hour-build-was-a-queue/
published: 2026-10-11
tags: ["ci-cd", "frontend-platform", "streaming", "testing", "developer-experience", "retrospective"]
image: https://cdn.fbritoferreira.com/images/the-25-hour-build-was-a-queue-1280.webp
---

In 2020, about three in ten end-to-end runs at NOW TV passed. The tests weren't the reason. A full run took 25 hours, so most engineers merged to master without running them, and the suite mostly met code after it had already broken. Thirty percent was the pass rate of a check nobody was using.

The 25 hours weren't compilation, and they weren't test execution either. They were waiting: for a Samsung television, an LG, an Apple TV, a Roku, or one of the games consoles to be free. The pipeline had hardware in it, and you can't spin up a second television the way you spin up a second container.

A pipeline with a shared physical resource in it is a queue, and a queue isn't fixed by making the work faster. It's fixed by measuring the wait, adding capacity where the wait is longest, and measuring again. In our case that cost 154 devices and a team of six over six months, and the run went from 25 hours to an hour and a half.

## Two PlayStations for five teams

The device pool when I started: 3 Tizen TVs, 4 webOS TVs, 10 Apple TVs, 10 Rokus, 2 PlayStations, 2 Xboxes, and about ten browser runners that nobody queued for, because a browser runs in a container and containers are cheap. Thirty-one physical devices. Five teams of 5 to 15 engineers, each with pipelines that needed the consoles.

Each device was a lock in a Concourse pool resource: a git repository of lock files, one per device, where a job acquires a lock on the way in and releases it on the way out. A console job holds its lock, which means it holds the console, for the length of a run. Lock files fail in their own way too: a job that dies mid-run leaves its lock behind, and the device sits idle until someone notices. With two PlayStations and five teams, the fifth team's job waits for four runs to finish before its own starts, and a team that pushes twice in an afternoon queues behind itself. A run that needs several device families waits for the slowest line of all of them.

Which is why the usual answer, run the tests in parallel, went nowhere. There was nothing to parallelise onto.

## Measure the wait, not the build

For the first three months this was one person's job, mine, reporting to a principal engineer. The first task wasn't buying anything. It was finding out where the time went, per device family, and separating two numbers the pipeline reported as one: how long a job waited for a device, and how long it ran once it had one.

That was harder than it sounds because of Concourse, the CI system that ran our pipelines. Concourse has an HTTP API, but the API isn't documented. As a contributor put it on the [project's own discussion board in 2022](https://github.com/concourse/concourse/discussions/8622), the API exists with no commitment to backward compatibility, and the reference is [a routes file in the source](https://github.com/concourse/concourse/blob/master/atc/routes.go). In 2020 it was the same. I read the routes, worked out which endpoints carried build and job timings, and built our own tooling on top to record wait time and run time by device family. Then the API moved under us. It isn't a public contract, so some of those endpoints changed across Concourse updates, and the tooling needed patching more than once in those months to keep the chart honest.

Every Friday that data became a short report for the NOW TV director: which stage was slowest that week, and where the next investment would shorten it most. Hardware spend stopped being an argument and became a line on that chart.

## Buy where the line is longest

PlayStation and Xbox went first. Two devices each for five teams was the worst ratio in the building, and the report said so every week until it changed. Then the TVs, the Apple TVs and the Rokus, as the report moved on to whichever line was longest next.

Six months in, the pool was 50 TVs split between Tizen and webOS, 45 Apple TVs, 65 Rokus, 10 PlayStations and 15 Xboxes: 185 physical devices, from 31. A full run took an hour and a half.

That is not a cheap fix, and it wasn't only hardware. Around month three the work stopped being one person: a team of six formed around the device pool and the tooling, and stayed. So the bill was 154 devices and a standing team, and I won't put a salary figure on the team, but six engineers for six months and beyond is the larger part of it. Every team asks for more hardware, so "buy more devices" isn't the lesson. The order is.

Buying in the order of the measured wait, consoles first, is what turned that money into time. I can't prove the counterfactual; what I can say is that the console line was the first thing the report flagged and the first thing we bought. That team became Web Core, the group that later owned the shared build and deployment tooling across the [territory consolidation](/blog/consolidating-12-streaming-apps-into-2/).

## The number that moved was the pass rate

25 hours to 1.5 is the headline. The number I'd defend first is a different one.

The end-to-end pass rate went from about 30% to over 95% across the five teams. The suite did change in those months: we added retries and fixed flaky tests, and both moved that number. Retries flatter a pass rate, since they hide a flake rather than fix it, and I don't have the split between the two. But nobody fixes flakes in a suite they never run. The fixes came after a run fit inside a working afternoon, which is when people started triggering it before they merged and seeing the failures. The queue was the enabler; the rest was ordinary work that had been impossible to start.

## What I'd do differently

Get the measurement out of Concourse in the first week, not the second month. The undocumented API cost more calendar time than any single device, and every week without the chart was a week the wrong thing could have been bought. And publish the Friday report to the teams, not only upward. The people waiting in the queue could have shortened it themselves, by batching pushes or moving tests off hardware, and they never saw the chart that would have told them where the line was.

## Queues, not pipelines

When a pipeline has a shared physical resource in it, treat it as a queue. Measure wait time and run time separately, per resource. Put capacity where the wait is longest, then measure again. Over those six months the build itself never got faster. Only the wait changed.
