---
title: "Owning the BFF: A Frontend Team Running GraphQL"
description: "At Sky the frontend team built and ran its own GraphQL BFF: response times halved at 99.9% uptime, and the failure modes were ours to fix. What ownership actually cost and what it bought."
author: "Filipe Brito Ferreira (Senior Front-End Platform Engineer)"
canonical: https://www.fbritoferreira.com/blog/owning-the-bff-what-a-frontend-team-learns-running-a-graphql-layer/
published: 2026-09-19
tags: ["graphql", "bff", "node-js", "architecture", "streaming", "frontend-platform", "retrospective"]
image: https://cdn.fbritoferreira.com/images/owning-the-bff-what-a-frontend-team-learns-running-a-graphql-layer-1200.webp
---

A single page in one of our apps used to fan out to dozens of REST endpoints: catalogues, profiles, entitlements, recommendations, each with its own shape and its own quirks. The frontend stitched all of it together, slightly differently in every app, and paid for that on every render. After we built our own GraphQL backend-for-frontend, those scattered calls collapsed into a handful of batched ones, response times halved, and the layer held 99.9% uptime across Sky GO, NOW TV, NOW, Peacock TV, and Sky Showtime. In October 2022 I [talked through the architecture on Syntax.fm](https://syntax.fm/show/529), and there's a [deeper write-up of it](/blog/supper-club-graphql-as-an-aggregation-layer-with-filipe-ferreira-of-sky-tv/) on this site already.

What I want to write down here is the other part: what changed because the frontend team owned the service itself. Whether to have a BFF was never really the argument at Sky. Who builds it was.

## The ticket queue is the tax

The default arrangement is that a backend team builds an API layer and frontend teams file tickets against it. Every data-shape change becomes a negotiation across a team boundary. The catch is that the frontend team holds the knowledge: we're the ones watching the waterfall, the bundle size, and the time-to-playable every day. Filing tickets to translate that knowledge to another team loses fidelity twice — once going in, once in the interpretation back.

So we built and ran the BFF ourselves.

The most immediate change was that schema design became a frontend decision. We shaped the schema around what screens needed rather than what backends happened to return, and when a screen needed data reshaped, the resolver changed that afternoon. No cross-team negotiation, no version dance. It took a while to trust how short that loop was.

The second change crept up on us. "The API is slow" is a phrase that can circulate in ticket systems for months. When the layer is yours, the person who notices the slow screen can open the resolver timing that explains it, in the same repo, on the same day. Most of the delivery speed didn't come from GraphQL at all — it came from frontends no longer waiting on anyone to investigate their problems.

The honest cost: frontend engineers had to become service engineers. On-call, deploys, capacity, Redis, resolver-level monitoring. That's a real investment in people who signed up to build interfaces, and I don't think the model works without it.

## The caching is where the wins actually lived

Halving response times sounds like a GraphQL feature. It wasn't. It happened because owning the layer gave us one place to cache properly, and we layered it:

- Per-request in-memory deduplication, so a single render that fans out to many resolvers doesn't hit the same backend service twice.
- DataLoader-style batching, collapsing what used to be 30+ scattered REST calls per page into a handful of batched calls.
- Redis for shared caching across instances, so one user's cold cache warms everyone's.
- CDN caching for public content queries, taking catalogue traffic off our infrastructure entirely.

The stack is the least interesting part of that list. What mattered was concentration: when every frontend does its own orchestration, each one builds a partial cache with its own invalidation bugs, and nobody owns the whole picture. We did.

## Failure modes, learned the direct way

Running a service whose users are your own colleagues concentrates the mind. They notice immediately, and they tell you at their desk.

The first lesson was that REST monitoring is useless here. GraphQL has one endpoint, so endpoint-level metrics tell you almost nothing — a query can be 95% fine and broken in the one resolver that matters. Per-resolver timing and error instrumentation wasn't a nice-to-have; without it the layer is a black box exactly when it's on fire.

N+1 problems don't disappear in GraphQL, they move. Fan-out means one query can mean hundreds of backend calls if a resolver fetches naively per item. Batching fixes it, but only everywhere it's applied — the failure mode returns the first time someone adds a resolver and forgets.

Cache invalidation under live content was the ongoing fight. Availability windows, entitlements, catalogue updates: cache too short and you've built an expensive pass-through, too long and customers see content they can't watch. We ended up tuning per data type — public catalogue data caches aggressively and long, anything entitlement-flavoured briefly or not at all. The rule I took away: classify every query by how wrong a stale answer can be, and let that decide its TTL.

Partial failure turned out to be the normal case. With a long list of backend dependencies, one of them is usually degraded, and the BFF has to decide per field whether to fail the query, serve stale cache, or return partial data with errors. The frontends consuming it had to be built for partial responses too. Designing that contract — what degrades gracefully, what must never return wrong data — was more important work than any happy-path feature.

One structural side effect surprised me: because breaking the schema breaks products we were also responsible for, we got disciplined about schema governance in a way external API teams rarely have to be. Typed schema-to-resolver-to-frontend caught changes at compile time, and explicit ownership rules kept the layer from becoming a knowledge silo.

## What I'd do differently

Lean on the CDN earlier. We tuned application-level caching thoroughly before fully exploiting the edge, and public catalogue data had no business hitting our infrastructure as often as it did. And formalise ownership sooner: "we built it, we run it" carries a small team a long way, but explicit rules about who could change which schema domains would have saved real coordination pain later.

## Carry the question, not the pattern

I don't think every frontend team should run a BFF. What convinced me long-term wasn't the response times — it was that the boundary is real ownership. At Roku I later architected shared media web frameworks adopted across 10+ device platforms and built multi-LLM routing over retrieval for support, and both times the deciding question was the one the BFF taught me: which team is closest to the knowledge this layer encodes? Give that team the layer. You'll make mistakes, but they'll be yours to fix, and that beats mistakes you can only negotiate about.
