Skip to content

/blogentry

August 9, 2026

⤢ Title

Cambridge Analytica: The Scandal Was Real. The Dataset Was Wrong.

I was inside Cambridge Analytica in March 2018. The story everyone knows names the wrong dataset. The real engine is still legal and still running.

cambridge-analytica / data / privacy / politics / psychographics


Rows of rolling archive shelves with numbered bays in a dimly lit records room

The first sign was an email. A Monday in March 2018, company-wide, subject line carefully bland, telling us a news story about Cambridge Analytica was about to air and that our CEO, Alexander Nix, had been filmed by an undercover reporter. Nix had spent four months taking meetings with a man he believed was a fixer for a wealthy client seeking election work in Sri Lanka. The man worked for Channel 4. There were cameras.

We’d launched Pangea the week before. A GDPR-compliant data platform, months of work, built so the company’s data operations would clear the strictest privacy law in the world before it came into force that May. We shipped it, and seven days later the world decided we were the data scandal of the decade.

The meetings started that morning and didn’t really stop. All-hands after all-hands, a rolling attempt to work out what was happening while it was happening. The office went quiet the way offices go quiet when everyone is reading the same thing at the same time. Outside, the story grew by the hour: 50 million Facebook profiles, then 87 million, psychological warfare, democracy hacked. Inside, the people who’d built the systems sat in a conference room listening to the world describe them.

And the description was wrong. Not wrong the way every company under siege claims the coverage is wrong. Wrong on the mechanism. The dataset the world spent 2018 furious about wasn’t the dataset that did the work. I was in the building for three years. Eight years on, I want to tell you what I actually saw.

The scandal was real. The dataset was wrong.

What Everyone Knows Happened

In the years since, I’ve had the story told back to me more times than I can count. In job interviews, at dinners. Someone asks where I worked before, I answer honestly, and then they tell me what happened at my old company. It’s always the same version.

A political consultancy paid a Cambridge academic, Aleksandr Kogan, to run a personality quiz on Facebook. The quiz harvested the quiz-takers and, thanks to the platform’s permission model, their friends, 87 million profiles by Facebook’s final count. Cambridge Analytica fed those profiles into personality models, built “psychographic” profiles of voters, and used them to swing the 2016 US election. Depending on who’s telling it, Brexit too.

The story had everything. A whistleblower with pink hair. A CEO on hidden camera offering honeytraps and Ukrainian sex workers. A Netflix documentary. Zuckerberg in front of Congress. It ended with Facebook paying a $5 billion fine and Cambridge Analytica filing for insolvency by May 2018.

I never argue with the outrage. The Kogan collection was a genuine consent failure. Facebook’s API allowed it, people’s data moved without their knowledge, and nobody in that chain comes out clean. If the scandal had been prosecuted as what it was (a platform whose permission model leaked friend data to thousands of apps for years), every word of the coverage would have held up.

Where I go quiet is the mechanism. The version people tell me needs the Facebook dataset to be the engine of everything that followed: harvest the likes, model the hidden fears, whisper to each voter’s weaknesses, collect the votes.

It wasn’t the engine. It was barely in the car.

The Facebook Data Wasn’t the Engine

When I joined SCL Elections, the London company behind Cambridge Analytica, in the summer of 2015, nobody handed me a stolen Facebook dataset. What they handed me was the voter file.

If you’ve never worked in American politics, the voter file is the part nobody believes. Every state publishes one: who’s registered, their address, their party registration where applicable, and whether they turned out to vote in each election. Not who they voted for. Whether they showed up. That file is public, legal, and for sale in cleaned-up commercial form to anyone with a chequebook. It was the first thing I learned about the business, and it reframes everything else.

The data I worked around came from the industry’s standard vendors: L2, Aristotle, Acxiom, Experian, Infogroup, Magellan Strategies. You don’t have to take my word for that. Nix told a parliamentary committee exactly that when it recalled him in June 2018, and when the UK regulator seized the servers it recovered datasets from those exact vendors. Aristotle has sold political and consumer data since 1983, and still sells it today. L2’s national voter file carries hundreds of attributes per voter: ethnicity (the kind of field European law treats as special-category), military status, net worth, pet ownership. All of it commercial. All of it still on sale today.

On top of the purchased data sat the part we built: surveys. More than 150,000 household surveys, continuously refreshed, anchoring OCEAN personality models: openness, conscientiousness, extraversion, agreeableness, neuroticism. The models predicted which traits dominated for a voter and matched the message framing: a security-framed ad for high neuroticism, a tradition-framed one for high conscientiousness. That’s psychographics. Survey ground truth, commercial data, regression. The published research on personality-matched ads finds lifts that are real but modest, and contested; field replications are mixed, which is the fight the two camps have been having ever since. At campaign scale, modest is exactly what you pay for. Boring, legal, effective at the margin: adjectives no headline ever used.

The tooling was ours too. Syphon, the data-visualisation platform, sat on the company website as a product. Anyone could have looked it up at the time. Pangea was the GDPR-compliant data platform we finished a week before the sting aired. I can’t link you to either; the archive crawlers didn’t keep the pages, so this part you’re getting from someone who was there.

And the Kogan data? By the time I arrived it was already the awkward drawer of the data estate. Facebook had shut the friends API the year before, the records were ageing, and the matching rates against voter files were poor. Nix later told Parliament it had proved “fruitless”, although MPs produced his own 2014 emails boasting the opposite, so treat his word the way this post treats everything he sold: carefully. My version is simpler. The pipeline I watched didn’t run on it.

Two streams. One: commercial data, surveys, voter files. That stream ran the campaigns. Two: a Facebook dataset with a rotten consent story and poor operational value. The public collapsed them into one, and every fix that followed aimed at the wrong stream.

The Narrators Weren’t in the Room

The world learned the Cambridge Analytica story from two people. I want to tell you where each of them sat, because I was sitting there too.

I never met Christopher Wylie. He was SCL’s research director from 2013 to July 2014. He was gone a year before I badged in, eighteen months before the Iowa caucus. Every piece of campaign work he later narrated on camera happened after he left. And a caveat before the next part, because receipts aren’t rebuttals: none of what follows invalidates what Wylie disclosed about the Kogan harvesting. The consent failure was real no matter who described it, and a man can be self-interested and still tell the truth.

Inside the building he wasn’t a ghost. I remember talking about him with our then head of data science, because there was a lawsuit: after leaving, Wylie had founded a competing firm, Eunoia Technologies, and pitched the same services, including to Corey Lewandowski, Trump’s eventual campaign manager, in spring 2015. SCL sued after a client forwarded a Eunoia proposal offering an identical service list. The case ended in 2015 with Wylie signing undertakings not to use SCL’s intellectual property. And per the ICO’s letter to Parliament, the Kogan Facebook data had been shared with Eunoia as well. When he appeared on my TV in March 2018 as the man exposing the toxic dataset, I already knew him as the man whose startup had held a copy. What the dates change is narrower than a gotcha: he wasn’t describing work he ever saw.

Brittany Kaiser I knew of differently. She joined in December 2014, recruited by Nix for her Washington connections (she’d spent the summer of 2007 on the Obama campaign’s new-media team, helping run the candidate’s Facebook presence) and rose to Director of Business Development. She sold. That was the job, and the thing she sold was the sales deck: the “5,000 data points on every American voter” pitch the regulator later singled out as exaggeration. When she became the second famous whistleblower, the capabilities she described to the documentary cameras were the ones from her own slides. And she did the whistleblowing part for real: she handed documents to Parliament and testified. But whistleblowing tells you someone opened the door, not which rooms they’d been in.

One narrator left before the work existed. The other sold the hype the regulator flagged. Neither fact makes them liars, and I’m not asking you to un-believe them. This isn’t a character judgment; it’s a map of where people sat. Between them they supplied the world’s entire mental model of what the data science floor was doing: a floor one of them never stood on and the other walked through on the way to client meetings.

It Worked, and Both Sides of That Fight Are Wrong

The night of the Iowa caucus, February 2016, was the first time I watched something we’d built show up in the world. Ted Cruz beat Donald Trump, an upset almost nobody’s polling had picked, with a campaign that had paid Cambridge Analytica close to $6 million for exactly the survey-anchored OCEAN targeting I described above. The campaign’s own staff credited the data operation. I’ll flag the obvious myself: a caucus win has many parents. Cruz’s ground game was famously good, his evangelical turnout operation better, and no campaign is a controlled experiment. What Iowa proves is thinner than either camp wants, and still worth something: the data operation was real, the campaign leaned on it, and the campaign won. I remember what that felt like on our side of it.

I bring it up because there are two camps on Cambridge Analytica, and they’re both wrong in different directions.

The panic camp believes the machine read voters’ minds through their Facebook likes. You’ve already seen that this rests on the wrong dataset. The backlash camp includes sharp people, and they concluded the whole thing was snake oil: psychographics doesn’t replicate at scale, Trump’s own digital director said he didn’t use it, the company was all sales patter. Reason ran the definitive version: the election interference operation that wasn’t.

Half of that critique lands, and I watched it land from the inside. The sales deck was hype. The gap between the deck and the floor was visible to anyone who worked on the floor. The regulator found the “5,000+ data points per individual” marketing claim didn’t survive contact with the seized servers, and Nix’s hidden-camera patter earned every ounce of ridicule it got. If you bought the pitch, you were oversold. Our clients were oversold.

But “the pitch was inflated” and “the capability was fake” are different claims, and the second doesn’t follow from the first.

On the Trump work I’ll give you the internal number, and tell you exactly how much it’s worth. The figure that went around the floor was that the campaign’s digital operation reached persuadable voters at roughly five times lower cost than the Clinton side. That’s folklore, not evidence. No audit, no public document; what the dashboards suggested and what people repeated. And it describes modelled targeting on the commercial data, the bread-and-butter kind every serious campaign now runs, more than it describes the personality layer specifically. The public record, for what it covers, points the other way on price: Facebook’s own ad data shows Trump paid higher average CPMs than Clinton through most of 2016. Both can be true, because CPM prices an impression shown to anyone and the internal number claimed to price a persuadable voter reached. But treat the five-to-one the way you’d treat any number a company tells itself: interesting, unverified, and flattering to the people repeating it.

The honest verdict: the marketing lied about the ceiling, and the floor is harder to prove than either camp admits. What I can tell you is what the buyers did. They kept coming back, and the industry that does this work today is bigger than Cambridge Analytica ever was.

The Regulator Already Told You

You don’t have to trust a former employee. In the same March the email arrived, the UK Information Commissioner’s Office raided the office and took the evidence with them: 42 laptops and computers, 31 servers, over 700 terabytes of data, more than 300,000 documents. They spent more than two years going through it. It remains one of the largest investigations a data-protection authority has ever run, and it was an investigation of the systems I’d worked around every day.

In October 2020, Commissioner Elizabeth Denham sent her closing letter to Parliament. It reads nothing like the 2018 coverage. The company’s data holdings were, in her words, mostly commercially available datasets: the L2s and Aristotles and Acxioms from earlier in this story. The models were built with commonly used, off-the-shelf analytical tools. The investigation found no evidence the Kogan Facebook data was used in the Brexit campaigns, and did not establish it was deployed in the US 2016 work either. The marketing claims were flagged as exaggeration. Two years with our servers, and the regulator’s letter lines up with both halves of what I just told you: real models on commercial data, inflated pitch on top.

One caution, because precision cuts both ways. “Found no evidence” and “did not establish” are non-findings, not exonerations. The letter can’t prove the Facebook data was never used anywhere; it says two years with the servers couldn’t show that it was. What ran in the campaigns, I’m telling you from memory. What the regulator tells you is that they went looking for the scandal’s version and couldn’t find it. Different kinds of claims. Weigh them differently.

The letter landed three weeks before a US presidential election, two and a half years after the scandal. It got a day of muted tech-press coverage and disappeared. There was no documentary about the letter. The story had finished in May 2018 when the company died; the verdict arriving later was an administrative detail.

I don’t blame anyone for missing it. But it means the most thorough independent examination of Cambridge Analytica that will ever exist, the one with the actual servers, agrees with the version I’m telling you, not with the one you remember.

The Investigation Was Its Own Kind of Theatre

The theatre had started the weekend before the raid. Facebook suspended Cambridge Analytica from the platform on a Friday night. By Monday its people were in our office alongside Stroz Friedberg, the forensics firm Facebook had hired to audit us, and the understanding inside the building was that an announcement was coming: suspension lifted, pending an independent investigation. It never went out. That same day the Information Commissioner went on Channel 4, live, saying she was seeking her own warrant and wanted the investigation for herself, and Facebook’s auditors stood down at the ICO’s request that evening. The independent forensic audit of the actual data never happened. What we got instead came four days later.

Here’s what the raid looked like from our side. The warrant was executed at eight on a Friday evening, about twenty enforcement jackets into an office that had mostly gone home for the weekend, and they left at three in the morning. If you want to image employees’ laptops while they’re in use, you come on a Tuesday morning. If you want footage for the evening news, you come after dark, when the only things left to seize are the ones bolted into racks. Among the “servers” they carried out that night, as we understood it inside, was the internet router. And, as I remember it, on the desk of our data compliance officer sat hard drives holding customer data. They left those where they were.

The part that stayed with me was the emails. The ICO didn’t run a forensics operation of this size in-house; the analysis was contracted out. And while the investigation ran, we were asked, over email, to hand over root passwords to servers so they could be passed to a third-party firm doing the work. This is my recollection, and you won’t find it in a press release. But sit with the picture: I’d spent that spring building a platform to clear the strictest privacy law in Europe, and the privacy regulator was requesting root credentials over email to forward to contractors.

None of this happened in a vacuum. GDPR was two months from coming into force, and the ICO was negotiating its post-GDPR powers in public with our case as the exhibit. Denham spent the week before the raid telling cameras she was waiting for a warrant and asking for a lower threshold and faster process. The Data Protection Act 2018 gave the ICO the stronger powers it asked for. The timing wasn’t lost on anyone inside our office.

You can hold all of this against their conclusions if you like. I read it the other way. An investigation this public, with two years, every disk it did seize, and a mandate riding on the outcome, was under every pressure to find the scandal’s version, and still told Parliament it couldn’t. A clumsy search plus a strong motive plus a non-finding is about as convincing as a non-finding gets.

The Fixes Fixed Nothing

The company died in May 2018. I watched the aftermath the way everyone else did, from the outside, except I knew what the fixes were supposed to be fixing.

Facebook paid the FTC $5 billion: the largest privacy penalty in history, about 23 days of the company’s profit at the time. It admitted no wrongdoing. Its stock rose 22% over the year of the settlement. The consent order restructured Facebook’s privacy compliance paperwork; it didn’t ban psychographic profiling, didn’t restrict political microtargeting, and didn’t touch a single data broker. In fairness, it was never meant to. A consent order against one platform can’t regulate an industry. A few state broker registries and California’s CCPA arrived in the years after; the trade continued through all of them.

The platform locked down its APIs, which sounds like the fix until you ask who used them. The Kogan-style friends access was already dead, announced closed in 2014 and shut for good in May 2015, three years before the scandal. What the post-2018 lockdown actually killed was researcher and academic access, the kind that had let outsiders study the platform. The commercial voter-data industry didn’t use Facebook’s APIs. It didn’t need them.

And that industry is fine. L2 sells the national voter file with hundreds of attributes today, to anyone. Aristotle has been in business since 1983 and still is. Campaigns buy phone-location data now, bought and sold through the same broker channels: the “data broker loophole” American law still hasn’t closed. Every US campaign since 2016, both parties, runs modelled persuasion targeting on commercial data. People I worked with scattered into that industry and kept doing the same work. The technique wasn’t buried with the company; the company was buried so the technique could be forgiven.

Here’s the detail I keep coming back to, eight years later. We built Pangea because GDPR was arriving and the company intended to operate under it (a law that reached the European side of the data estate, not the US voter files). The compliant platform shipped a week before the sting aired. The scandal killed the one company that had just finished doing what the scandal’s aftermath never made anyone else do.

What You Do With a Corrected Record

I’m not asking you to feel differently about Cambridge Analytica. The consent failure was real, the sting tapes were real, and a company that lets its CEO talk like that on camera doesn’t get to complain about how it died.

I’m asking for something smaller: get the mechanism right. The mind-control machine built on stolen Facebook data was mostly theatre. The real machine was duller. A campaign buys a file from a commercial broker. The file has a row for every registered voter in America, and your row costs a fraction of a cent. Nobody hacked your social graph, and nobody needed to. We bought our rows from that shelf, and campaigns are buying theirs from it right now.

Eight years on, the record deserves better than the Netflix version, so I’ve started writing the longer story down: what those three years actually looked like from the inside, starting long before the week the email arrived. The rest is becoming a book.