Clearview AI in 2020: Three Billion Scraped Faces, One Leaked Client List

📋 Key Takeaways
  • What happened?
  • The scraped three billion
  • The paper trail
  • The client list as confession
  • Platforms respond with paper
8 min read · 1,556 words
Educational & Ethical Use Only — This article is provided for educational and ethical cybersecurity research purposes only. The techniques described should only be used on systems you own or have explicit permission to test. Always follow responsible disclosure and the laws applicable to you. Mitigations are included so engineers can harden real systems.

What happened?

In late February 2020, facial recognition startup Clearview AI discovered that a misconfigured server had left its internal documents exposed to the internet – and among them was the company’s entire client list. The leak landed days after a major newspaper investigation had revealed that the company had scraped more than three billion facial images from social networks and the open web, without consent, to sell a face-search engine to law enforcement and, it turned out, many others. The exposed list named retailers, banks, investors, and government entities – customers who had quietly sampled a tool most people had never heard of and never agreed to feed. Within days, platforms fired off cease-and-desist letters, European regulators opened complaints, and the most aggressive company in facial recognition lost control of its own story. This is how three weeks in early 2020 reshaped the debate on face surveillance.

Quick Answer: Clearview AI built a face-search engine from more than three billion images scraped from social media and websites, marketing it to police and private firms. In late February 2020, a misconfigured server exposed the company’s client list and internal documents, revealing users including major retailers, banks and investors. The exposure followed a January 2020 investigation into the scraping and preceded cease-and-desist demands from Facebook and other platforms plus a wave of privacy complaints across multiple countries.

The sequence compressed three revelations into one news cycle. First, the scraping: reporting established that Clearview had collected images from Facebook, YouTube, LinkedIn, Instagram, Venmo and countless public web pages – mass collection that no single person had authorized and no platform had licensed. Second, the clientele: the company pitched itself as a law enforcement tool, but its sales footprint reached into retail loss prevention, banking security, and the novelty curiosity of wealthy individuals associated with investors. Third, the exposure: the server misconfiguration put client names, account counts, and search volumes into public view – the company that collected everyone’s faces could not secure its own configuration. Each revelation amplified the others, and the combination moved the story from tech press to global policy discussion within days.

The scraped three billion

The scale was the shock. Not three thousand faces, not three million – three billion and growing, harvested by automated crawlers treating the public web and social platforms as an all-you-can-take biometric buffet. Legally, Clearview argued that public availability equals free use; critics, including the platforms themselves, answered that platform terms prohibit mass collection and that individuals never consented to inclusion in a searchable biometric index. The distinction between a human viewing a public photo and a company indexing every face on earth for permanent retrospective search is the entire policy argument in one sentence. Clearview’s technology meant that any photo of anyone – protest, coffee shop, ex-partner’s post – could become a query returning identity, collected without their knowledge. Democracies had debated face surveillance hypothetically for years; February 2020 made it operationally real and commercially distributed.

data-hmmnm-seam="2">

The paper trail

Date Event
2020-01-18 Major newspaper investigation reveals Clearview AI’s scraping of over three billion facial images and marketing to law enforcement
2020-02-05 Server misconfiguration exposes Clearview’s internal client list; discovery by a researcher begins the disclosure process
2020-02-27/28 Reporting documents the client list contents: retailers, banks, investors, and government users among tens of thousands of accounts
2020-02-27 Facebook sends a cease-and-desist letter demanding Clearview stop scraping and delete collected data; other platforms follow with demands
2020-03-05 Privacy complaints filed in France, Germany, and other jurisdictions; Canada and Australia open joint investigation that year
2020-10 Canada’s privacy commissioner rules Clearview’s scraping illegal under Canadian law, in the first major regulator verdict
data-hmmnm-seam="3">

The client list as confession

What made the leak detonate was specificity. An abstract startup scraping faces is a policy story; a list naming Macy’s, Walmart, Kohl’s, and Bank of America as users is a front page. Volume mattered too – tens of thousands of accounts had run hundreds of thousands of searches, which redefined the tool from pilot experiment to habitual infrastructure. The list documented a federated surveillance supply chain: a single company amassing the database, thousands of institutions querying it, zero subjects consenting. For the companies named, the exposure forced public statements distinguishing trials from contracts, terminated pilots from active use – corporate reputational triage none had prepared for. For privacy advocates, the list was discovery material: proof that face search had already distributed far beyond policing into commerce before any democratic decision had authorized it.

data-hmmnm-seam="4">

Platforms respond with paper

Facebook’s cease-and-desist letter demanded Clearview stop collecting images from its services and delete what it had taken – a demand with limited immediate teeth but significant legal positioning value. Google’s YouTube, LinkedIn, and Venmo sent similar notices in the following days. The platforms’ posture drew its own criticism: years of permissive API access and business partnerships had built the surveillance economy being condemned, and enforcement against a single scraper left the business model untouched. But the letters mattered. They established, on the record, that mass scraping violated terms of service, giving regulators a peg for unfair-competition and computer-fraud theories, and they signaled to the venture ecosystem that face-scraping carried platform risk. Clearview responded as it would throughout: characterizing its collection as lawful public-data use and continuing operations pending litigation – a stance multiple regulators would later reject with orders and fines.

data-hmmnm-seam="5">

The transatlantic regulatory cascade

Europe reacted with the machinery of GDPR. Complaints filed in early March 2020 by privacy organizations in France and Germany argued that biometric data processing requires explicit consent and that scraping billions of faces violated lawful-basis requirements outright. The cascade continued: UK information commissioner scrutiny, an Australian-Canadian joint investigation announced within months, and national orders to delete data. The October 2020 Canadian federal ruling that Clearview’s collection violated Canadian law set the template other regulators echoed with multimillion-euro fines in the years following (France’s CNIL among the most persistent, issuing escalating penalties for non-compliance). The deterministic thread through every proceeding: biometric data is special-category material, consent cannot be manufactured by public availability, and retrospective face search is mass surveillance regardless of the sales pitch.

  • Misconfigurations cut both ways: the company indexing everyone’s faces exposed its own secrets through a cloud misconfiguration; security posture is measured at the weakest bucket.
  • Public is not consent: regulators across jurisdictions converged on the principle that publicly accessible biometric data still requires a lawful basis for processing.
  • Client lists shape policy: naming commercial users converted an abstract rights debate into concrete institutional accountability, accelerating regulation by years.
  • Platform terms are leverage: cease-and-desist letters built the legal record regulators and litigants needed; paper battles preceded and enabled the enforcement wave.

FAQ

What was exposed in the Clearview AI leak?

A misconfigured server left Clearview’s internal documents accessible, including its client list, account information, and search-volume counts. The list revealed tens of thousands of accounts spanning law enforcement plus major retailers, banks, and entities associated with investors – exposure first surfaced by a researcher and confirmed in late February 2020 reporting.

How did Clearview AI get three billion face images?

Automated scraping of publicly accessible web pages and social media services, including major platforms, without user consent or platform licenses. Clearview maintained the collection was lawful use of public data; platforms and, later, regulators disagreed, with multiple authorities ordering cessation and deletion.

The exposed client list documented law enforcement agencies across countries as the core market, alongside commercial users including major retail chains and banks sampling loss-prevention applications, plus accounts tied to investors and founders’ networks. The breadth – not merely policing – was the leak’s central revelation.

What did Facebook and other platforms do?

They sent cease-and-desist letters in late February and early March 2020 demanding Clearview stop scraping their services and delete previously collected images. Clearview contested the demands, and the dispute migrated into litigation and regulatory proceedings that continued for years, with regulators ultimately imposing fines and deletion orders in several jurisdictions.

What was the regulatory outcome?

Complaints and investigations multiplied through 2020: Canada’s privacy commissioner ruled the scraping illegal in October 2020 with Australia joining the findings, European authorities including France’s CNIL issued orders and escalating fines for continued processing, and the UK opened its own actions. The collective precedent: mass biometric scraping lacks lawful basis under modern privacy regimes.

Clearview’s February 2020 collapsed the distance between surveillance theory and deployed reality. Before it, face recognition was a debate about government cameras and hypothetical databases; after it, the database existed, was searchable, had thousands of institutional users, and had been built from everyone’s shared online lives without asking. The consent gap – the space between what is technically collectable and what is democratically authorized – became the defining question of the biometric decade that followed, powering Illinois’ BIPA litigation wave that would later pin Clearview with a settlement covering millions of Americans, the EU AI Act’s subsequent restrictions on retrospective face search, and deletion orders across continents. For security practitioners, the case taught configuration humility alongside policy: the exposure that detonated the story was mundane, a bucket left open. For societies, it taught that privacy law either evolves to meet scraped-scale biometrics or becomes decorative. The company that bet the farm on “public means permitted” spent the following years learning, jurisdiction by jurisdiction, that consent is the one credential that cannot be scraped.

data-hmmnm-seam="end">

Prabhu Kalyan Samal

Application Security Consultant at TCS. Certifications: CompTIA SecurityX, Burp Suite Certified Practitioner, Azure Security Engineer, Azure AI Engineer, Certified Red Team Operator, eWPTX v3, LPT, CompTIA PenTest+, Professional Cloud Security Engineer, SC-900, SC-200, PSPO I, CEH, Oracle Java SE 8, ISP, Six Sigma Green Belt, DELF, AutoCAD. Writing about ethical hacking, security tutorials, and tech education at Hmmnm.