Arup’s HK$200M Deepfake Call: The CFO Fraud Manual Rewritten

📋 Key Takeaways
  • What happened?
  • The paper trail
  • Why the video call beat the email
  • The tooling collapse
  • What Arup changed
9 min read · 1,735 words
Educational & Ethical Use Only — This article is provided for educational and ethical cybersecurity research purposes only. The techniques described should only be used on systems you own or have explicit permission to test. Always follow responsible disclosure and the laws applicable to you. Mitigations are included so engineers can harden real systems.

What happened?

In February 2024, a finance employee at the Hong Kong office of Arup – the British engineering firm whose projects span skylines on six continents – joined a video conference call with the company’s UK chief financial officer and four other colleagues. Everything about the call looked right: the faces, the voices, the org-chart logic of who was speaking. Everything about it was manufactured. The employee was the only human on the line; the other participants were deepfakes assembled by fraudsters who had spent the prior week conditioning the target through phishing messages posing as the UK CFO. After the call, the employee executed the requested transfer – fifteen transactions totaling roughly HK$200 million, about US$25.6 million – into local accounts controlled by the gang. When this post publishes on 25 November 2024, the case has spent a year as the world’s reference standard for deepfake-enabled fraud: the largest publicly documented, and a preview of a threat class that AI video and voice cloning made available to any criminal with a laptop.

Quick Answer: In February 2024, an employee at Arup’s Hong Kong branch was deceived into paying out approximately HK$200 million (about US$25.6 million) in a multi-stage fraud whose centerpiece was a deepfake video conference. The scam began with phishing messages impersonating the UK-based CFO, then escalated to a group video call in which the CFO and several colleagues appeared – all of them AI-generated fakes, per Hong Kong police, built from publicly available footage of the real executives. Believing the instructions were legitimate, the employee made 15 transfers to five local bank accounts controlled by the fraudsters. The deception came to light only when the employee followed up with the corporation’s headquarters afterward. Hong Kong police detailed the case publicly on 4 February 2024 at a briefing that made international headlines; no arrests were announced at the time, the funds were largely dispersed through the territory’s financial plumbing, and the case became the canonical example of deepfake-enabled CEO/CFO fraud – demonstrating that “verify the person” fails as a control when the person, the face, and the voice can all be synthesized on demand.

The engineering of the scam deserves cold study because its stages generalize. Stage one was conditioning: phishing messages established the pretext – a confidential transaction, the authority of the UK CFO – so that the video call arrived as confirmation rather than surprise. Stage two was the seal: a group call with multiple recognizable colleagues, which flips the victim’s default from skepticism to deference, because multi-party social proof is how humans sanity-check unusual requests. Stage three was urgency and secrecy, the standard fraud stack that suppresses the impulse to double-check through a side channel. The deepfake layer did not invent any of this; it patched the classic Business Email Compromise formula’s one weakness – the meeting request that would previously have exposed the impersonation. With synthetic faces and voices, the richest verification channel available to a remote workforce became just another forgery surface.

The paper trail

Date Event
2024-01 mid→late Phishing messages impersonating Arup’s UK CFO target the Hong Kong finance employee, building the confidential-transaction pretext
2024-01→02 The group video conference: the employee joins a call with the “CFO” and four colleagues – all deepfakes per police; only the victim is real
2024-01→02 The group video conference occurs
2024-01-15→20 Fifteen transfers totaling roughly HK$200 million (about US$25.6 million) go to five local Hong Kong bank accounts controlled by the gang
2024-02-04 Hong Kong police brief the case publicly; deepfake fraud is now a named, quantified category in the territory caseload
2024-02→11 Global coverage cements the Arup case as the reference incident for synthetic-media financial fraud; banks and enterprises re-examine verification protocols
2024-11-25 This post publishes with deepfake fraud attempts rising across the industry and the case still the largest publicly documented
data-hmmnm-seam="2">

Why the video call beat the email

BEC fraud is a mature industry with known immune responses: finance teams learned to treat email payment instructions as unverified by default, callback procedures spread, and the “out-of-band verification” checkbox became standard. The Arup case worked because it targeted the upgraded protocol itself. A video call is treated as the strong verification – the thing you escalate to when email seems risky – so the fraudsters supplied one, complete with multiple faces the victim recognized and voices calibrated to match. The psychological payload of multiple apparent colleagues cannot be overstated: humans resolve uncertainty socially, and a conference full of confident, familiar people telling a finance professional that the transfer is expected is a stronger compliance signal than any single authority figure. The industry’s uncomfortable post-mortem conclusion was that the verification channel matters less than its independence – any channel the attacker can fully synthesize, including face and voice, must not double as the authentication mechanism.

data-hmmnm-seam="3">

The tooling collapse

What made 2024 the inflection year was tooling economics. High-quality face-swap and voice-clone services crossed from research demos to commodity subscriptions – tens of dollars a month, minutes per clone, requiring only the publicly available conference footage, earnings-call appearances, and podcast interviews that every senior executive accumulates. The Arup gang needed thirty seconds of a CFO’s voice from any public source to pass a phone check, and a handful of photographs to animate a convincing call presence. Defensive technology scrambled to keep pace – liveness detection, media-forensics tooling, C2PA-style content provenance – but the honest assessment of the year was that detection is a losing race at the point of interaction, and that process redesign (segregated approval paths, transaction limits, callback to known numbers) buys more protection per dollar than forensic AI. The fraud crews read the same product announcements as everyone else; by late 2024, deepfake job interviews, deepfake KYC onboarding, and voice-clone family emergencies were all appearing in caseloads, each a variation on the Arup architecture: synthesize the trust anchor, then request the transfer.

data-hmmnm-seam="4">

What Arup changed

Arup itself handled the aftermath with notable transparency for a private firm: it confirmed the incident, quantified the loss tolerance without drama, and treated the case as an industry warning rather than a reputation shield. Inside the firm and across the engineering and professional-services sector, the response pattern was procedural: dual-control payment verification that no single meeting can satisfy, out-of-band callbacks to pre-registered numbers for any transfer above threshold, and – the hard one – training that explicitly revokes the video call’s status as verification. The last is culturally difficult, because it asks organizations to institutionalize distrust of their own executives’ faces. The ones that managed it reframed the rule positively: the more seniors the request involves, the more automated and impersonal its verification must be – precisely because senior presence is now the cheapest thing to fake and the most psychologically potent thing to deploy.

  • Synthesizable channels cannot authenticate: face and voice are now forgeable at commodity prices; any verification protocol that terminates in a video call or voice match is a protocol with a bypass.
  • Multi-party presence is the payload: the group-call format weaponizes social proof; the more familiar faces on the screen, the stronger the engineered compliance signal.
  • Conditioning precedes the con: the phishing phase built the pretext the call then “confirmed” – treat any urgent confidential financial narrative introduced by message as pre-fraud grooming.
  • Process beats forensics at the point of transfer: dual control, thresholds, and pre-registered callback channels stopped this fraud class where real-time deepfake detection could not.

FAQ

How were the deepfakes actually made?

Per Hong Kong police descriptions, the fraudsters combined publicly available footage and imagery of the executives with commodity face-swap and voice-clone tooling to populate the live call. No breach of Arup’s systems was reported – the raw material for executive deepfakes is public media: earnings calls, conference talks, corporate videos, press photographs. That public-source dependency defines both the threat’s reach (every organization’s leadership is instantiable) and its limit: quality varies with available footage, which is why the call was engineered to be brief, authoritative, and socially dense rather than long and conversational.

Was the money recovered?

Only fragments, in the immediate aftermath. The fifteen transfers moved through five local accounts and were dispersed quickly into the territory’s financial system – the standard money-laundering pattern of layered hops designed to outrun freeze orders. Hong Kong police said at the briefing that some sums were recovered and the investigation continued, but the overwhelming majority of the HK$200 million was gone within days. The recovery arithmetic is the quiet driver of prevention economics: once a deepfake-assisted transfer clears, the money is structurally closer to gone than any forensic breakthrough is to fast.

How should a finance team train for this now?

Three controls, in order of effectiveness. First, decouple verification from identity performance: payments above threshold require approval through a system – not a meeting – with dual control and a pre-registered callback to a number on file, independent of whatever contact details the request itself supplied. Second, rehearse the narrative, not just the media: the scam’s signature was urgency plus confidentiality plus authority, so training drills exactly that stack arriving by message, call, or video. Third, make executives harder to clone and easier to verify: minimize high-quality public voice/face corpus where feasible, agree family and organisational code words for sensitive requests, and treat any request to bypass a control “because I am asking personally” as the red flag it has always been – now with a synthetic face attached.

data-hmmnm-seam="5">

Legacy: the year trust got a render farm

The Arup case closed 2024 as the deepfake threat’s reference incident – the one regulators cited, vendors demoed against, and board PowerPoint decks led with. Its lasting contribution was a reordering of intuition: the most secure-feeling channel in remote work, the live video call with known colleagues, became the most convincingly forgeable. Every control redesign that followed shared one principle – authentication must live in channels attackers cannot render, whether that is hardware keys, pre-registered numbers, or workflow systems that do not care how senior the face on the screen looks. The victim in Hong Kong was not careless by the standards of the era before; they followed the escalation protocol of a modern finance function and were defeated by a technology that counterfeits its assurances. That is the Arup legacy: proof that the cost of faking trust collapsed to commodity prices in 2024, and that the only durable defense is to stop asking faces – real or rendered – to vouch for money.

data-hmmnm-seam="end">

Prabhu Kalyan Samal

Application Security Consultant at TCS. Certifications: CompTIA SecurityX, Burp Suite Certified Practitioner, Azure Security Engineer, Azure AI Engineer, Certified Red Team Operator, eWPTX v3, LPT, CompTIA PenTest+, Professional Cloud Security Engineer, SC-900, SC-200, PSPO I, CEH, Oracle Java SE 8, ISP, Six Sigma Green Belt, DELF, AutoCAD. Writing about ethical hacking, security tutorials, and tech education at Hmmnm.