Tool Hijacking: The 2024 Papers That Predicted Agent Attacks

By November 2024, AI-agent security research had already documented the attack class that production incidents would later make infamous. InjecAgent (March 2024, ACL Findings) benchmarked 1,054 indirect-injection scenarios across 30 agents, finding ReAct-prompted GPT-4 attacked successfully roughly a quarter of the time. Breaking Agents (July 2024) demonstrated malfunction amplification through agentic loops. Together with 2023’s foundational indirect-prompt-injection work, they mapped how tools, descriptions, and fetched content become command channels. This survey walks the papers, the hijack taxonomy, and the controls that predate the incidents.

Continue ReadingTool Hijacking: The 2024 Papers That Predicted Agent Attacks

Kadrey v. Meta: Piracy Allegations and the Llama Paper Trail

In mid-December 2024, unsealed filings in Kadrey et al. v. Meta Platforms alleged the company torrented LibGen’s pirated library while engineers warned it ‘doesn’t feel right’ on corporate laptops, stripped copyright management information with purpose-built scripts, and used a dataset a memo called ‘we know to be pirated’ — with CEO approval over executive objections. The proposed DMCA and CDAFA claims reframe the AI-copyright fight around distribution and concealment rather than fair use alone. This account walks the exhibits, the legal architecture, and the compliance lessons for every AI data program.

Continue ReadingKadrey v. Meta: Piracy Allegations and the Llama Paper Trail
Read more about the article Agentic AI Security: Attack Surface in Autonomous Systems
Agentic AI Security: Attack Surface in Autonomous Systems

Agentic AI Security: Attack Surface in Autonomous Systems

A practical guide to agentic AI security covering goal hijacking, tool misuse, identity and privilege abuse, memory poisoning, multi-agent trust issues, and defense frameworks for autonomous AI systems.

Continue ReadingAgentic AI Security: Attack Surface in Autonomous Systems

Google AI Overviews: Prompt Injection Hits the Homepage of the Internet

When Google rolled AI Overviews into US search in May 2024, satirical sources got quoted as fact at national scale — glue on pizza, rocks as vitamins — and security researchers reframed the comedy as indirect prompt injection: retrieved content steering the answer in Google’s own voice. This piece tracks the launch-week failures, the overview-bait SEO economy that followed, the manual-removal treadmill, provenance-aware retrieval as the real fix, and why RAG systems inherit the trust profile of their worst-cited source.

Continue ReadingGoogle AI Overviews: Prompt Injection Hits the Homepage of the Internet

KnowBe4 vs a Fake North Korean IT Worker: The AI-Era Insider Case Study

In July 2024, security-awareness firm KnowBe4 hired a remote principal software engineer who turned out to be a North Korean IT worker using an AI-groomed persona, a US PPPoE front, and a stolen identity. Detected within 32 minutes of suspicious activity and fully rigged with granular session logging, the case became the definitive inside look at DPRK pension applicantFraud — from laptop farms to paycheck revenue streams funding weapons programs. This piece reconstructs the fraud chain, the detection story, and the hiring controls that failed.

Continue ReadingKnowBe4 vs a Fake North Korean IT Worker: The AI-Era Insider Case Study

Ray AI Framework’s ‘Won’t Fix’ CVEs: A Control-Plane Debate

When Protect AI disclosed five Ray vulnerabilities in March 2024 — including critical RCE via the unauthenticated control plane — Anyscale’s ‘won’t fix, trusted-networks design’ stance ignited the year’s sharpest debate over AI infrastructure responsibility. This piece unpacks the job-submission RCE, the exposed-cluster census, the bounty economics, what Anyscale later shipped anyway, and the hardening playbook that became standard for every exposed ML control plane.

Continue ReadingRay AI Framework’s ‘Won’t Fix’ CVEs: A Control-Plane Debate

Windows Recall: The Privacy Debate Before Launch

Announced May 20, 2024 as a Copilot+ flagship, Windows Recall promised searchable memory of everything on screen — and researchers found the archive in a plaintext SQLite database any user-context malware could read, with a runtime API to match. This account covers Kevin Beaumont’s teardown, the TotalRecall extraction tool, the threat-model fallacies in each Microsoft defense, the June climb-down to opt-in plus Windows Hello and encryption, and the rare process win of an architecture changed before deployment.

Continue ReadingWindows Recall: The Privacy Debate Before Launch

ChatGPT’s November 2023 DDoS Outages, Explained

For much of 8 November 2023, ChatGPT and parts of OpenAI’s API cycled in and out of service under a denial-of-service wave claimed by Anonymous Sudan, with smaller recurrences through the month. OpenAI confirmed the DDoS, rolled global WAF rules, and absorbed a false-positive tax on legitimate users. Nothing was breached — the story is availability risk wrapped around AI dependence. This post walks the campaign’s anatomy, Microsoft’s Storm-1359 telemetry link, and the business-continuity lessons for anyone running on AI vendors.

Continue ReadingChatGPT’s November 2023 DDoS Outages, Explained