Tool Hijacking: The 2024 Papers That Predicted Agent Attacks
By November 2024, AI-agent security research had already documented the attack class that production incidents would later make infamous. InjecAgent (March 2024, ACL Findings) benchmarked 1,054 indirect-injection scenarios across 30 agents, finding ReAct-prompted GPT-4 attacked successfully roughly a quarter of the time. Breaking Agents (July 2024) demonstrated malfunction amplification through agentic loops. Together with 2023's foundational indirect-prompt-injection work, they mapped how tools, descriptions, and fetched content become command channels. This survey walks the papers, the hijack taxonomy, and the controls that predate the incidents.
