September 28, 2026
You Approved $20, It Ran $2,000: What Loopjacking Means for Every 'Human-in-the-Loop' AI Guardrail
Enterprises rolling out AI agents have converged on one safety net above all others: put a human in the loop before anything consequential happens. A wire transfer, a database deletion, a production deploy — the agent proposes, a person reviews, and only then does the action fire. It’s the control that lets risk and compliance teams sleep at night, and it underpins how Promact has argued businesses should keep human control over AI.
New research published this month says that safety net has a hole in it — and the hole isn’t theoretical. In a paper titled “Loopjacking: Hijacking Human-in-the-Loop Approval”, security researcher Adithyan Arun Kumar demonstrated that a human can approve exactly what they see on screen, while the system executes something else entirely. Not a hypothetical edge case — a reproducible exploit, tested and confirmed across two of the most widely deployed open-source agent frameworks.
What Loopjacking Actually Is
The paper’s core claim is simple to state and unsettling to absorb: human approval is only a meaningful security boundary if the operation shown to the reviewer is the same operation the system later executes. Loopjacking is the name for what happens when that binding breaks — when a human approves what they understand to be operation A, but the implementation carries out a materially different operation B.
The research identifies two distinct ways this binding fails, as detailed in the full paper on arXiv:
Representation-based attacks
Here, operation B already exists in the system at approval time — it’s just not shown to the reviewer, or it’s misrepresented. The human is asked to approve something that looks benign, without visibility into what will actually run.
Post-approval state substitution
This is the more alarming variant. The human sees the correct operation A, reviews it accurately, and approves it in good faith. But between that approval and execution, mutable workflow state is swapped out, and operation B runs instead — under the authority of the approval the human already gave.
The $20 Approval That Became $2,000
The proof-of-concept that made this research spread fast is deliberately simple. A test scenario staged a pending call to mock_wire_transfer(20, approved-vendor) — a $20 transfer to a known, approved vendor. The human approver reviewed exactly that request and signed off on it.
Before execution, a separate message arrived through the agent-to-agent (A2A) communication channel and replaced the pending call with mock_wire_transfer(2000, attacker-sink). The system dispatched the substituted call. The mock ledger recorded a 2,000-unit transfer to an attacker-controlled destination — authorized, on paper, by a human who had approved a $20 transfer to someone else entirely. The researcher’s own write-up of the A2A attack walks through the exact sequence, and a public evidence archive on GitHub preserves the requests, decisions, and exact software versions tested, so the finding isn’t a one-off claim — it’s independently reproducible.
Which Systems Were Affected
This wasn’t a single misconfigured demo. The researcher reproduced post-approval state substitution across seven separate releases of Agno AgentOS, running through version 3.0.9, and across twelve tested versions of a LangGraph Agent Server composition, through version 0.14.0. A representation-mismatch variant was separately reproduced in OpenClaw version 2026.2.23. Coverage from The Hacker News-adjacent security press and the Agentic Security newsletter treated the breadth of affected frameworks as the real story: this is a class of vulnerability in how “approval” gets wired into agent architectures, not a bug in one vendor’s code.
Notably, the research also ran a positive control: OpenAI’s Agents SDK (versions 0.22.0 and 0.22.2) resisted the attack. Its approach serializes the exact per-call action into run state at the moment of approval, alongside the approval metadata itself. Any attempt to mutate the pending action after that point is rejected before dispatch. That design detail matters more than it sounds — it’s the difference between an approval that binds to a specific, immutable action and one that merely binds to a request that can still change underneath it.
The Fix Is Already Shipping
To its credit, the Agno team moved fast. Maintainers merged nine authorization guard fixes in a single pull request shortly after disclosure, each fix reproduced on the vulnerable branch head, corrected, and pinned by a regression test. A follow-up PR closed additional high-severity authorization findings uncovered during the same audit. The pattern the fixes converge on is exactly what the paper prescribes: complete, canonical rendering of what’s being approved, an exact comparison against that canonical version at the moment of use, and hard prevention of unauthorized mutation to pending state in between. Teams running Agno AgentOS in production should confirm they’ve pulled the patched release before treating any approval workflow as trustworthy.
Why This Matters Beyond These Two Frameworks
It would be a mistake to read Loopjacking as “patch these two libraries and move on.” The finding cuts at something Promact has flagged before: action-layer governance, not prompt-layer controls, is where enterprise AI agent safety actually has to live. An approval dialog is a UI-layer control. If the system underneath doesn’t cryptographically or structurally bind that approval to one specific, immutable action, the dialog is theater — a compliance checkbox rather than a security boundary.
This also complicates the case for bounded-autonomy architectures that lean on human checkpoints as their primary safety mechanism, the kind of design Promact covered in Proofpoint’s bounded-autonomy bet on enterprise AI security. Bounding what an agent can do without sign-off is still sound design. But if the sign-off itself can be silently redirected, the bound is only as strong as the plumbing connecting approval to execution — and this research shows that plumbing has failed in production-grade, widely deployed frameworks.
Zoom out further and this fits a pattern Promact has tracked all year: as enterprises lose track of the sheer number of agents they’re running, the controls meant to keep any single agent in check are getting less scrutiny than the agents themselves. Loopjacking is a reminder that a governance program built entirely around “a human approves it” is only as strong as its weakest binding layer — and most teams have never audited that layer at all.
What Enterprise Teams Should Do Now
Security teams running agent frameworks with human-approval gates have concrete, immediate steps available:
- Inventory every approval gate. For each one, ask: is the exact action serialized and locked at approval time, or is state still mutable afterward?
- Check framework versions against the disclosure. If you’re running Agno AgentOS or a LangGraph Agent Server composition, confirm you’re on a patched release, not just a “recent” one.
- Demand canonical rendering from vendors. Any tool that presents an approval UI should be able to state, precisely, how it guarantees the approved action and the executed action are the same object — not just the same-looking request.
- Treat approval binding as a first-class security control, subject to the same audit rigor as authentication or encryption — not a UX feature bolted on to make agents feel safer.
The uncomfortable truth in this research is that “human-in-the-loop” has become a phrase enterprises use to feel safe, more than a technical guarantee they’ve verified. Loopjacking shows the gap between those two things can be worth 100x the approved amount — and it can be exploited without the human ever knowing they were wrong.
Frequently Asked Questions
What is Loopjacking, in plain terms?
Loopjacking is a class of security failure in AI agent systems where a human approves one action, but a different, more consequential action executes instead — because the system never cryptographically or structurally locked the approval to a single, unchangeable operation.
Which AI agent frameworks were shown to be vulnerable?
Researchers reproduced the attack in seven releases of Agno AgentOS (through version 3.0.9) and twelve versions of a LangGraph Agent Server composition (through version 0.14.0), plus a related representation-mismatch variant in OpenClaw version 2026.2.23.
Has this been fixed?
Agno’s maintainers merged nine authorization guard fixes addressing the issue shortly after disclosure, with a follow-up patch closing related high-severity findings. Teams should confirm they are running the patched release rather than assuming an update has already reached their deployment.
Is my organization at risk if we use human-in-the-loop approval for AI agents?
Only if your approval mechanism doesn’t bind the exact, serialized action to the approval decision at the moment of review. OpenAI’s Agents SDK was tested as resistant because it locks the specific action into run state before dispatch — that’s the design pattern to look for or demand from vendors.
Does this mean human-in-the-loop controls are pointless?
No — it means the control has to be implemented correctly at the system level, not just presented correctly at the UI level. An approval step is only as trustworthy as the guarantee that what was shown is what gets executed.
How is this different from prompt injection?
Prompt injection manipulates what an agent decides to do by corrupting its inputs. Loopjacking operates after a decision has already been made and approved — it manipulates what actually executes, independent of the reasoning that led to the approval.
Sources
- Loopjacking: Hijacking Human-in-the-Loop Approval (arXiv:2609.21081) - The original research paper describing the vulnerability class and attack variants.
- Full paper HTML version - Detailed methodology and reproduction data.
- Loopjacking in A2A Implementations (researcher’s own write-up) - First-person account of the $20-to-$2,000 wire transfer proof-of-concept.
- Evidence and reproduction archive on GitHub - Public archive preserving requests, decisions, and exact tested software versions.
- Agno AgentOS PR #10270: authorization guard fixes - The maintainers’ fix, with nine guard-style corrections from the audit.
- Loopjacking discussion on Hacker News - Community and practitioner reaction to the disclosure.
- The Agentic Security Newsletter, Week of September 21, 2026 - Industry security-newsletter coverage placing the finding in context.
- A $20 Approval Ran as $2,000 in A2A Test - StartupHub.ai - News coverage summarizing the proof-of-concept.
Have a project like this in mind?
Tell us what you're building — we'll help you scope it and ship it.
Talk to usKeep reading

September 27, 2026
The Agents Nobody Counted: What Dataiku's New Management Platform Means for Enterprise AI Governance

September 25, 2026
Jev's Record Launch: What TypeSafe AI's 'System One' Judgment Model Means for Your AI Automation Bill

September 25, 2026