Back to blogs

September 28, 2026

You Approved $20, It Ran $2,000: What Loopjacking Means for Every 'Human-in-the-Loop' AI Guardrail

AI agent securityenterprise AI governancehuman-in-the-loopAI agentsagentic AIAI risk management
You Approved $20, It Ran $2,000: What Loopjacking Means for Every 'Human-in-the-Loop' AI Guardrail

Enterprises rolling out AI agents have converged on one safety net above all others: put a human in the loop before anything consequential happens. A wire transfer, a database deletion, a production deploy — the agent proposes, a person reviews, and only then does the action fire. It’s the control that lets risk and compliance teams sleep at night, and it underpins how Promact has argued businesses should keep human control over AI.

New research published this month says that safety net has a hole in it — and the hole isn’t theoretical. In a paper titled “Loopjacking: Hijacking Human-in-the-Loop Approval”, security researcher Adithyan Arun Kumar demonstrated that a human can approve exactly what they see on screen, while the system executes something else entirely. Not a hypothetical edge case — a reproducible exploit, tested and confirmed across two of the most widely deployed open-source agent frameworks.

What Loopjacking Actually Is

The paper’s core claim is simple to state and unsettling to absorb: human approval is only a meaningful security boundary if the operation shown to the reviewer is the same operation the system later executes. Loopjacking is the name for what happens when that binding breaks — when a human approves what they understand to be operation A, but the implementation carries out a materially different operation B.

The research identifies two distinct ways this binding fails, as detailed in the full paper on arXiv:

Representation-based attacks

Here, operation B already exists in the system at approval time — it’s just not shown to the reviewer, or it’s misrepresented. The human is asked to approve something that looks benign, without visibility into what will actually run.

Post-approval state substitution

This is the more alarming variant. The human sees the correct operation A, reviews it accurately, and approves it in good faith. But between that approval and execution, mutable workflow state is swapped out, and operation B runs instead — under the authority of the approval the human already gave.

The $20 Approval That Became $2,000

The proof-of-concept that made this research spread fast is deliberately simple. A test scenario staged a pending call to mock_wire_transfer(20, approved-vendor) — a $20 transfer to a known, approved vendor. The human approver reviewed exactly that request and signed off on it.

Before execution, a separate message arrived through the agent-to-agent (A2A) communication channel and replaced the pending call with mock_wire_transfer(2000, attacker-sink). The system dispatched the substituted call. The mock ledger recorded a 2,000-unit transfer to an attacker-controlled destination — authorized, on paper, by a human who had approved a $20 transfer to someone else entirely. The researcher’s own write-up of the A2A attack walks through the exact sequence, and a public evidence archive on GitHub preserves the requests, decisions, and exact software versions tested, so the finding isn’t a one-off claim — it’s independently reproducible.

Which Systems Were Affected

This wasn’t a single misconfigured demo. The researcher reproduced post-approval state substitution across seven separate releases of Agno AgentOS, running through version 3.0.9, and across twelve tested versions of a LangGraph Agent Server composition, through version 0.14.0. A representation-mismatch variant was separately reproduced in OpenClaw version 2026.2.23. Coverage from The Hacker News-adjacent security press and the Agentic Security newsletter treated the breadth of affected frameworks as the real story: this is a class of vulnerability in how “approval” gets wired into agent architectures, not a bug in one vendor’s code.

Notably, the research also ran a positive control: OpenAI’s Agents SDK (versions 0.22.0 and 0.22.2) resisted the attack. Its approach serializes the exact per-call action into run state at the moment of approval, alongside the approval metadata itself. Any attempt to mutate the pending action after that point is rejected before dispatch. That design detail matters more than it sounds — it’s the difference between an approval that binds to a specific, immutable action and one that merely binds to a request that can still change underneath it.

The Fix Is Already Shipping

To its credit, the Agno team moved fast. Maintainers merged nine authorization guard fixes in a single pull request shortly after disclosure, each fix reproduced on the vulnerable branch head, corrected, and pinned by a regression test. A follow-up PR closed additional high-severity authorization findings uncovered during the same audit. The pattern the fixes converge on is exactly what the paper prescribes: complete, canonical rendering of what’s being approved, an exact comparison against that canonical version at the moment of use, and hard prevention of unauthorized mutation to pending state in between. Teams running Agno AgentOS in production should confirm they’ve pulled the patched release before treating any approval workflow as trustworthy.

Why This Matters Beyond These Two Frameworks

It would be a mistake to read Loopjacking as “patch these two libraries and move on.” The finding cuts at something Promact has flagged before: action-layer governance, not prompt-layer controls, is where enterprise AI agent safety actually has to live. An approval dialog is a UI-layer control. If the system underneath doesn’t cryptographically or structurally bind that approval to one specific, immutable action, the dialog is theater — a compliance checkbox rather than a security boundary.

This also complicates the case for bounded-autonomy architectures that lean on human checkpoints as their primary safety mechanism, the kind of design Promact covered in Proofpoint’s bounded-autonomy bet on enterprise AI security. Bounding what an agent can do without sign-off is still sound design. But if the sign-off itself can be silently redirected, the bound is only as strong as the plumbing connecting approval to execution — and this research shows that plumbing has failed in production-grade, widely deployed frameworks.

Zoom out further and this fits a pattern Promact has tracked all year: as enterprises lose track of the sheer number of agents they’re running, the controls meant to keep any single agent in check are getting less scrutiny than the agents themselves. Loopjacking is a reminder that a governance program built entirely around “a human approves it” is only as strong as its weakest binding layer — and most teams have never audited that layer at all.

What Enterprise Teams Should Do Now

Security teams running agent frameworks with human-approval gates have concrete, immediate steps available:

  1. Inventory every approval gate. For each one, ask: is the exact action serialized and locked at approval time, or is state still mutable afterward?
  2. Check framework versions against the disclosure. If you’re running Agno AgentOS or a LangGraph Agent Server composition, confirm you’re on a patched release, not just a “recent” one.
  3. Demand canonical rendering from vendors. Any tool that presents an approval UI should be able to state, precisely, how it guarantees the approved action and the executed action are the same object — not just the same-looking request.
  4. Treat approval binding as a first-class security control, subject to the same audit rigor as authentication or encryption — not a UX feature bolted on to make agents feel safer.

The uncomfortable truth in this research is that “human-in-the-loop” has become a phrase enterprises use to feel safe, more than a technical guarantee they’ve verified. Loopjacking shows the gap between those two things can be worth 100x the approved amount — and it can be exploited without the human ever knowing they were wrong.

Frequently Asked Questions

What is Loopjacking, in plain terms?

Loopjacking is a class of security failure in AI agent systems where a human approves one action, but a different, more consequential action executes instead — because the system never cryptographically or structurally locked the approval to a single, unchangeable operation.

Which AI agent frameworks were shown to be vulnerable?

Researchers reproduced the attack in seven releases of Agno AgentOS (through version 3.0.9) and twelve versions of a LangGraph Agent Server composition (through version 0.14.0), plus a related representation-mismatch variant in OpenClaw version 2026.2.23.

Has this been fixed?

Agno’s maintainers merged nine authorization guard fixes addressing the issue shortly after disclosure, with a follow-up patch closing related high-severity findings. Teams should confirm they are running the patched release rather than assuming an update has already reached their deployment.

Is my organization at risk if we use human-in-the-loop approval for AI agents?

Only if your approval mechanism doesn’t bind the exact, serialized action to the approval decision at the moment of review. OpenAI’s Agents SDK was tested as resistant because it locks the specific action into run state before dispatch — that’s the design pattern to look for or demand from vendors.

Does this mean human-in-the-loop controls are pointless?

No — it means the control has to be implemented correctly at the system level, not just presented correctly at the UI level. An approval step is only as trustworthy as the guarantee that what was shown is what gets executed.

How is this different from prompt injection?

Prompt injection manipulates what an agent decides to do by corrupting its inputs. Loopjacking operates after a decision has already been made and approved — it manipulates what actually executes, independent of the reasoning that led to the approval.

Sources

Have a project like this in mind?

Tell us what you're building — we'll help you scope it and ship it.

Talk to us

Keep reading

Promact team

We are a family of Promactians

We are an excellence-driven company passionate about technology where people love what they do.

Get opportunities to co-create, connect and celebrate!

Join Us

Vadodara

Headquarter

B-301, Monalisa Business Center, Manjalpur, Vadodara, Gujarat, India - 390011

+91 (932)-703-1275

Pune

46 Downtown, 805+806, Pashan-Sus Link Road, Near Audi Showroom, Baner, Pune, Maharashtra, India - 411045

USA

4056, 1207 Delaware Ave, Wilmington, DE, United States America, US, 19806

+1 (765)-305-4030
Promact global office locations on world map