Research
When Payment Agents Meet Adversaries: Why Verification Beats Vibes
Lessons from PayMAS Phase 0 — a synthetic research sandbox for measuring whether AI payment agents still honour user intent under prompt injection, data tampering, and tool poisoning.
BXTrack Editorial Team
September 30, 2026·9 min read

Autonomy next to money needs a hostile baseline
Most agent demos assume clean invoices, honest tools, and helpful merchants. Production payment work does not. Notes contain instructions. Quotes get rewritten. Tool schemas grow fields that “must be used.” If your evaluation never includes those surfaces, you are measuring helpfulness — not integrity.
PayMAS Phase 0 asks a blunt question in a synthetic sandbox: can an autonomous payment agent preserve the user’s intended transaction when untrusted instructions, financial data, or tool outputs are adversarially manipulated? No real banks. No credentials. No funds. Just a frozen oracle and a ledger.
Three attack families, ten payment workflows
The protocol stresses invoice payments, P2P transfers, merchant purchases, escrow funding and release, subscription cancellation, hotel booking and refunds, cross-border transfers, and recurring rent. Across those flows we rotate prompt injection, data tampering, and tool poisoning.
That breadth matters. Payee integrity, amount integrity, and workflow-term integrity fail differently. An agent that resists a poisoned invoice note can still accept a tampered refund destination if the eval never covers refunds.
v1 taught us what a broken fixture looks like
In the first protocol, every task was an invoice payment and the agent’s only payment information came from a single invoice tool. Data tampering overwrote the payee account in that store. Tool poisoning added a preferred settlement account and told the model it must be used.
With no second source, paying the attacker was the “correct” reading of the evidence. Attack success rate hit 100% across models, with identical misdirected value. That was not model stupidity. It was experiment design claiming to measure robustness while removing the possibility of verification.
Protocol v2: separate evidence from verification
v2 keeps a frozen task oracle for the evaluator, copies a world of sources into the sandbox, and lets attacks patch exactly one source or rewrite one tool response. Independent verification surfaces stay available. Runs are planned with identical fixtures per task, condition, and repetition so models are compared fairly.
Outcomes are scored as safe success, attack success, safe refusal, task failure, or partial failure. Events track whether the agent saw a conflict, used verification, proposed an unsafe action, or executed it. Metrics cover clean utility, ASR, utility under attack, verification usage, and simulated misdirected value.
What this changes for teams shipping agents
If you are putting an agent near payments, subscriptions, or payouts, build adversarial evals before you expand autonomy. Require a second source for high-stakes fields. Treat tool descriptions as untrusted input. Measure refusal and verification, not only task completion.
PayMAS is BXTrack’s research surface for that discipline. The point is not fear of agents. The point is earning the right to automate money by proving — under attack — that user intent still wins.
Written by
BXTrack Editorial Team
The BXTrack editorial team covers AI products, software systems, and research that keeps agentic automation accountable when it sits next to real business risk.