Reliability and evidence#

This page reports what has been measured, including a payment that was lost, a mitigation that has not yet been validated, and an estimate that turned out to be wrong and was retracted.

A reliability page that reports only successes is indistinguishable from one that was never tested. This one is written so that a reader can tell the difference.

Counts increase on every CI run and every scheduled probe, so every figure here is a snapshot with a date attached. Re-run scripts/collect-probe-results.sh for current numbers.


What is proven#

A stock, unmodified x402 client completes real settled payments against Turnpike on stellar:testnet, in continuous integration, on infrastructure I do not control.

The client is installed from public npm at a pinned version with no patches, and the conformance harness imports nothing from this repository. If the harness passes, it passes for the same reason any third party’s client would.

Pinned versions, from conformance/package.json and verified against what is installed:

PackageVersion
@x402/fetch2.21.0
@x402/core2.21.0
@x402/stellar2.21.0
@stellar/stellar-sdk16.1.0

The CI-verified payment#

The strongest single piece of evidence, because it is the hardest to fake:

Transactionb983f668aae5de59c88e512c510f1e95ca6654780eba2a3c1b903359934cd0ea
Ledger4104581 — successful
Fee paid byGDUUUKVFFCCNMBGH7IYV42RQDSZS4J4KJC6WQQFIL76YE7WNSTALW66N

That fee-paying account did not exist before the CI run started. The runner created it via Friendbot, paid a resource server, and settled on testnet. There are no long-lived funded secrets in the workflow, which is why anyone can fork the repository and reproduce the result rather than taking my word for it.

Asset coverage#

AssetTransactionNote
Native XLM (SAC)f6c6fbcc19d2661d5b9a0d977f562a371411b5974c621e2b79688906bce31fd6Default demo path — ledger 4094092, fee paid by GCGPGN6N…NOGN, 20 554 stroops
Testnet USDC538e1d8355e772cac97a8e3720e0f94ea1201941cf4d06f16f369eb885bc8cd3SEP-41 path — ledger 4094044, buyer 20.00 → 19.99, seller 0.00 → 0.01

The default path uses native XLM so that a setup script can fund every account via Friendbot: clean clone, one command, real settled payment, no manual steps, and it survives testnet resets. The USDC transaction was a one-off manual run — testnet USDC requires a web faucet visit — executed through the same code path with only an environment variable changed. USDC is not the default and does not work without that faucet step.

Reproducibility#

docker compose up from a clean clone brings up the facilitator and a demo resource server. A preflight check validates Node version, the Docker daemon, and port availability, distinguishing “an x402 stack is already running here” from “something else holds this port,” and a harness failure prints the two likely causes — testnet reset, lagging RPC — with exact recovery commands. Every one of those checks exists because it was a failure I hit myself.


The measured record#

Snapshot as of 13 August 2026, 23:12 UTC.

EnvironmentPaymentsSettledFailedSkew retries fired
Local, healthy window303000
CI on push, one payment per run151500
Probes on GitHub runners, 9 runs × 1513513411 (exhausted)
Total18017911

Every one of these spent real testnet XLM. Nothing is mocked; the conformance harness contains no stubs. The probe and CI rows are traceable to CI artifacts; the local row is not independently traceable and is reported separately for that reason.

Settlement latency, measured across the 135 probe settlements: minimum 1.3 s, median 3.9 s, maximum 7.2 s. An earlier internal figure of “9–15 seconds” came from a handful of runs on a home connection and overstated it; the distribution above is what CI runners actually observe.

A separate cheap sampling of the failure predicate — reading ledger height twice 1.2 seconds apart, 299 times, without spending anything — returned zero cases that would have triggered rejection, with a maximum observed delta of zero.


The upstream defect#

The single failure above is not mine, and finding it is one of the more useful things this project has produced so far.

The mechanism. The public Soroban testnet RPC endpoint load-balances across nodes that can be up to three ledgers apart. @x402/stellar tolerates a divergence of two. When a client reads ledger height from one node and the facilitator reads from another that is three behind, a perfectly valid payment is rejected as signature_expiration_too_far. The failure predicate reduces to clientLedger − facilitatorLedger > 2.

The observed failure. During a manually dispatched probe on GitHub-hosted runners with fresh Friendbot accounts, one payment out of the fifteen in that run failed:

14:57:35.974  /verify  retry attempt 1   invalid_exact_stellar_signature_expiration_too_far
14:57:37.047  /verify  retry attempt 2   invalid_exact_stellar_signature_expiration_too_far
14:57:38.123  /verify  rejected          invalid_exact_stellar_signature_expiration_too_far

Why the original mitigation was wrong. All three attempts landed inside 2.9 seconds — shorter than one ledger close, which is roughly five seconds. The lagging node never had time to advance. Retrying at 750ms intervals could only help by chance of hitting a different node in the pool. The retry was the wrong shape, not merely too short: no number of attempts at that spacing would reliably work.

The fix. Two attempts spaced at least one full ledger close apart, roughly six seconds, both environment-configurable. Configuration validation rejects any retry budget whose count multiplied by delay exceeds 24 seconds, because the resource server’s facilitator client gives up at 30 seconds and converting a recoverable rejection into a client timeout would be a worse failure than the one being fixed. The unit test asserts retry spacing, not attempt count — an attempt-count test would have passed against the broken implementation.

What is validated, and what is only simulated. Being precise here, because the distinction is the whole point:

ClaimStatus
The fault is real and loses paymentsValidated live — one payment lost on a GitHub-hosted runner, 2026-08-12, log above
The old 750ms retry could not recover itValidated live — same run; three attempts inside 2.9s, all exhausted
The 6s backoff spans a ledger closeUnit-tested — test/retry.test.ts asserts the spacing with fake timers, not just the attempt count
The 6s backoff recovers a paymentSimulated only — test/skew-recovery.test.ts drives /verify and /settle through a modelled 3-ledger divergence that clears as ledgers close, and both recover
The 6s backoff recovers a real degraded windowNot observed. No such window has occurred since the fix landed

The simulation models the measured mechanism — divergence beyond a 2-ledger tolerance, clearing as the trailing node catches up — rather than mocking the outcome, and it also reproduces the observed failure: when the divergence never clears, the endpoint gives up and still explains itself. What it cannot prove is that a real degraded window behaves the way the model says. Treat it as evidence that the code path works, not as evidence about the network.

The probe keeps running against real testnet, and scripts/collect-probe-results.sh now labels each skew event RECOVERED or EXHAUSTED, so the first live recovery will be unambiguous when it happens. Until then this row stays as it is. Note also that the fault appears bursty rather than uniformly rare: two events inside one hour on 12 August, then nothing across two days of five-times-daily sampling. A long clean streak is therefore weaker evidence than its size suggests. I cannot manufacture a degraded window on demand, and I am not going to claim a validation I do not have. Scheduled probes run five times daily at spread hours specifically to catch one.

Upstream. The defect, its reproduction, the measured failure rate, the observed ledger spread, exact package versions, and four ordered candidate fixes have been prepared as an issue for the x402 Foundation repository. The strongest of the four is to carry the expiration bound in extra, so the client signs against the same number the facilitator will check, eliminating the class of failure rather than tolerating it.


A retraction#

An earlier internal estimate put the skew failure rate at “roughly 1 in 4.” That figure was an impression formed during an interactive debugging session, not a measurement, and it is retracted.

What is defensible from that window: a pool spread of three ledgers was observed, and two skew events occurred across roughly eight to ten payment attempts — one outright failure, one absorbed by a retry. That is a small denominator and it is flagged as such.

What the 30 clean local runs demonstrate is narrower than it may appear: they show that when the RPC pool is healthy, the failure rate is indistinguishable from zero. They do not demonstrate the retry working, because the retry never fired. Thirty runs in a healthy window cannot bound the behaviour of a bad one.

One anomalous sampling round reported a 232-ledger spread and was excluded as an artifact — the sampler process was suspended mid-round, and neighbouring rounds show normal progression. It is recorded here rather than quietly dropped.


Known limitations#

One in-flight settlement at a time. The current implementation runs single-signer. Settlement takes a median of 3.9 s and up to 7.2 s, during which the signing account cannot start another. One overlapping settlement was observed to produce a hash that never reached a ledger. The channel account pool is the fix and it is Tranche 1 work; until then the harness makes at most one on-chain settlement per run, partly to avoid this: the stock-client payment settles, and the replay check re-submits that same payload expecting rejection. If the payment fails before settling, the replay is what settles — either way, one settlement, never two, because a Soroban authorization entry is single-use.

/verify staleness. For a few seconds after settlement, /verify may still report a settled payload as valid, because the RPC read has not caught up. Settlement still fails, so this is not a double-spend, but it is a real caveat for anyone building on /verify alone.

Skew mitigation unvalidated in the wild. Simulated and unit-tested, never observed recovering a live degraded window. As above.

Testnet only. Mainnet is Tranche 3.

Terminal transcript, not asciinema. The recorded demo session is a text transcript rather than a replayable cast.


How to check any of this yourself#

  1. Clone the repository and run the stack. One command, no secrets required.
  2. Fork it and let CI run. The workflow needs no secrets — accounts are created per run via Friendbot — so a fork produces an independently-generated settled payment.
  3. Open any transaction hash above on a Stellar testnet explorer.
  4. Run scripts/collect-probe-results.sh against accumulated probe runs for the current cumulative counts.

I would rather you verify than believe me.