Bazaar discovery#

Status: designed, not yet built. This is the largest single piece of proposed grant work, delivered across Tranches 1 and 2. Nothing described here is running today.

Discovery is what turns x402 from a payment rail for pre-arranged relationships into an open market. It is also, by some distance, the hardest part of this RFP — and the part most likely to be under-implemented, because a catalogue that merely lists things is easy and a catalogue that answers questions well is not.


What discovery must do#

An agent needs a capability it does not have. It should be able to describe that need, receive a ranked set of paid resources that plausibly satisfy it, evaluate their prices and terms, pay for one, and use it — without a human intermediating.

That decomposes into three problems, in increasing order of difficulty: recording what exists, ensuring the record is trustworthy, and returning the right thing for a query.


The catalog#

Resources enter the catalogue from the payment flow itself. When a seller returns a 402, its body carries resource metadata: a description, a MIME type, an output schema where one exists, an optional route template. After a payment against that resource settles, the cataloging worker records it.

This has a useful property: the catalogue only ever learns about resources that someone actually paid for. A listing is evidence of a working, priced, reachable endpoint — not of an intention to offer one. Spam that nobody pays for never enters the index.

Cataloging is asynchronous and runs after settlement, off the payment latency path, for the reasons set out in Architecture.

/resources returns the catalogue with pagination and filtering. /search is described below.


Listing integrity — the sharpest boundary in the system#

Here is the problem, stated precisely.

The resource metadata that reaches the facilitator has passed through the client. In the x402 flow, the client echoes the seller’s resource block back as part of the payment payload. A hostile client can therefore submit a listing describing a resource that is not theirs.

Left unaddressed, that permits a straightforward set of attacks: claim a competitor’s URL and attach a misleading description; overwrite an existing listing with worse terms; poison the index with entries designed to rank highly for valuable queries and deliver something else; or simply flood the catalogue.

The catalogue is only useful if an agent can trust it. So integrity is not a hardening pass to be added later; it is a precondition for the feature having value at all.

The model#

Ownership derives from settlement, not from claims. A listing is bound to the payTo address that actually received a settled payment. The chain, not the client, establishes who owns a listing.

First-writer-wins on (url, payTo). The first settled payment to a given recipient for a given URL establishes the listing. Subsequent writes from the same recipient update it; writes from a different recipient for the same URL do not overwrite it.

Conflicts are quarantined, not resolved. When two distinct recipients claim the same URL, neither silently wins. The conflict is recorded and the entry is held out of the ranked index pending review. Silent resolution would be a vulnerability: it would let an attacker take a listing by making the system pick.

Percent-decode before traversal checks. Route templates are decoded before they are validated for path traversal, not after. Validating the encoded form and decoding later is a classic ordering bug and it would be exploitable here.

Invalid fields are soft-dropped and reported. A listing with one bad field is not silently accepted with the bad field intact, nor rejected wholesale. The offending field is dropped, the rest is kept, and the drop is reported back to the submitter through the extension-responses header — so a seller with a genuine formatting mistake finds out rather than wondering why their listing looks wrong.

How this gets verified#

An adversarial test suite, running in CI, is a deliverable in its own right — not a byproduct. It demonstrates that a hostile client cannot claim a listing it did not settle to, cannot overwrite another recipient’s listing, cannot traverse via an encoded route template, and cannot poison the index with malformed fields. Completion is measured by that suite passing, not by the feature appearing to work.


Search and ranking#

This is where discovery layers usually stop short. Listing and filtering are straightforward; returning the right resource for a natural-language need is not, and an agent that gets bad results will stop using the catalogue.

Retrieval#

Lexical. BM25 over a synthesised per-resource document — description, path structure, MIME type, schema field names, seller-supplied tags, concatenated into one searchable text. Lexical retrieval handles the case an agent frequently needs: an exact term, an API name, a specific format.

Semantic. Dense vector retrieval over the same synthesised document, stored in pgvector. This handles the case lexical search fails: an agent asking for “current conditions outdoors” should find a resource described as “weather API,” despite sharing no terms.

Fusion. The two ranked lists are combined with reciprocal rank fusion. RRF is chosen deliberately over score-weighted blending: it requires no calibration between two scoring systems with incomparable scales, and it degrades gracefully when one retriever performs badly for a given query.

Re-ranking. A cheap re-rank pass over the fused top-k, incorporating signals the retrievers do not see — price, recency, and observed reliability.

Evaluation — the part that makes this real#

Any team can claim their search works. The differentiator is measuring it.

A golden query set, committed to the repository, pairing representative agent queries with the resources that should be returned. NDCG@10 computed against it, reproducible by anyone via a single script in CI. Ranking changes are then evaluated against a number rather than an impression, and a regression is visible before it ships.

A live signal loop. The payment flow generates a relevance label that most search systems have to infer: searched → paid → did not retry. An agent that searches, pays for a result, and does not immediately search again probably got what it needed. That signal feeds re-ranking without any explicit feedback mechanism.

Completion for the ranking deliverable is defined as published NDCG@10 scores against the committed golden set — not as “search returns results.”


Explicitly out of scope for this grant#

An on-chain resource registry. The RFP names this as a stretch goal, and it should stay one. Putting the catalogue on-chain adds ledger rent and eviction management, and roughly doubles the per-payment cost for a benefit — censorship-resistant listings — that does not yet have demand behind it. The catalogue’s integrity already derives from on-chain settlement; storing the index itself on-chain is a different and much more expensive property.

Batch settlement and auth-capture. Both deferred by the RFP.

If the discovery layer proves out and demand for a trustless registry materialises, that is a strong candidate for a subsequent proposal — with the delivery record of this one behind it.