← Research
Working paper · Public draft

Evidence before trust: a buyer-centred model for autonomous service verification

A design argument for keeping evidence and judgment separate when autonomous buyers evaluate machine services.

Author
Atinamos Labs
Published
16 September 2026
Version
0.1
Status
Public draft · not peer reviewed
ABSTRACT

Autonomous buyers need more than a payment rail and a service address. They need a defensible basis for deciding whether a particular service should be used under a particular buyer's constraints. This paper sets out an evidence-first model in which observable service behaviour is recorded separately from the policy that interprets it. The aim is not to produce a universal trust score, but to make later decisions more inspectable.

Research question

Can an autonomous buyer discover a machine service, decide whether to use it, pay within defined limits, receive a result and retain enough independent evidence for a later policy decision — without relying on a central authority to label the service “trusted”?

The distinction matters because the technical ability to transact is not the same thing as a reason to transact.

Evidence and trust are different things

A single score compresses several different questions into one number. Was the service reachable? Did it request the expected payment? Was payment settled? Was anything returned? Was the result correct for the test that was run? Those are observations. Whether those observations are acceptable is a buyer-policy question.

Atinamos therefore treats evidence as an input to a decision rather than the decision itself. Two buyers can inspect the same evidence and reasonably reach different conclusions because their spend limits, risk tolerance, service requirements or evidence thresholds differ.

A bounded evidence model

The useful unit is not “is this service trustworthy?” but “what was observed in this defined interaction?” A research record should therefore identify the operation under test, the conditions that mattered, the evidence captured and the point at which interpretation begins.

Depending on the experiment, useful evidence can include request and response characteristics, payment conditions, execution state, fulfilment, independently observable settlement, timing and a correctness test where one can be defined without pretending that every service has an objective quality metric.

Buyer policy stays with the buyer

The evidence layer should not quietly become a policy authority. A buyer may care about price ceilings, data handling, accessibility, response time, independent settlement evidence or previous test history. Another buyer may care about a different subset.

This separation makes the system less convenient than a single score, but it also makes the decision easier to explain: the evidence says what was observed; the policy says why that evidence was or was not sufficient for this buyer.

What a passing test does not prove

A successful bounded test is not proof that a service is universally safe, honest, high quality or suitable for every buyer. It does not guarantee future behaviour, establish every aspect of identity, or remove the possibility that a service behaves differently outside the tested conditions.

Research outputs should state those limits explicitly rather than letting a narrow observation harden into a broad reputation claim.

Research programme

The practical work around this model focuses on the parts that fail when a diagram becomes a real transaction: discovery, exact request construction, spend controls, challenge handling, payment boundaries, replay protection, fulfilment, independent evidence collection and publication of failures as well as successes.

The purpose of those experiments is not to prove that one architecture has “won”. It is to find which boundaries remain reliable when autonomous systems have to spend real value and justify what they did afterwards.

Limitations and open questions

The model still has difficult questions. Evidence can go stale. Services can adapt to known tests. Some outputs are difficult to score for correctness. Independent observation has cost. Reputation can leak back into evidence systems through ranking and presentation. A buyer policy can also be badly designed even when the evidence underneath it is sound.

These are research problems, not footnotes. Later versions of this paper should change as experiments produce stronger evidence or expose weaknesses in the model.

How to cite

Atinamos Labs (2026). Evidence before trust: a buyer-centred model for autonomous service verification. Version 0.1. https://atinamoslabs.co.uk/research/evidence-before-trust/

Revision history

v0.1 · 16 September 2026 — Initial public draft.