Research note 001 · 5 September 2026

What would make an AI-agent limit earnable?

A public question prompted by AI underwriting — not a case study, endorsement or claim about any company's internal practice.

An AI agent can be technically capable of acting long before a risk-owner will permit it to act at a meaningful economic limit. The unresolved question is what evidence should move that limit — if any.

Why this question is concrete

Armilla AI publicly describes itself as a Lloyd's-backed MGA providing AI liability insurance and performance warranties for generative AI and AI agents. Its team identifies Andrew Correll as Director of Underwriting. Public statements from Armilla and Correll describe work across AI performance warranties, third-party liability and agentic systems.

Armilla's public MKIII example describes a different but adjacent mechanism: validation of an AI loan-decision model's performance, fairness and robustness, followed by a warranty that triggers if accuracy drops below verified thresholds. That is a useful distinction. A verified pre-deployment or model-performance threshold is not the same thing as a longitudinal record of real operating outcomes after deployment.

Those facts do not tell us what limit Armilla sets, what a particular customer has been offered, or what evidence its underwriters require. EARNED has no access to that information and makes no claim to it.

The hypothesis

For one bounded agent action, independently verifiable longitudinal outcomes may change the authority, exposure limit or insurance terms a professional risk-owner is willing to permit.

This is a falsifiable hypothesis, not a product claim. A sensible answer may be that outcome history is insufficient, irrelevant, or outweighed by controls, task design, portfolio effects, policy wording or other evidence.

What a decision test would hold constant

A controlled decision exercise begins with a hypothetical AI agent and a defined authority boundary. Its identity, mandate, policy controls, counterparties, transaction limits, revocation path and audit trail remain constant.

Only one class of information is added: a longitudinal record that separates confirmed successful outcomes, reversals, losses, policy exceptions, human interventions, internal-only observations and unresolved evidence. Nothing unresolved is converted into positive evidence.

The decision question

First, a risk-owner records the maximum authority, exposure or terms they would permit with the control baseline alone. Then the same owner sees the outcome record and makes the same decision again.

The result is not a score. It is a delta, if one exists:

Does the decision move? If so, which evidence moved it? If not, what was missing or irrelevant?

What would count as useful evidence

  1. A real decision owner: someone who can set, insure or materially constrain the relevant boundary.
  2. An explicit before and after: a number, threshold, term or documented accept/reject decision — not general interest.
  3. A stated reason: the specific evidence that changed the decision, or the reason it did not.
  4. A real constrained case, if available: the next test must be about that case, not an abstract platform.

An invitation to correct the premise

If you underwrite, insure, govern or deploy AI systems that take consequential actions, the most useful response is a correction: which part of this decision model is wrong, incomplete or non-decisive in practice?

EARNED is an early-stage independent research initiative. This note is not advice, a solicitation of insurance, an audit, a certification, or an assessment of Armilla AI.

Public sources