Two pytest results for the same Fetch.ai uAgent handler that raises on every message: standing the agent up the usual way reports PASS because the framework catches the exception into a log, while uagent-testkit's deliver() reports FAIL with the ValueError surfaced

You wrote a Fetch.ai uAgent. It has an @on_message handler that does something worth getting right — parses a request, updates storage, sends a reply. Now you want a test for it, the kind you can run on every commit. And you find there is no obvious place to put one.

The framework was not built with a test seam. There is no agent.handle(msg) you can call and assert against. So the path most people take is to stand the agent up the way it runs in production and talk to it over a socket: register on the Almanac, resolve an endpoint, on some paths fund a wallet, then asyncio.sleep() your way through the test hoping the reply arrives before the timeout. It is slow, it is flaky, and it is miserable in CI — a network hiccup fails a test that has nothing to do with the network.

That is the annoying problem. Underneath it is a worse one.

A green test that should be red

uAgents wraps every handler in a try/except. When your handler raises, Agent._handle_message catches the exception and writes it to a log instead of letting it propagate. At runtime that is exactly right: an agent processes messages from peers it does not control, and one malformed message must not take the whole process down.

But there is no strict mode to turn that containment off, and it does not know it is running under pytest. So the same swallow happens in your test. Consider a handler with an ordinary bug — it assumes a field is numeric:

@agent.on_message(model=Order, replies=Receipt)
async def on_order(ctx, sender, msg):
    ctx.storage.set("orders", (ctx.storage.get("orders") or 0) + 1)
    total = int(msg.amount)              # msg.amount arrives as "12.50"
    await ctx.send(sender, Receipt(total=total))

Send that agent an amount of "12.50" and int() raises ValueError on every call. Stand the agent up and drive a message through it, and the traceback lands in a log line, the handler returns as if it were “handled,” no reply is sent — and your test, which was waiting for the process not to fall over, goes green. The bug is real and shipping. The signal that would have caught it was absorbed one layer down.

Route the message straight to the handler

The fix is to stop talking to the agent over a socket and call the handler directly, with a context that records instead of transmits. That is the whole idea behind uagent-testkit, a pytest harness we built while shipping uAgents of our own:

from uagent_testkit import harness

async def test_order_is_handled():
    h = harness(agent)
    result = await h.deliver(Order(amount="12.50"), sender="agent1qexample")

    result.assert_replied_with(Receipt)       # it declared replies=Receipt; did it send one?
    assert h.storage["orders"] == 1

No Almanac, no endpoint, no wallet, no sleep(). The message goes to the registered handler; ctx.send is intercepted and its payloads collected on result.sent; storage is a fresh in-memory store. And crucially, deliver() re-raises whatever the handler raised. Run the test above against the buggy handler and it does not go green — it fails on the ValueError, at the line that caused it. When the failure is the thing you are testing, you opt back into the framework’s behaviour explicitly:

result = await h.deliver(Order(amount="12.50"), raise_errors=False)
assert isinstance(result.error, ValueError)

The quiet failures it turns loud

The swallowed exception is the headline, but the same theme — a real failure that a stand-it-up test reports as success — runs through the rest of the framework. Each of these is an assertion instead of a log line nobody reads:

  • A broken reply contract. A handler marked @on_message(replies=Receipt) that never actually sends a Receipt is a bug on a live network and a shrug in a normal test. assert_reply_contract() makes it fail.
  • State leaking between tests. The stock KeyValueStore writes a JSON file to your working directory, so one test’s storage bleeds into the next and order-dependent passes hide real bugs. The harness swaps in an in-memory store that starts empty every time.
  • An hour-long interval you can’t wait for. An @on_interval(period=3600) or an @on_event("startup") handler is real logic that never runs in a message test. h.tick(), h.startup() and h.shutdown() run them on demand, period ignored.
  • An echo loop between two agents. A keyword responder whose answer contains its own trigger word will talk to another agent forever. AgentNetwork plays the conversation out deterministically and raises ConversationTooLong with the transcript attached, so a loop that costs real money in production fails in milliseconds here.
  • A handler reachable by senders production would reject. Handlers without allow_unverified=True only accept verified agent addresses; the harness enforces that, so a test can’t pass for a path the live network refuses.

Agents that read the chain, tested offline

uAgents in the BNB Chain and ASI ecosystems often read on-chain state — a BEP-20 balance, a token’s decimals — and intercepting ctx.send does nothing about a handler that reaches out to an RPC endpoint. ChainDouble closes that gap by patching whichever of requests, httpx and aiohttp the agent imports:

chain = ChainDouble()
chain.add_token(FET_BSC, symbol="FET", decimals=18,
                balances={wallet: 5_500_000_000_000_000_000})
with chain.install():
    result = await h.deliver(CheckBalance(wallet=wallet))
assert result.reply(BalanceReport).amount == 5.5

It answers eth_call (balanceOf, symbol, name, decimals, totalSupply), eth_getBalance, eth_getCode and the rest from the state your test set. The two things it refuses are the point: reading a value you never stubbed raises UnstubbedCall rather than quietly returning zero, and any non–JSON-RPC HTTP from inside a handler raises NetworkCallBlocked rather than hitting a real API from CI.

What it won’t pretend to do

A test tool earns trust by being honest about its edges, so this one is blunt about them. It does not simulate Almanac resolution, envelope signing, or on-chain settlement. Those are integration concerns, and faking them would hand you a green suite that proves nothing about whether your agent actually registers, signs, or settles — false confidence is worse than no test. Keep one thin integration test on a real testnet for the wire, and use this for everything above it: routing, replies, storage, protocol contracts, and inter-agent flow.

It is free and Apache-2.0, it depends only on uagents itself, and installing it registers the pytest fixtures automatically — there is nothing to wire up:

pip install uagent-testkit

The source and the full API are on GitHub, and it is on PyPI. If you have shipped a uAgent, the fastest way to see the point is to point it at a handler you believe already works — the interesting result is the one that doesn’t go green.

We build small tools that make the failure show up where the bug is, not three steps later. That’s what we do at Rebel Studios.