The first week with a new agent
Run it alongside the human for a week and compare. Do not let it act unsupervised on day one.
Get an agent into real use without discovering its failure modes in production.
How the first week with a new agent works, in one picture
The same argument as the text, as a chain. Each step is what makes the next one possible.
- 1
Day one: read-only, output to you
Whatever the agent is eventually going to do, start it in a mode where it produces a draft and a person sends it. Almost every agent supports this, and where one does not, that is worth knowing on day one rather than day thirty.
You are not testing whether it works. You are finding the category of thing it gets wrong.
- 2
Days two to five: run it in parallel and count disagreements
Your own workflow, not a dashboardKeep doing the work the way you do it now, and let the agent do it too. Compare. Count how often you would have sent what it drafted.
A number under about eight in ten means it needs different instructions or different data, not more time.
- 3
Check the spend on day three, not day thirty
Metered listings can be a great deal and can also run away, and the difference usually shows up within seventy two hours of somebody discovering it.
Look at the actual number early, while the refund window is still open.
The refund window on most listings is fourteen days. A month-end surprise is outside it.
- 4
Day seven: decide, and write down why
Keep it, change how it is used, or return it. All three are fine outcomes and the third is why the window exists.
Write one line about what it was good and bad at. In six months, when you are looking at its replacement, that line is the most valuable thing you own.
You have a decision, a reason, and a spend figure from real use.
Read next
Run it alongside the human for a week and compare. Do not let it act unsupervised on day one.
The near-monopolies and the commodities, side by side, because they look identical from outside and they are not.