An agent pilot is one AI agent, built to do a single job, with the outcome it will be judged on agreed in writing before it is built and two weeks of measurement afterwards. The fee is fixed. We do not sell agent building by the hour, and we do not start until the data the agent would be reading has been checked, because an agent reading incomplete records produces confident wrong answers faster than a workflow does.
Why start with one agent rather than several?
Because one agent can be judged and several cannot. With one there is a baseline, a single thing that changed, and a number at the end of two weeks. With five there is an impression.
It also keeps the first transaction small. If the agent does not earn its place you have spent one fixed fee finding that out, which is a considerably better outcome than discovering it across a programme.
Why will you not build before checking the data?
An agent is only as good as what it can read. It needs the right properties populated, associations that actually connect the records it reasons across, and lifecycle stages that mean what they say. Where those are missing the agent does not fail loudly. It produces plausible output that is wrong, which is worse.
So the pilot follows the Customer Platform Assessment, which is where the candidate agents are identified in the first place, along with the data each one depends on. If you have had the assessment, the pilot is the next step and the candidates are already named. If you have not, that is where to start.
Which agent is usually worth piloting first?
These six recur across the businesses Flowbird works with, which is why they tend to be where the assessment lands:
1. **Customer health scoring.** Reads engagement, support history and renewal data, and flags accounts drifting before the renewal conversation does.
2. **Closed-lost postmortem.** Reads the deal record and the correspondence on it and writes up why it actually went, rather than the reason somebody picked from a list.
3. **Lead qualification with web research.** Checks an inbound lead against your ideal customer profile using public sources, so the first human touch lands on the ones worth touching.
4. **Onboarding handoff notes.** Turns a won deal into a written handover the delivery team can act on without re-interviewing the customer.
5. **Demo preparation.** Assembles what is already known about an account into a short brief before the call, so nobody opens a meeting reading the record for the first time.
6. **A data quality monitor.** Watches for the duplicates, blanks and stalled records that quietly degrade everything else, and reports rather than silently fixing.
Yours may not be on this list. The assessment decides it against your own data, not against this page.
What does "proved" mean here?
Four things, agreed before building rather than argued about afterwards:
**The outcome.** One sentence naming what should be different, and how it will be observed.
**The baseline.** What that number is today, measured before the agent exists.
**The credit ceiling.** A limit you set, so consumption cannot surprise you.
**The window.** Two weeks running against real work, rather than a demonstration on selected records.
Who pays for the AI credits?
You do, and they stay yours. The credits sit in your own subscription rather than being resold through us, so you can see what the agent consumes and you keep the relationship with the platform. Many businesses already have credits included in a plan they are not using, and that is often the cheapest place to start.
The assessment sizes expected consumption before anything is built, and the pilot measures what it really cost, so the figure you plan the next one with is your own rather than a vendor's.
Which platforms can you build an agent in?
Where you run HubSpot, the agent is built in HubSpot, using its own agent tooling and its own credits. Where you run Pipedrive, ActiveCampaign or Workbooks, it is built in Make, which connects to the CRM you already have rather than asking you to move to one that suits us. On FlowFlex it is built into the platform directly.
In every case it is built under your account. It does not leave when we do.
What happens after the pilot?
You get a recommendation, and it may be to extend it, to change it, or to stop. Whichever it is, the agent is handed over documented: its instructions, the knowledge it draws on, the properties it depends on, and what it measurably cost to run.
That documentation is deliberate. The same agents recur across businesses, so writing each one down properly means the next business that needs the same thing gets a shorter build rather than a rebuild from nothing.
What if the pilot does not work?
Then it did not work, and you have a written reason. Two weeks of measurement against an agreed baseline is enough to tell the difference between an agent that is not earning its place and one that needs a better prompt or better data underneath it, and the report says which.
We would rather tell you that than extend a programme on the strength of a good demonstration. Agents and automation both depend on the data underneath them, which is the same argument CRM automation makes, only sharper.
What does a pilot cost, and what do you get for it?
A fixed fee for one agent, built in your own platform, with the outcome it will be judged on agreed before anybody starts building it.
One agent, built where your CRM already lives, with a defined outcome, a credit ceiling and two weeks of measurement against it. We do not sell agent building by the hour, and we do not start before the data the agent would be reading has been checked.
Agent Pilot
£2,500, after the assessment
One agent, scoped to a single job it can be judged on. Built in HubSpot where you run HubSpot, and in Make where you run Pipedrive, ActiveCampaign or Workbooks.
You set the credit ceiling before anything starts, and we measure what it actually consumed against what it actually changed. If it does not earn its place, that is a finding, and it cost one fixed fee to get it.
One agent, with the outcome it will be judged on agreed in writing before any building starts
A credit ceiling you set, and the measured consumption against it
Built in your own platform under your own account, so it does not leave when we do
Its instructions, the knowledge it needs and the properties it depends on, documented and handed over
A two-week measurement window against the baseline, rather than a demonstration
A recommendation at the end, which may be to extend it, to change it or to stop