AI implementation
Most of what a business repeats follows a rule and belongs in ordinary software. A smaller share needs a decision each time: what to tell a customer, which request to take first, where an answer is buried in old records. That share is where AI belongs, and only under conditions — proven against your own history before it starts, watched after launch, and never in front of a customer without a person approving the words.
The fit
This engagement is for businesses where part of the weekly grind is thinking work — reading, wording, sorting, finding — and it lands on the people who have better uses for the hour. If several of these sound like your week, the fit is probably there.
The method
Six decisions separate AI that quietly helps from AI that gets turned off after one bad week. This is how we make each of them.
Not every step deserves a model. Work that follows a written rule — copy this field, send that reminder on day three — runs as ordinary software, which costs less and behaves the same on Friday as it did on Monday. AI is kept for the steps where someone currently has to think: wording a reply, judging what a request is really about, deciding what an old record means. Sorting the two apart is the first working session, and the AI list comes out of it shorter than it went in.
Anything written for a customer is produced as a draft and stops at a named person on your team, who edits, approves or discards it before it goes anywhere. This is not a temporary training phase; it is the permanent design. The speed gain survives the checking: working from a competent draft is far quicker than writing from scratch, and the reviewer stays close enough to the output to notice the day quality slips.
Every AI step carries a confidence threshold. When the model is unsure, or a request looks unlike anything it was tested on, the item routes to a person with the context gathered so far attached — no answer is forced out. We design against the confident wrong answer, because a polite "someone will get back to you" costs minutes, while a mistake delivered smoothly can cost the customer. The chain always ends with a person, never with a guess.
Before an AI step touches live work, it is scored against examples drawn from your records: replies your team actually sent, requests they actually sorted, questions with answers we can check. Each task gets its own test set and its own passing bar, and a step that misses the bar does not launch — the task stays with a person or gets a simpler design. When a failure shows up later in production, it joins that test set, so the same mistake cannot pass quietly twice.
Going live is the start of the measurement, not the end of it. The system keeps a record of what each AI step did and why, and we watch how often reviewers rewrite drafts — the honest signal of quality — alongside cost and volume. New models from the providers are treated as candidates, not upgrades: each one runs the existing test sets first and takes over a task only where it scores at least as well as the current one. The previous model stays installed, so a change that misbehaves is rolled straight back.
The most capable models cost many times more per request than the small ones, and most tasks in a business do not need the most capable. Each step runs on the smallest model that passes its tests: sorting a message into six categories is small-model work even when drafting a sensitive reply is not. Because the tests are per task, moving down a size is a decision backed by evidence, not a hope. You see the projected monthly running cost per task before launch, so the operating bill is a figure you approved rather than one you discover.
The engagement
We go through the candidate work with you and split it: rule-following steps to ordinary software, genuine judgment calls to the AI list. You get that list with a projected running cost and a build order, and the fee for each phase is agreed in writing before that phase begins.
For every AI step we assemble the test set from your records and measure candidate models against it. You see the results in plain terms — how often the draft matched what your team would have sent — and nothing moves to a build until it clears the bar.
The system goes live with every review gate on. We watch the first weeks alongside your team, adjusting thresholds where real traffic differs from the history, until the volume reaching people is right.
Code, accounts, documentation and the test sets transfer to you, and the people who will run it are trained on real work. Support after handover is offered and priced on its own; the system keeps working whether we stay involved or not.
What transfers
Every engagement is built to be handed over. These are yours once the work is paid for, not licensed back to you.
Next step
Bring the piece of your week that needs judgment but eats time: the inbox, the replies, the hunting through records. On a short intro call we will tell you honestly whether AI belongs in it, and where we would begin.
Book a callAlso in the catalog