Turn AI into a working advantage for your business.
With 17+ years delivering software and complex technology projects, I help businesses identify valuable AI opportunities, define the requirements, build and test the solution, and put it into real use.
From global automotive and insurance platforms to AI systems now operating inside a growing business, my work has always sat between business needs, technology, delivery and quality.
Selected organisations, clients and programmes I have worked with directly or through consulting and delivery partners.
Most AI projects don't fail because the technology was wrong. They fail because nobody was precise about the problem before the building started — the opportunity was assumed rather than sized, the requirements were vague enough to satisfy on paper, and no one agreed what "working" would mean. The model was never the hard part.
I work end to end on that whole arc: understanding how the business actually operates today, identifying where AI earns its place and where it plainly doesn't, turning that into requirements precise enough to build and test against, then designing, building and shipping the thing. Seventeen years of delivering software and complex technology projects means I've seen how this goes wrong often enough to steer around it.
Discover Define Design
Before anything is built, it's worth knowing what the current process really costs. Most teams have a rough feeling about where time disappears and are wrong about the specifics. This phase replaces the feeling with a number, and separates the work AI should touch from the work it shouldn't.
A requirement that can't be tested isn't a requirement, it's a hope. This is where the problem becomes precise enough that a developer can build it, a tester can verify it, and everyone recognises the result as what they asked for.
The riskiest assumption should be tested first and cheaply. A working prototype settles arguments that documents extend — and it surfaces the problems that only appear once something is real.
Build Test Handover
Building the thing, and connecting it to the systems that already run the business. Most of the difficulty here is integration, not invention — the AI is rarely the part that resists.
AI systems fail differently from conventional software: they're non-deterministic, they degrade quietly, and they can be confidently wrong. That needs a testing approach built for it, not a regression suite borrowed from elsewhere.
A system nobody uses is a cost, not an asset. The last phase is the one most projects skip, and it's the one that decides whether the work survives contact with the organisation.
Often it isn't, and you should want someone who will say so. Plenty of problems presented as AI problems are really process problems, data problems, or a report nobody reads. The opening phase exists to establish which one you have — and if the honest answer is that a well-built form and a clear workflow would solve it for a fraction of the cost, that's the recommendation you'll get.
It starts with understanding how the work is done today and where the time actually goes — usually one to two weeks. From there, a prototype of the riskiest part before committing to a full build, because that's what settles arguments cheaply. Then build, test and handover. Small automations can be live in weeks; a system that touches several existing platforms takes longer, mostly because of the integrations rather than the AI.
Yes, and that's usually the better arrangement. Very little of this work happens on a blank page — most of the difficulty is connecting to platforms that already run the business and were not designed with this in mind. I'm comfortable working inside an existing codebase, alongside an existing team, and to whatever delivery process you already use rather than imposing another one.
It will, so the design has to assume it. That means deciding in advance where a human stays in the loop, what the system does when it isn't confident, and how a wrong answer gets caught before it reaches somebody who'd act on it. A system that's right 95% of the time and silent about the other 5% is more dangerous than one that's right 80% of the time and says so.
Sometimes you should, and if your team has the capacity and the experience I'll say that. What I typically bring is having watched this go wrong before — knowing which requirements will turn out to be vague, which integration will take three times longer than estimated, and which impressive-sounding use case should be the last thing you attempt rather than the first. That's worth most at the start, which is also when it's cheapest to apply.
Quality isn't a stage at the end of a project — it's a property of how the project was run. By the time testing is a phase you're squeezing to hit a date, the important decisions have already been made. The teams that ship reliably aren't testing harder at the end; they're deciding differently at the start.
I've spent most of seventeen years on this: manual and exploratory testing where judgement matters, automation where repetition does, and test strategy for platforms where being wrong carries real consequences. Much of it in insurance, and specifically life insurance — an industry where a defect isn't an inconvenience, it's a payout calculated wrongly, a policy priced incorrectly, or a compliance obligation quietly missed.
Discover Define Design
Not everything deserves the same level of testing, and pretending otherwise is how budgets get spent in the wrong places. The first job is working out where failure would actually hurt.
A test strategy is a set of deliberate choices about what to check, how, and when — not a document written to satisfy an auditor. Done properly it makes the release decision obvious rather than political.
Coverage is not a percentage, it's a question: which of the things that could go wrong would we catch? Designing for that produces far fewer tests than teams expect, and far better ones.
Build Test Handover
Scripted tests confirm what you already thought to ask. Exploratory testing finds what nobody thought to ask, which is where the expensive defects live. It's a skill, not a fallback for teams without automation.
Automation pays back on the tests you'll run a thousand times, and loses money on the ones you'll run twice. Knowing the difference is most of the value — an automation suite nobody trusts is worse than none, because it produces confident false assurance.
The goal isn't to be the person who finds the bugs. It's to leave behind a team that catches them earlier than you would have.
Usually not more testing — usually better-aimed testing. Teams with capable testers frequently still lack an agreed strategy about what deserves scrutiny and what doesn't, so effort spreads evenly across things that carry very different risk. The gain tends to come from redirecting the same people, not adding more of them.
No, and trying to is one of the more expensive mistakes available. Automation pays back on checks you'll run hundreds of times and loses money on checks you'll run twice. It's also worth being blunt about the failure mode: a large automated suite that nobody trusts is worse than having none, because it produces confident assurance that isn't real.
Conventional software is deterministic — same input, same output, so a test either passes or fails. AI systems aren't. They can produce a different answer to the same question, degrade quietly as the world changes around them, and be wrong with complete confidence. That needs a different approach: testing ranges and behaviours rather than exact outputs, monitoring for drift after release, and deciding explicitly where human judgement is required.
Yes — much of my experience is in insurance, and specifically life insurance, where a defect can mean a payout calculated wrongly or a compliance obligation missed. That environment demands traceability from requirement through to test evidence, and being able to demonstrate not just that something was tested but why it was tested that way. If you operate somewhere similar, that discipline is already second nature.
A test strategy your team agrees with rather than tolerates, coverage aimed at genuine risk, whatever automation earns its place, and quality gates built into your delivery pipeline so verification happens without anyone having to remember. The aim is to leave a team that catches problems earlier than I would have, not a dependency on me.
Most organisations now have access to the same AI tools. Very few have people who know what to ask of them, where the output can be trusted, and where it quietly can't. The gap isn't licences — it's judgement, and judgement is transferable if someone who has built the habit sits down with the people who haven't.
I work with teams and individuals to turn AI from something they've heard about into something they use on Tuesday. That means starting from the work they actually do, not from a demo, and being straight about the limits — the fastest way to lose a team is to oversell it once.
Discover Define Design
Teams are rarely where they think they are — some are further along than they claim, others have been quietly using tools nobody sanctioned. Knowing the real starting point prevents training people in what they already know.
The best first use case is small, frequent, low-risk and annoying. Teams tend to nominate their most impressive problem, which is usually the worst place to start.
A tool bolted onto an unchanged process usually adds a step. The value arrives when the workflow itself is redrawn around what's now cheap.
Build Test Handover
Nobody learns this from a slide deck. It lands when someone sits with you on your own real task and you watch the thing work — then do it yourself while they're still in the room.
What one person works out on Tuesday should be available to everyone by Thursday. Without somewhere for it to live, every team member solves the same problem independently.
Enthusiasm after a workshop is not adoption. Six weeks later is when you find out whether anything changed.
Access isn't capability. Most teams have people using these tools for a narrow set of tasks they worked out alone, with no shared sense of where the output can be trusted and where it can't. The valuable part isn't the tool — it's judgement about when to reach for it, how to check what comes back, and which work should stay well away from it.
A course teaches the tool in the abstract and gets forgotten within a fortnight. This works from the tasks your team already has in front of them — their real work, their real constraints — and the measure of success is whether they're still doing it differently six weeks later, not whether they enjoyed the session.
Then that's the starting point rather than an obstacle, and it's a good sign somebody has thought about it. Constraints usually shape which tools and which use cases are viable, and occasionally rule out an entire category. Working within them from day one is far better than building a workflow that has to be dismantled later.
Against the baseline taken before starting. Enthusiasm immediately after a session tells you nothing — adoption is what's still happening once the novelty has gone. That means agreeing up front what would count as success, and being willing to report honestly when a use case didn't earn its place.
Less than most training, because it happens on work they were going to do anyway. Typically short working sessions rather than days out, then follow-ups spaced far enough apart that habits have had a chance to form or fail. If it needs a large block of everyone's calendar, it's probably designed wrong.
“We had the licences and no real idea what to do with them. Flo worked out where AI genuinely helps us and where it does not, then built the thing that proved it. Reporting that used to take days takes an afternoon.”
“Every client report used to be rebuilt by hand and never came out looking quite the same twice. Now it is identical every time and I get my evenings back. Finally something that actually matters to how we work.”
“Flo sat with the whole studio, worked out what we actually do all day, and only then brought AI anywhere near it. No hype, no jargon. Months on, people reach for it without being told to.”
“Content was our bottleneck. We now produce more in a week than we used to manage in a month, in both languages, and it still sounds like us rather than like a machine. Fast, and genuinely ours.”
“I used to lose the first two days of every week to research. That is about five minutes of reading now, and the rest of the week is actual work. I would not go back to doing it by hand.”
“He handed me something I could run on my own from the first day — one form, one prompt, and a playbook that answers the question before I ask it. I have never had to chase him to keep it going.”
“We were drowning in reports. Flo built us something that turns the data straight into a finished, properly formatted document first time. What took the better part of a week now takes a morning.”
“Flo works the way you want a senior engineer to work. He asked the awkward questions early, found problems in our flow we had stopped being able to see, and then quietly got on with fixing them. Great to work with.”
“We had tried AI twice before and both times it stalled. This time it stuck, because he started with how we already work instead of with the tool. It has saved us most of a role in admin.”
Flo Hofmann — freelance software delivery and QA expert.
Pricing is based on scope, not hours. After an initial conversation you get a clear proposal that breaks down what is included, what the deliverables are, and what it costs — no vague day-rate estimates that balloon later. For ongoing retainer work, a monthly scope is agreed and adjusted each quarter if needed.
Both, depending on what suits the project. Fixed-price works well when the scope is well-defined — discovery, a design sprint, a handoff package. Time-and-materials makes more sense for embedded retainer work or longer-running projects where scope naturally evolves. The right model gets recommended for your situation.
Typically a deposit to secure the dates, then progress payments tied to milestones rather than calendar dates. Invoices are issued on completion of each milestone with fourteen-day terms. For retainers, invoicing happens monthly in advance.
It gets flagged before any work starts on it, with an estimate of the impact on cost and timeline. Nothing outside the agreed scope gets built quietly and invoiced later. Small changes are usually absorbed; anything larger becomes an explicit decision you make.
Every engagement starts with a short planning pass that maps the outcome, the constraints, and the sequence of work. That becomes a written plan with milestones you can hold the work against, and it gets revisited whenever something material changes.
Interviews with the people who will use the thing, a review of whatever analytics and support data already exists, and a walkthrough of the current experience. Discovery ends with a written synthesis and a prioritised list of the problems worth solving.
Two structured review rounds per deliverable are included, which in practice is more than enough because feedback is gathered continuously rather than saved up for a big reveal at the end.
A named decision-maker, access to whoever will use the product, and any existing research, analytics or brand material. Everything else can be worked out together in the first week.
Yes. Retainers work well once a product is live and needs continuous design attention. A monthly scope is agreed up front and reviewed each quarter so it keeps matching what the team actually needs.
You receive the working files, a documented design system where one exists, and a walkthrough session with your engineering team. Follow-up support for the first few weeks after handoff is included by default.