I Tried & Failed to Build My Own ChatGPT. So I Built Biggie a Boxing Ring instead.
I read about a fella called Andrej Karpathy building his own ChatGPT from scratch.
For about a hundred quid.
Apparently he is one of the Gods of Ai that teaches us mere mortals what there is to know.
I thought to myself: I'll have some of that.
I'll build Biggie from the ground up.
That would definitely be cool.
I think now would be a good time to remind you that I was pushing Zuckerberg and McConaghey out of the picture with misplaced confidence and enthusiasm only a few blog posts ago.
So you can probably tell what's coming next.
A dose of reality that I badly needed.
I looked at what it actually takes.
The hardware - a LOT of hardware £€$
The hours.
The money.
I have none of the three.
Still in the middle of year one of identical twin girls (+ a 2.5 year old and a 4 year old) & very little sleep 😅
Yes Mr Karpathy, you're safe for now.
Building your own ChatGPT from scratch is a rich mans game, or at least for someone who has a weekend and more tech skills. That's not me right now.
So I asked a better question.
Not: am I able to build the whole thing? Obviously not.
But: What part of this is actually useful to me?
And the answer was the bit I hadn't even been looking at.
Not the LLM model itself. It was the way he tested it & measured its effectiveness. How we took a step by step approach to proving it was good instead of just hoping it was.
Something called a synthetic harness 🤷🏻♂️
Turns out that - according to the Jedi council of myself & Claude code - this we COULD actually build. So we did.
Testing Biggie via a harness turned out to be very important. I realised once I started that for most apps, testing their chat system is low grade quality assurance. For what I'm building it's one of the central features aka a surface where trust and safety can break down.
And that - I had to get as right as is humAInly possible (see what I did there? 🙈😅)
Because with pain, words do damage. Im reminded of one of my favourite quotes when writing this:
"Handle them carefully, for words carry more power than atom bombs" - Pearl Strachan Hurd
She's dead right.
16 years of listening to how this played a part in creating suffering in people's lives proves just how right she was. The wrong sentence and delivery of that sentence can potentially set someone back months.
And I'd built a thing that talks to people in pain, on their worst days, without me actively in the room.
And more importantly you cannot test "does this frighten vulnerable people"… on actual vulnerable people.
There's this little thing called ethics ya see….
So what this synthetic harness allows me to do was build a population of people who don't really exist. Fake chronic pain sufferers.
Niamh, Rory, Mairéad, Cathal, Mark, Ethan etc.
Each with a history. A fear. A way of showing up and seeking help. Different horses for different courses as we say in Ireland.
I also built nasty ones. Personas I designed to trip Biggie up. To argue. To demand a diagnosis. To bait him into catastrophising. To drag him off-script.
This was weirdly a bit of fun - to try to figure out how to build the type of character that would drive Biggie towards the type of responses that I never wanted to see in the app.
What a thought experiment this was.
Then I set them loose on him in our virtual "boxing ring". Time and time again. Round after round with Tyson, Ali and the likes.
And we scored every answer against the things that actually hold water in chronic pain care - using the results to figure out how we could sharpen how Biggie responds again and again, until he held up the standard I wanted to deliver.
Safe, providing clarity and as trustworthy as possible.
We put Biggie through round after round of this - and lifted his pass rate in our safety testing by roughly 70%. And that's against the full set, baiters and all: the nastiest boxing ring with the toughest opponents I could build for him.
Is he "done"? Not a chance.
But I'd rather measure him against the hardest crowd imaginable and know exactly where he stands than score him soft and kid myself. So this work carries on in the background - and I'll keep this post updated as that number climbs. (Last updated: June 2026.)
So I couldn't build my own LLM like I naively thought I would be able to.
But I built a virtual boxing ring around the one I'm using and drilled and constrained him until he held his own against all manner of contenders.
And the one that matters most to me - the persona in real despair from their pain. Biggie holds space, and he knows when to step back and point toward real-world help rather than proceed. The very hardest crisis turns are still where he might overstep the odd time, and that - above everything else in here - is the bar I'm still raising.
All in the name of Primum Non Nocere - First Do No Harm.
That's important to me.
I wanted to get this right before I let a single real human near it. And I think I did my damn best at this.
But… Synthetic users can only catch what Biggie says.
They can't feel it if the floor they're standing on is quietly rotting underneath him. And something had been hiding in the dark of my own code for months doing exactly that.
I just didn't know it yet…