A Conversation with Sanmi Koyejo
There’s a boxing gym in Austin, Texas that you’d easily miss if you weren’t looking for it. No AC. Grimy walls. A coach named Bruce who yelled at everyone. It was the kind of place you went to when the polished gyms felt too curated.
Sanmi Koyejo trained there. So did I. We both had no idea at the time, where life would take us.
Sanmi is now an assistant professor of computer science at Stanford, where he leads the Stanford Trustworthy AI Research group and runs Virtue AI, a company working on trust and safety for AI systems. He’s writing a book on machine learning from human preferences. And we both went to International School of Lagos (ISL).
When I found out all of this, I knew this conversation had to happen.
The Gap Between What We Built and What We Need
One of the sharpest things Sanmi said came when I asked about benchmarks and measurements in AI.
He explained that the measurement systems that guided the field for decades were built for a very specific purpose: tell researchers, roughly, whether they were doing better or worse. Not to make decisions. Not to be interpreted by policymakers or CEOs or regulators. Just to give a signal.
But now those same numbers are being trumpeted by major AI labs when they release new models. Governments are making policy based on them. Investors are moving billions based on them. And the systems were never designed for that.
“The gap that we’re seeing,” Sanmi told me, “is that the reason we were building measurement systems before – and the kinds of work that we did – don’t actually answer the call for what measurement is required to do now.”
Fifty Percent Data, Ten Percent Algorithms
I asked him what percentage of real AI work is data and engineering versus the algorithms that get all the attention.
He didn’t hesitate: roughly fifty percent data and engineering, ten percent the stuff academics actually teach.
“The algorithmic – that’s the thing I’m most excited about,” he said. “But the things that have led to meaningful, significant outcomes are a combination of data and engineering.”
It’s one of those honest moments that only comes from someone deep enough in the field to see the gap between what’s celebrated and what actually works.
The Exploration Is the Job
But what stayed with me most was Sanmi’s description of research as a way of life.
“I say this a lot to my PhD students: in terms of making money, in terms of most of the obvious signals for career success, you’re probably better off actually not doing a PhD and not being a researcher. I think you should do research if you value the exploration process. Because you’ll do a lot of that in a way where you won’t get – there’ll be long periods where you don’t get signal that you’re going in a good direction.”
That’s rare honesty for an academic. But it’s also the most useful thing you could say to someone deciding whether research is for them.
The exploration is the job. The long periods without feedback are the job. If you can’t live in that uncertainty and still find meaning in the process, research will break you.
Watch the Full Conversation
We also talked about the fire alarm at a conference in College Station that led to my career at Rockwell (a serendipity story that still makes me laugh), what trustworthy AI actually means technically, and what he’d do with his time in a utopian world.
The full episode is embedded above (paid subscribers get early access to Practice Ground episodes).
Also on YouTube and Spotify/Apple Podcasts.
From that boxing gym in Austin to Stanford.
Some paths you can’t plan — you can only be ready when the moment comes.
— Nifemi