Most bad AI leadership hires are not sourcing failures. They are evidence failures. The signals boards trust most, lab pedigree, paper counts and a brand-name employer, have almost no track record of predicting performance in a discipline this young. At Olofsson & Company we build the evidence somewhere else: in the operating decisions a candidate has made, in peer-level reference work, in regulatory literacy, and in treating the first 100 days as part of the search rather than someone else's problem.
The comfortable story about failed hires
When a senior hire does not work out, the post-mortem usually lands on chemistry. Wrong fit, wrong moment, bad luck. It is a comfortable story because it makes the failure unforecastable.
The research does not support it. Leadership IQ tracked more than 20,000 new hires and found that 46 per cent failed within 18 months, with 89 per cent of those failures traceable to attitudinal and behavioural factors rather than a gap in technical skill [1]. Heidrick & Struggles reviewed 20,000 of its own executive placements and found roughly 40 per cent of senior executives were pushed out, failed or quit inside the same window [2].
Those figures describe hiring in well-mapped territory, for roles the market has understood for decades. AI leadership is harder than that.
Why AI leadership breaks the usual assessment
Executive assessment runs on comparison. You look at what a candidate did before, you compare it to what has worked in similar seats, and you form a view. The method depends on a reference class: a body of prior outcomes stable enough to learn from.
AI leadership barely has one. MIT's State of AI in Business 2025 study reviewed more than 300 disclosed enterprise initiatives and found that around 95 per cent of generative AI pilots delivered no measurable return [3] [4]. Read that as a hiring signal rather than a technology story and it becomes uncomfortable. If nineteen in twenty initiatives produce nothing, then most candidates who can honestly say "I led AI at a large company" led something that did not work. The credential is real. The outcome behind it is not.
This is the trap. A CV in this field is a record of proximity, not of results. Proximity to a well-funded lab, to a flagship model, to a transformation programme that was announced loudly and quietly wound down. Boards reading those CVs against a 2015 mental model of what a strong technology executive looks like are reading a signal that has never been validated.
What track record actually means in a field under ten years old
The honest answer is that it cannot mean years, and it cannot mean citations. It has to mean decisions.
We look for four kinds of evidence, and we press hard on each one.
Something that survived production. Not a pilot, not a demo, not a proof of concept that impressed a steering committee. A system that stayed up while the data underneath it drifted, while latency budgets tightened, and while somebody had to explain an output to a regulator or a customer. Ask what broke and how they found out. Candidates who have genuinely operated a model in production answer with monitoring, evaluation harnesses and rollback procedures. Candidates who have not answer with architecture.
Something they killed. The most valuable line on an AI leader's record is a project they stopped. MIT's data is blunt on this point: buying from specialist vendors or building through partnerships succeeded roughly 67 per cent of the time, while internal builds succeeded about a third as often [3]. A leader whose instinct is to build everything in-house is not demonstrating ambition. They are demonstrating an expensive default. We want the candidate who can describe, in detail, the thing they chose not to build and why.
A judgement call that cost them something. Deprecating a model the business liked. Telling a founder the roadmap needed six more months. Escalating a fairness problem that delayed a launch. These are the moments that reveal whether someone will hold a line when the pressure is commercial rather than technical.
Evidence calibrated to your stage. A research director from a well-resourced lab and a first head of AI at a Series B company are not interchangeable, and the failure mode runs in both directions. The lab leader stalls when there is no data platform to build on. The scrappy operator stalls when the organisation needs governance, headcount planning and a five-year architecture. Stage-fit calibration is not a soft consideration. It is the single most common reason a technically excellent hire underperforms.
Vetting that tests judgement instead of pedigree
Interviews reward preparation. Most senior AI candidates have rehearsed their story, and a well-rehearsed story tells you almost nothing.
Three changes make the process diagnostic.
Give the candidate your actual mess. Not a case study, not a take-home exercise. The real problem your team argued about last quarter, with the real constraints, the messy data situation and the political dimension included. Watch how they frame it before they solve it. Strong candidates spend most of the conversation narrowing the problem. Weaker ones reach for an architecture in the first five minutes.
Ask them to argue against a decision they made. A candidate who can construct the strongest case against their own past choice is showing you calibrated confidence. A candidate who cannot is showing you that the decision was never really examined.
Test the translation layer. Ask them to explain a model failure to a hypothetical audit committee, then to a hypothetical customer. Communication is not a soft skill in this role. It is the mechanism by which technical risk reaches the people accountable for it. We covered the broader shape of this problem in The AI Leadership Readiness Matrix, which sets out why organisational maturity matters as much as candidate quality.
Reference work is where the assessment actually happens
Interviews are performed. References are observed. Most search processes waste them anyway, because they call the two names the candidate offered and ask whether the person was good.
Peer-level references carry more signal than managerial ones. A manager saw output. Peers saw process: how the person behaved when a launch slipped, whether they shared credit, what happened when their model failed in front of the business. We work through three to five peers per finalist, and we ask about failure before we ask about achievement.
Contribution claims need triangulating. Ask several references the same question about the same project, independently, and the gap between "I led" and "I was on the team" shows up quickly. This is not about catching people out. It is about establishing what a person actually owns, which is exactly what you are buying.
And ask stage-specific questions. "Would you hire them again" is nearly useless. "Would you hire them again for a company at this stage, with this data maturity, reporting to a board with this level of technical fluency" is the question that matters.
Regulatory literacy has become part of the risk
For anyone hiring into Singapore or the wider region, this is now technical due diligence rather than a compliance footnote.
Financial services candidates should be able to speak to the Monetary Authority of Singapore's FEAT principles on fairness, ethics, accountability and transparency without being prompted [5] [6]. Anyone leading generative AI deployment should know the IMDA Model AI Governance Framework for Generative AI and what it expects around testing and disclosure [7]. Anyone building for European customers needs a working grasp of the EU AI Act obligations phasing in across 2026 and 2027, because the design decisions that determine compliance are made long before legal reviews them [8].
There is a practical dimension too. Hiring a senior AI leader from outside Singapore means Employment Pass and COMPASS eligibility, which shapes both who can realistically start and when [9]. A search process that surfaces a perfect candidate who cannot be onboarded for five months has not reduced risk. It has moved it.
Speed and rigour are a sequencing problem
The standard framing puts speed and rigour in tension. Move fast and you skip diligence. Diligence properly and you lose the candidate to a faster offer.
We think the tension is mostly misplaced. Speed belongs in market mapping. Rigour belongs in assessment. Those are different phases and they draw on different capabilities.
Our AI-led mapping platform is built for the first half of that: covering a full market, including the senior people who are not looking and will never see a job posting, in days rather than weeks. That is where a traditional relationship-only network loses time, because it can only surface the people it already knows. Buying back three or four weeks at the front of a search means you can spend them where they compound, on structured assessment and real reference work, and still close ahead of a slower process.
What we do not do is compress the assessment. A shortlist delivered in 72 hours is a sourcing achievement. It is not a hiring decision.
The first 100 days belong inside the search
A search that ends at signature has optimised for the wrong milestone.
Three things materially reduce post-hire failure, and all three are cheap compared to a replacement search. Agree a 90-day milestone tied to a business metric rather than an onboarding checklist, because "reduce inference latency on the recommendation service by 30 per cent" tells you something that "completed orientation" never will. Confirm before the start date that the data access, compute budget and executive sponsorship the role requires actually exist, since new AI leaders more often stall on organisational friction than on technical difficulty. And run a structured review at 90 days while the situation is still recoverable. Misalignment caught at three months is a conversation. At nine months it is a severance negotiation.
The question boards should be asking
The useful question is not "is this candidate good". Almost every finalist in a senior AI search is good by some measure.
The question is: what evidence do we have that this person's judgement holds under our specific constraints, and how much of that evidence came from a source with an incentive to impress us? If the honest answer is a CV, two friendly references and a strong interview, the risk has not been assessed. It has been deferred to the first quarter.
Olofsson & Company works with founders, boards and CHROs across Singapore and APAC on exactly this problem: AI and technology leadership searches where the cost of getting it wrong is measured in lost quarters rather than lost fees. If you are opening a search where the reference class is thin, that is the conversation worth having early.
Frequently asked questions
How can an executive search partner reduce the risk of a bad AI hire?
By replacing credential-based assessment with evidence-based assessment. That means structured probing of production decisions rather than pedigree, peer-level back-channel references rather than nominated ones, stage-fit calibration against the hiring company's actual maturity, and staying engaged through a 90-day review so misalignment surfaces while it is still fixable.
Does a PhD or a strong publication record predict success in an applied AI leadership role?
For research-heavy roles in well-funded labs, it matters a great deal. For applied and commercial roles, it predicts very little on its own. Operational evidence, models running in production, evaluation systems built, projects deliberately stopped, is the stronger signal.
How long should a senior AI leadership search take?
Expect eight to fourteen weeks for a genuinely senior mandate in Singapore, and factor in Employment Pass timelines for overseas candidates. Faster mapping can compress the front end, but assessment and reference work should not be shortened. Related reading: what it actually costs to hire AI leadership.
Sources
- Leadership IQ, Why New Hires Fail
- PrimeGenesis on the Heidrick & Struggles study of 20,000 executive placements
- Forbes, MIT finds 95% of GenAI pilots fail
- State of AI in Business 2025, report summary
- Monetary Authority of Singapore, FEAT principles
- MAS FEAT principles, full document
- IMDA, Model AI Governance Framework for Generative AI
- EU AI Act implementation timeline
- Singapore Ministry of Manpower, Employment Pass eligibility and COMPASS
