Case studies of AI-driven transformation — question bank
37 standalone questions on nine real deployments — a public reversal, a legal precedent, a $62m failure, and how to read any case study without being fooled.
Audience: senior business leaders. No technical background assumed.
Source: the case-studies handout — nine real deployments including a public reversal, a legal precedent and a $62m failure, plus the six questions for reading any case study without being fooled.
How to use this: every question stands alone. Pick an option, then read the answer. The cases run in order, then the two pattern lists, then the interrogation method — which is the most durable part.
A · Step 1 — A reversal, in full
Q1A fintech launched an AI customer-service assistant and published month-one figures: 2.3 million chats, work equivalent to 700 agents, 67% of conversations automated, resolution in under 2 minutes against 11 for humans, and "on track to add $40 million" in profit. Which of those is not a result?
Answer: (3)"On track to add" is a forecast, and it was repeated globally for years as though it were an achievement. Projections get quoted as results, and this is the most-cited example of it.
Q2Fifteen months later the company reversed course and began rehiring human agents. Which of these was not among the stated reasons?
Answer: (3)Cost was not the problem — the technology kept working. What broke was quality on the tail of hard cases, and the governance exposure of letting a model handle disputes on its own.
Q3The most important sentence in that case is that satisfaction dropped even when the AI's answer was factually correct. What does that establish?
Answer: (2)No accuracy metric would have caught this, which is precisely why accuracy is a dangerous single measure. A system can be improving on the metric you watch while degrading the thing you actually sell.
Q467% automation and two-minute resolution are averages. Why did the damage sit outside them?
Answer: (2)The 33% that were not automated were the entire problem. An average is a summary of the cases you already handle well.
Q5The case is widely written up as an AI failure. Why is that reading wrong?
Answer: (2)The reversal is the admirable part. And the question it leaves for the room is sharper than the case itself: if your AI deployment were quietly damaging your best customer relationships, how long would it take you to find out — and would anyone be rewarded for saying so?
B · Step 2 — Adoption is engineered
Q6A wealth manager reached over 98% adoption among adviser teams — against an industry norm where most enterprises never get past pilots. What did they do differently?
Answer: (2)Not because the research was secret, but because finding it was too slow. The AI did not add a capability; it removed a friction everyone already resented — so adoption did not need to be sold.
Q7Document access went from 20% to 80%. Why is that number unusually useful in a case study?
Answer: (2)"40% faster" is meaningless without a starting point. Stating a humiliating baseline is a mark of honesty, and it is rare enough to be worth pointing at.
Q8Their meeting-notes tool drafts, and the adviser reviews and edits before anything is finalised. In a regulated business, what is that?
Answer: (3)Keeping the human at the point of consequence is not caution slowing things down; it is the thing that allowed the deployment to happen. The transferable rule: your first deployment should attack something your people already complain about — enthusiasm you do not have to manufacture is worth more than any change programme.
C · Step 3 — The right shape for a first major deployment
Q9An airline's virtual agent handles about 40,000 queries a day across 1,300+ question types, with 13 million+ conversations resolved. Which feature makes this the right shape for a first major deployment?
Answer: (2)Answering known questions about bookings, refunds and schedules — not "transform the airline." Fare rules and baggage policies are documented, so this is retrieval against a knowable corpus. A wrong answer escalates to a human and nobody is harmed.
Q10The airline reports a 97% success rate. What is the right follow-up question?
Answer: (2)It is the company's own definition of success, and companies define resolution generously. Percentages hide scope: resolved out of attempted queries, or out of all contacts including those that bounced to a human?
Q11Contrast that airline case with the oncology-recommendation failure. What is the lesson?
Answer: (2)One task was narrow, bounded, documented and recoverable. The other was unbounded, contested, life-critical and had no clean ground truth. The technology did not decide the outcome; the task selection did.
D · Step 4 — When the opponent adapts
Q12A bank deployed an agentic system that detects emerging fraud patterns in payments data and generates the detection rules to intercept them. What makes this case different from the three before it?
Answer: (2)This is the adversarial case: rules lose to adversaries, because a fixed rule tells the attacker exactly where the wall is. Note what the AI is actually doing — not just detecting fraud but writing the rules, which closes the loop between pattern discovery and defence and compresses a cycle that used to take analysts weeks.
Q13The bank also sends around 40,000 proactive customer warnings a day. What does that third line represent?
Answer: (2)Hard rules for the certain cases, a model for the suspicious remainder, and customer warnings as a third line. When your problem has an adversary, "we deployed a model" is not an answer. The question is how fast your loop turns relative to theirs.
E · Step 5 — When the training *is* the deployment
Q14A large bank launched a GenAI Academy training 35,000 staff across 15+ programmes. Why does this case belong in a leadership module rather than a technology one?
Answer: (2)It is the 70% of the 10-20-70 rule, funded and staffed, in an organisation of that scale. The sequencing signals what the leadership believed the hard part actually was.
Q15Among that bank's reported figures — sub-90-second query responses, sub-200-millisecond fraud scoring, and false positives down by one-third — which is the impressive one?
Answer: (3)That is the metric a customer actually feels, and the one that used to make fraud systems hated internally. A system that flags fewer legitimate transactions while catching more fraud is doing the genuinely hard thing — the two usually trade against each other.
F · Step 6 — Portfolio, not project
Q16A bank runs 450+ AI use cases in production, with a stated ambition of 1,000. What does that number actually require?
Answer: (2)Almost every organisation runs between two and ten initiatives, treating each as a project with its own business case, vendor and integration. You cannot run 450 that way. The 450th use case cannot possibly have cost what the 1st did — that is what a platform buys, and it is the only way the arithmetic works.
Q17That same bank reports fraud false positives cut by about 50% and detection up about 30%. How should those be treated?
Answer: (2)The handout flags them explicitly. Being clear about which numbers are audited, which are company-stated and which are forecast is the discipline the whole document is teaching.
G · Step 7 — The unglamorous 70%
Q18One pharmaceutical company appears in the list despite having the fewest hard numbers. Why?
Answer: (2)The approach was not to procure a tool and announce it, but to ground a generic assistant in company context and then run an adoption programme — proactively surfacing and answering employee doubts rather than waiting for resistance to appear.
Q19What outcome did that approach avoid, and what is the term used for it?
Answer: (2)And the mechanics transfer: champions give peer credibility (people believe a colleague in their own function over a central team); forums surface failures early; governance built in from the start prevents legal killing it three weeks before launch; celebrating specific uses turns abstract encouragement into a copyable example. None of that is technology. All of it is why the technology got used.
H · Step 8 — You own what your AI says
Q20A passenger was told by an airline's website chatbot that he could claim a bereavement discount retroactively. The airline's published policy said the opposite. What was the airline's legal defence?
Answer: (2)The tribunal's response is worth reading aloud: "This is a remarkable submission. While a chatbot has an interactive component, it is still just a part of Air Canada's website. It should be obvious to Air Canada that it is responsible for all the information on its website."
Q21Damages were CAD $812.02. Why is a trivial sum one of the most important cases in the set?
Answer: (2)The company's real policy was published. The chatbot contradicted it, and the chatbot won — against the company.
Q22How does that precedent scale?
Answer: (2)The exposure question the room should sit with: if your customer-facing AI made a wrong promise to 10,000 customers this month, who would find out, and when?For most, the honest answer is "when the complaints arrive" — which means the exposure is already priced in before anyone knows.
I · Step 9 — The $62m plumbing lesson
Q23A cancer centre and a technology company spent $62 million over roughly five years building an oncology recommendation system. How many patients were treated using it?
Answer: (3)Five years, $62 million, zero patients. And almost none of the failure was about the quality of the AI.
Q24What effectively killed it?
Answer: (2)"It did not fail at medicine. It failed at plumbing." Set it directly against the wealth manager, whose headline achievement was moving document access from 20% to 80% — integration. Same era, same category of technology, opposite outcome, and the difference was not model quality.
Q25What was the second, governance failure in that case?
Answer: (2)Scope also expanded repeatedly while value remained unproven, and a university audit found procurement irregularities, cost overruns and delays. A body that can kill things is exactly what was missing.
J · Step 10 — The two patterns
Q26Which feature did every success in the set share?
Answer: (2)Along with enormous volume so small gains compound, a measurable and often embarrassing baseline, integration into the system of record, a human at the point of consequence, and investment in adoption rather than just the tool.
Q27Which feature did the failures share?
Answer: (2)With scope expanded before value was proved, averages that concealed a damaging tail, no owner willing or able to stop it, the vendor's roadmap substituted for a plan, and accountability assumed to sit elsewhere.
Q28What is conspicuously absent from both lists?
Answer: (3)Not one of these outcomes turned on whether the underlying AI was good enough. That is the 95%-of-pilots-fail finding arriving again through nine independent doors.
K · Step 11 — Reading a case study without being fooled
Q29You are shown a case study. What is the first of the six questions?
Answer: (2)Vendor case studies select for success and are approved by the customer's marketing team. Several of the best-known cases in this set originate from the AI vendor's own website — which does not make them false, but tells you which direction the errors run.
Q30"Our AI achieved a 40% improvement." Which of the six questions bites hardest?
Answer: (2)"40% faster" is meaningless without the starting point. The 20%→80% example is the counter-model: it states where it began, which is why it can be believed.
Q31Why does "how long after launch?" matter so much?
Answer: (2)Novelty, careful launch cohorts and the easiest cases all flatter an early number. The sequel is where the truth is, and almost nobody looks for it.
Q32What is the meta-question, after the six?
Answer: (2)If your evidence base is entirely composed of successes, you are not looking at evidence — you are looking at a marketing category. And the finding that ~95% of pilots produce nothing tells you exactly how much of the real distribution you are being shown.
L · Step 12 — Interrogate four claims
Q33"Our AI reduced customer service costs by 40%." What is the single question to ask?
Answer: (2)Cost falls immediately and churn appears two quarters later. This is precisely the arc of the fintech reversal, expressed as a question you can ask in the meeting where the claim is made.
Q34"After deployment, our sales team's productivity rose 3×." What do you ask?
Answer: (2)Usually it means "3× more outreach" — activity, not outcome. If revenue did not move, you have automated the wrong thing very efficiently.
Q35"The model is 94% accurate — better than our human reviewers at 91%." What is the strongest question, and why is it the least obvious?
Answer: (2)Human errors are usually random; model errors are usually systematic — the same subgroup, every time. 94% that is always wrong about the same customer segment is far worse than 91% scattered. An aggregate accuracy number cannot tell you whether your errors are concentrated on the people who can least afford them.
Q36"We've deployed AI across 30 processes company-wide." What do you ask?
Answer: (2)Deployed ≠ used. This is the pilots-versus-production ratio compressed into a single question — and it usually converts an impressive number into a much smaller one, in front of the person who quoted it.
Q37Three questions to take away. Which is not one of them?
Answer: (4)Copying a published deployment means copying the marketing version of it. The other three all point at the same discipline: look for the tail, the detection lag, and the incentive to stop — none of which appears in any case study you will be shown.