The Broken Economics of Demos and Pilots

Why the two rituals every hospital uses to evaluate software are the two least reliable sources of evidence available to it, and what we do instead.

18 min read

Pensieve

Title card for The Broken Economics and Business Model of Demo and Pilot: a glowing copper puzzle piece resting on an illuminated circuit board, beside the line that a demo is a claim about the future while a use case is an observation of the past.

You’ve got to start with the customer experience and work backwards to the technology.

— Steve Jobs, Apple Worldwide Developers Conference, 1997

A note before the argument

We build one thing. A single operating picture of a hospital: every bed, every admission, every consultant, every drug batch, every claim, described once and connected to everything it touches. We work on that problem and nothing else.

What follows is a criticism of how software of this kind is sold, worldwide. We have a commercial interest in the conclusion, so we have kept our own opinions to a minimum and put the evidence in front. Every number here comes from a named source, and is well cited.

Our position is that the demonstration and the pilot, the two instruments almost every hospital in the world uses to evaluate software, are the two weakest sources of evidence available. We think the reasons are structural, and that the same failure repeats across countries, health systems and decades. We would rather be argued out of this than sell into it quietly.

Two rituals

A hospital software purchase, in Chicago or Manchester or Riyadh or Bengaluru, passes through two rituals.

The first is the demonstration. A meeting room. A projector that takes four minutes to connect. A sales engineer who is the most skilled user of that product in the country. A database of two hundred fictional patients, every one of whom has a complete record, a correctly spelled name, a valid policy number and no duplicate. Registration to discharge in eleven minutes. Nine people nod. Someone asks about accounting integration. The answer is yes.

The second ritual arrives when the buyer is careful. The careful buyer understands the criticality, and asks for a pilot. One ward. Three months. Limited scope. Low risk.

The careful buyer is right about the demonstration. The evidence on pilots is where this gets uncomfortable, because it is far worse than most people assume, and it has been consistent for fifteen years.

What a demonstration can evidence

A demonstration establishes that software can complete one defined task, once, under controlled conditions, on data constructed for the purpose.

It cannot test the buyer’s own data, with its duplicate identifiers, its several spellings of the same surname and its years of accumulated legacy. It cannot test the fourth month, after the novelty fades and a ward quietly returns to paper. It cannot test two in the morning on a Sunday with one nurse on the floor. It cannot test a senior clinician who has decided not to use it. It cannot test four hundred outpatient registrations before noon. It cannot test whether a claim can be defended to a payer eleven weeks after discharge.

What it does test is interface polish, presenter fluency and the quality of the fictional data. Vendors respond rationally to the thing they are measured on, and so the industry has become extraordinarily good at all three.

The pilot is a demonstration with a longer runtime

The failure pattern was first documented at scale outside healthcare.

  • A widely cited McKinsey survey of connected-technology programmes found that fewer than 30 percent of pilots were beginning to scale, that 84 percent of companies had been stuck in pilot mode for more than a year and that 28 percent had been stuck for more than two.4
  • Practitioners named the condition pilot purgatory. The pilot is never cancelled. It is never finished either.4

The pattern has repeated with every technology wave since. MIT’s Project NANDA traced it step by step across enterprise adoption in 2025, and the collapse between one stage and the next is the whole finding: 60 percent of organisations evaluated a system, 20 percent reached a pilot, and 5 percent reached production.1

The same study reported that 95 percent of organisations obtained no measurable return against 30 to 40 billion dollars of investment, and attributed the failure to misalignment with daily operations rather than weak technology.12 One chief information officer quoted in the report had seen dozens of demonstrations that year and found one or two useful.1

Other datasets, gathered independently, land in the same place:

  • IDC, in research with Lenovo, found that 88 percent of proofs of concept never reach production: for every 33 launched, four went live.3
  • S&P Global reported in 2025 that 42 percent of companies had abandoned most of their initiatives, up from 17 percent a year earlier, with 46 percent of pilots scrapped between proof of concept and broad adoption.5
  • Gartner has projected that 60 percent of projects will be abandoned through 2026 where the underlying data is unready.5

Turned around and read as survival rates, those studies put the share of work that reaches production or scale at 12 percent, 5 percent, 36 percent, 30 percent and 54 percent. Different researchers, different industries, different decades. The figures measure different points in a lifecycle and are directional rather than strictly comparable, and they still converge.

Healthcare has its own diagnosis for this

Hospitals are not exempt. Global health researchers have written about this failure for fifteen years and gave it a clinical name: pilotitis.6

A 2020 paper in the Journal of Medical Internet Research describes how large sums went into digital health interventions across the global south, and how many failed to reach scale despite successful pilot studies.6 It identifies two forces. Technology succeeds when it supports the way people already work. And usage decays over time, a pattern researchers call the law of attrition, when a tool is designed without understanding the real gap at the front line.

A 2026 review in Frontiers in Digital Health describes the same syndrome across developing countries with unusual precision.7 A hospital runs a pilot for a year with outside support. The funding ends. Servers go offline. The trained staff move on. Priorities shift. The paper’s judgement deserves repeating: the difficulty is rarely proving that a technology works. The real difficult part is building an environment in which it can keep working.

Field data records the same outcome. Of 36 mobile health projects reviewed in sub-Saharan Africa across 2008 and 2009, 23 never advanced past the pilot phase.8 Two thirds of them died in the exercise that was meant to prove them safe.

The economics that produce this behaviour

Enterprise software is sold slowly and at high cost. Benchmark analysis by Ebsta and Pavilion across hundreds of thousands of business-to-business opportunities put the median win rate near 19 percent.12

If four pursuits in five are lost, every deal that closes must pay for the four that did not. Every demonstration delivered to a hospital that chose a competitor is recovered from the hospital that said yes. This business behaviour has two consequences.

Capital goes to selling. Project NANDA observed that over half the budget in the category it examined went to sales and marketing.1 Money spent on winning never reaches the product.

The product is designed for the wrong person. Expensive acquisition forces the vendor to close with the fewest possible decision-makers, which means the executive office rather than the ward. Software is shaped to impress the person who will never open it, and handed to the person who will live in it.

That second mechanism explains most of the unusable clinical software in the world. It is a procurement artefact rather than an engineering one.

What the buyer pays

The outcomes have been measured repeatedly by people with no stake in the answer. Zylo, working from tens of billions of dollars of software spend under management, puts licence utilisation near 54 percent, and average annual waste per organisation at close to 20 million dollars.11 The other 46 percent is bought, paid for, and never opened.

A 2023 systematic review in the Journal of the American Medical Informatics Association, examining clinical information system implementation in developing countries, cites failure rates for such systems of up to 70 percent.9

And the software that survives is not well liked by the people who carry it. Melnick and colleagues, writing in Mayo Clinic Proceedings, measured how physicians rate electronic health records using the standard System Usability Scale. The mean score was 45.9. For scale, Microsoft Excel scores 57 on the same instrument, and the average across more than 1,300 studies from other industries is 68. A score of 45.9 sits in the bottom 9 percent of those results and earns a grade of F.10

The same study found a steady relationship with burnout: every one point of better usability was associated with 3 percent lower odds of a physician being burned out.10 Poor software in a hospital is a clinical workforce problem with a measurable dose response.

Underneath all of this sits an asymmetry that buyers rarely price.

When a purchase fails, the vendor loses

Direct. A renewal.

Operational. One logo on a slide.

Institutional. Nothing.

When a purchase fails, the hospital loses

Direct. The licence fee and the implementation year.

Operational. Thousands of staff hours, and the parallel registers that never went away.

Institutional. The confidence of the people who were asked to change, and the credibility of the management that chose it.

The party with the smaller exposure controls the evidence.

The bill for a pilot, itemised

A pilot is sold as the low-risk option. Here is what it costs, and none of it appears on the quotation.

  • Staff run two systems at once for three months. Every registration is entered twice. Every discharge is reconciled by hand.
  • The best people are assigned to it, because nothing can be piloted with the weakest. The people who can least be spared lose their evenings.
  • The pilot ward is chosen for being easy, which means the result says very little about the difficult wards where the money is actually leaking.
  • The scope is deliberately small, so the parts that matter most — billing against clinical documentation, pharmacy against orders, occupancy against reality — are excluded precisely because they are hard.
  • When three months end without a verdict, which is the usual outcome, the institution has spent a quarter, learned little, and taught its own staff that new systems arrive, consume their time and disappear.

That last item is the most expensive, because it is charged again to the next project, and to the one after that.

India, where the clock makes this worse

Every health system faces this problem, but India’s situation is unique. The national digital health programme has introduced a strict deadline for a decision that already lacks strong evidence. Evaluating any system under a tight deadline is the worst possible approach.

The Ayushman Bharat Digital Mission has moved with real speed. In May 2026 the Ministry of Health announced that more than 100 crore health records had been linked to ABHA accounts, roughly double the figure from February 2025, with more than 450 health technology solutions integrated into the ecosystem.13 State authorities have issued compliance directives to private hospitals empanelled under PM-JAY, with de-empanelment as the consequence of failing to integrate.14 NABH cycles impose dates of their own. These mandates are legitimate and overdue, and the infrastructure being built is among the most ambitious in the world.

The effect on buying behaviour is predictable. A hospital that must be compliant by a stated date stops evaluating and starts procuring. The question in the room changes from whether this will work here to whether this can be live by March. Vendors understand this perfectly. A deadline-driven buyer is the most demonstration-susceptible buyer that exists, because there is no time left to test anything, and the only evidence available is the performance in the meeting room. A hospital that has already spent a quarter on an inconclusive pilot has less time remaining than when it started.

The pilotitis literature was written with India substantially in view. The 2020 JMIR paper argues the case using Indian examples, and it makes an observation any hospital owner will recognise: a minimum viable product is normal practice in the software industry and unacceptable in medicine.6 Speed does not have to displace rigour. The correction it proposes is a real-world testing environment rather than a longer showcase.

Working backwards from the hospital

Everything above describes an industry that starts from its own product and works forward, searching for an institution that will take it. The alternative is the one this piece opened with: start at the other end.21

Not from the budget, and not from the signature. From the operational problem that is costing the institution money and sleep this month. Every hospital leader already knows what it is. It is the discharge that takes four hours because one administration cannot be confirmed. It is the eleven percent of claims returned on documentation grounds. It is the pharmacy count that has never once matched the shelf. It is learning what happened last night three days later, from a report assembled by hand.

There is something worth saying about the people inside these buildings, because the industry that sells to them has been careless about it. The most sophisticated system in any hospital today is not software. It is the nursing superintendent who knows which beds are truly free and which are held for an admission arriving at four. It is the billing head who knows which payer delays and why. It is the store keeper who knows what the register claims and what is actually on the shelf. It is the quality manager who carries the institution’s standards in her head between assessments.

People invent workarounds because they are smart, and there are optimal ways to do things in much less time. Every trick shows where the software broke. People and their own workarounds are the actual design document in the building. That distributed model of the hospital is what actually runs it, and it walks out of the door at the end of every shift. When a senior person resigns, part of it is gone for good. The work is to put that model into the institution itself, so that the people are supported by it rather than serving as its only copy.

Both routes end in a decision. Only one of them supplies evidence before the decision is made, and on that route each completed use case makes the next one faster.

What we do instead

We are stating this with complete clarity so there is no room for ambiguity, and you can hold us strictly accountable to every word.

We do not demonstrate on sample data. There is no fictional patient anywhere in our process.

We do not ask for a pilot. No parallel entry. No three-month observation ward.

We solve one real problem first. Together we select the operational problem that is costing the most right now. We agree on the current number, the target, the method of measurement and the person who owns it. All of it is written down before any work begins. That is the point at which we make ourselves falsifiable.

We work on real data. We take a defined slice, usually the last full month, exactly as it is, with every gap left in. We do not clean it in private first and we do not hide what we find. The state of the master data is the real obstacle between an institution and every benefit any vendor has ever promised it. Keeping that data clean is a default layer in Pensieve.

We do the heavy work ourselves. Your staff keep working as usual, apart from some basic cooperation on agreed terms. Nurses treat patients and billing processes claims. Nobody gets homework, and no one enters data twice for us. We come to the floor on-site during actual hospital hours, including night shifts, where software is truly judged.

The output belongs to the hospital. The working solution, and everything learned about the data, stays with the institution whether or not it buys anything from us.

At the end of it there is no promise about the future. Something has already changed in the building. If the number moved, that is evidence. If it did not, that is evidence too, and it was obtained in days rather than in a quarter.

The sequence is documented at scale

Tampa General Hospital began with a single use case on Palantir’s platform in 2021 and expanded to more than a dozen across the health system.15 The outcomes below are published by the hospital, in its own press releases of 2024 and 2025, and are customer-reported rather than independently audited.16

  • Patient placement time: down 83 percent.
  • Post-anaesthesia care unit holds: down 28 percent.
  • Mean length of stay for sepsis patients: down 30 percent.

Palantir applies the same logic at the front of the relationship. Its bootcamps take a customer from zero to a working use case in one to five days, on the customer’s own data and the customer’s own problem.17 Reporting places conversion from those sessions to signed contracts near 75 percent, against an industry norm of six to twelve month cycles and roughly one win in five.1812 The lesson is about what counts as evidence rather than about speed.

The engineers we send

This is the part that is difficult to copy, because it requires having someone to send.

Most vendors have sales engineers, who demonstrate, and implementation partners, who configure and then leave. When real data breaks the standard workflow, and it always does, the available responses are a change request, a quotation and a place in a queue.

We send forward deployed engineers. They work directly with end users to understand their needs, design and build product features, and deploy software in the field.19 They embed with the customer, work on real data, write production code and own the outcome. This model has spread across the enterprise technology industry for a simple reason.20 It works where nothing else does.

Ours are engineers of the standard we would put on our core platform. That is deliberate and it is expensive.

What one of them actually does looks like this. She sits in the billing office for a day or two and watches a claim being assembled, and finds the four undocumented workarounds the team invented in 2019 that no requirements interview would ever have surfaced. She stands at the nurses’ station at two in the morning, because that is when the system is tested hardest and staffed thinnest. She figures out the way in Pensieve, against real data, and shows the result to the charge nurse before lunch. When a consultant says the order screen costs him ten minutes he does not have, she fixes it with Pensieve that afternoon and asks him to try again.

She carries the hospital’s problem the way the hospital carries it. She is measured on whether the number moved, and on nothing else.

There is a second thing she does, and it matters more than it appears. Everything learned in one hospital returns to the platform. Where a solution built for one institution proves general, the durable part becomes product, and every hospital that follows inherits it. The field is the research function, and forward deployed engineering therefore operates as a product function rather than as a service wrapped around one. Every hospital that follows inherits what yours taught us.

This is why we treat it as part of the product and price it that way. A platform that requires an engineer to reach production, and does not include one, is a construction kit with an invoice attached.

Hospitals have been sold software for decades by people who have never watched a discharge. We would rather send the people who build the product to stand where the work happens.

What this costs us

We would rather write this here than have it discovered later.

The model is expensive per hospital, because we spend on engineers instead of presenters. It limits how many institutions we can take on in a quarter, which means we sometimes have to ask a hospital to wait. We uncover issues early, precisely when standard vendors project false confidence. When honest results show you should buy less or walk away, we tell you directly, even at the cost of our own contract. You do not like mess, and neither do we.

We accept that. Comparable settings show failure rates approaching 70 percent,9 and a failure costs everyone far more than a lost sale. We would rather lose a hospital in week two than in month fourteen. A company that grows by expanding inside institutions where its software demonstrably works spends far less on winning new ones. The arithmetic is sound at our size and it stays sound at a hundred times our size.

The standard we ask to be held to

One question. It works on every vendor in this market, and it should be asked of us first.

Show me this working on my data, from my last month, with my staff in the room, against a number we agree beforehand.

If the answer is a longer demonstration, a larger demonstration, a demonstration with more modules, or a three-month pilot in the easiest ward, that is useful information about the vendor.

If the answer is yes, evidence is about to replace a promise.

Peer-reviewed findings (Mayo Clinic Proceedings, JAMIA, JMIR, Frontiers in Digital Health) are distinguished throughout from analyst and vendor datasets (McKinsey, IDC, Gartner, S&P Global, Zylo, Ebsta and Pavilion) and from customer-reported outcomes (Tampa General), which are directionally useful rather than independently audited.

References

  1. The GenAI Divide: State of AI in Business

    MIT Project NANDA2025mlq.ai

    Source for the evaluation-to-production collapse, the 95 percent no-return finding, and the share of category budget going to sales and marketing.

  2. MIT report: 95% of generative AI pilots at companies are failing

    Fortune2025finance.yahoo.com

    Press coverage of the NANDA findings.

  3. The enterprise AI pilot purgatory problem: what the statistics actually tell us

    SoftwareSenisoftwareseni.com

    Summarises IDC research with Lenovo on the share of proofs of concept that reach production.

  4. What is pilot purgatory, and why every start-up should know this term

    Carl Vausecarl-vause.medium.com

    The widely cited McKinsey survey data on connected-technology pilots, and the origin of the term.

  5. Four AI failure modes

    Writer2025writer.com

    Compiles the S&P Global 2025 abandonment data and the Gartner projection through 2026.

  6. Regulatory Sandboxes: A Cure for mHealth Pilotitis?

    Journal of Medical Internet Research2020jmir.org

    Peer reviewed. The pilotitis literature, the law of attrition, and the argument against a minimum viable product in medicine.

  7. From pilot to policy: why AI health interventions fail to scale in developing countries

    Joseph JFrontiers in Digital Health2026pmc.ncbi.nlm.nih.gov

    Peer reviewed. The anatomy of a pilot that ends when its funding does.

  8. Field data on mobile health projects in sub-Saharan Africa

    2013d-nb.info

    Citing Lemaire 2011 and Tomlinson et al. 2013: of 36 projects reviewed, 23 never advanced past the pilot phase.

  9. Clinical information system implementation in developing countries: a systematic review

    Journal of the American Medical Informatics Association2023academic.oup.com

    Peer reviewed. Cites failure rates for such systems of up to 70 percent.

  10. The Association Between Perceived Electronic Health Record Usability and Professional Burnout Among US Physicians

    Melnick ER, Dyrbye LN, Sinsky CA, et al.Mayo Clinic Proceedings2020mayoclinicproceedings.org

    Peer reviewed. Mean System Usability Scale score of 45.9, and the dose response with burnout.

  11. SaaS Management Index and shelfware analysis

    Zylozylo.com

    Licence utilisation near 54 percent, from tens of billions of dollars of software spend under management.

  12. Sales cycle length guide

    ORM Technologyorm-tech.com

    Reports the Ebsta and Pavilion B2B benchmark data on median win rates.

  13. Ayushman Bharat milestone: 100 crore health records linked with ABHA accounts

    IANS2026-05ianslive.in

    Ministry of Health and Family Welfare statement on ABDM crossing 100 crore linked health records.

  14. ABDM compliance for AB-PMJAY empanelled hospitals

    Nirmitee2026nirmitee.io

    Reporting on the compliance directives issued to empanelled private hospitals, and the consequence of failing to integrate.

  15. TGH selects Palantir AI software for connected care coordination

    Tampa General Hospital2024-06tgh.org

    Customer reported, not independently audited.

  16. Tampa General Hospital honored with Healthcare Innovation Award by the Joint Commission

    Tampa General Hospital2025-09tgh.org

    Customer reported, not independently audited.

  17. Deploying Full Spectrum AI in Days: How AIP Bootcamps Work

    Palantir Technologiesblog.palantir.com

    The vendor account of taking a customer from zero to a working use case on their own data in one to five days.

  18. Who gets an FDE and who does not: the great B2B AI debate right now

    SaaStrsaastr.com

    Analysis of the bootcamp model and its reported conversion rate.

  19. Forward deployed engineer: the role guide

    Netgurunetguru.com

    Overview of the role and its origin, including the description of the mandate.

  20. Forward deployed engineers

    The Pragmatic Engineernewsletter.pragmaticengineer.com

    On how the model has spread across the enterprise technology industry.

  21. Working backwards

    John Gruber1997daringfireball.net

    Transcript reference for the Steve Jobs remark at the Apple Worldwide Developers Conference, 1997.