Essay 04

Before you measure AI, fix the work around it.

Some AI pilots do not fail because the model is bad. They fail because the work around it was never ready.

Because when an AI pilot does not create clear value, the answer is not always “the AI was not good enough”. Sometimes the answer is more boring.

The documents were messy. The permissions were wrong. Nobody owned the process. The policy was outdated. The team never agreed what good looked like.

Then we ask AI to perform magic on top of that.

And when the measurement is weak, everyone starts arguing about the tool.

AI exposes the state of the work.

This is one of the things I keep seeing.

AI does not only produce outputs. It reveals the quality of the system around the work.

If knowledge is duplicated across SharePoint sites, Teams chats, old PDFs, and one person’s private folder, Copilot will feel inconsistent. If an agent is connected to unclear policy pages, the agent will answer with confidence and still make people nervous. If a workflow has six exception paths nobody wrote down, automation will make that visible very quickly.

That is useful. Annoying, but useful.

It means the pilot can create value even before the final AI result is perfect, because it shows where the organization is not ready yet.

But we should measure that too.

Sometimes the first measurable AI outcome is not time saved. It is the discovery that the process was never clean enough to automate.

The boring foundations are part of the value case.

I do not think we should treat data, governance, and process ownership as separate from AI value.

They are part of it.

If you improve metadata, reduce duplicate content, clean permissions, clarify retention, add sensitivity labels, document decision rules, or define who owns a knowledge source, that work can make the AI result better. It also makes the human process better.

So when we measure an AI initiative, I want to see the foundation work in the same picture.

  • Which data sources were cleaned?
  • Which permissions were corrected?
  • Which policies were rewritten because the agent could not answer safely?
  • Which prompts became reusable work instructions?
  • Which review steps stayed manual because the risk was too high?
  • Which process owner now has a clearer decision rule?

Those are not side notes. They are part of the delivery.

A simple example

Imagine an internal HR policy agent.

The first demo looks good. People ask questions. The agent answers quickly. The dashboard shows usage. Everyone smiles a bit.

Then testing starts.

One answer cites an old policy. Another answer mixes two locations. A third answer is technically correct but written in a way nobody in HR would approve. Suddenly the project becomes less about the agent and more about the knowledge base.

That can feel like failure.

I do not think it is.

It is a measurement moment. You just learned that the organization cannot safely scale this agent until policy ownership, review, and publishing are clearer.

Weak claim: “The HR agent answered 1,000 questions.”

Better claim: “The pilot found 18 unclear policy pages, fixed 11, and reduced repeat HR questions by 22% after review.”

Measure readiness, not only impact.

For some pilots, the right decision is not scale or stop.

It is: fix the foundation, then measure again.

That is not as flashy as announcing ROI. Fine. I would rather have a slower honest result than a faster fake one.

A useful readiness score should include things like:

  • clear owner for the process,
  • clear owner for the data or knowledge source,
  • documented baseline,
  • quality review method,
  • security and access check,
  • known exception paths,
  • cost model,
  • decision rule before the pilot starts.

If half of that is missing, I would be careful with big value claims. Not because I doubt AI. Because I respect it enough to measure it properly.

The consultant problem

There is also a consulting trap here.

It is tempting to make the AI story sound cleaner than the work really is. A good slide wants a simple before and after. The real project usually has a baseline that was guessed, a data source that was half-ready, and a process owner who changed the scope halfway through.

That is normal.

I just want us to write it down.

Because that is how the next pilot gets better. Not by pretending the previous one was perfect.

What I would do next

Before measuring any AI use case, I would run a short foundation check.

One page. No drama.

What is the task? Who owns it? Where does the input data live? Can AI access it correctly? What is sensitive? What is the baseline? Who reviews the output? What would make us stop, change, retest, or scale?

If we cannot answer those questions, we can still experiment.

But we should be honest about what kind of experiment it is.

It is not yet an ROI case. It is a readiness test.

And honestly, that might be exactly what the organization needs first.