Essay 01

Your AI pilot was successful. According to what?

I love AI. I just do not want us to clap because the dashboard went up.

This is a question I have been thinking about more and more.

Because we are doing a lot with AI right now.

Microsoft 365 Copilot is being rolled out. Agents are popping up everywhere. We are building things in Copilot Studio, Microsoft Foundry, Power Platform, and Azure AI.

And honestly, I love it.

I really believe we are only at the beginning of what we can do with AI.

But I am also that person who, when someone says “our AI pilot was a success”, immediately thinks: nice. But how do we know?

Maybe that sounds annoying. Fair. I can live with that.

Usage is a signal. It is not the whole story.

Let us say we give a group of employees Microsoft 365 Copilot.

They use it. Feedback is positive. Adoption looks good. People tell us it saves them time.

Great. Seriously. Those are good signs.

But now I want to know what happened next.

What actually changed?

An agent used 5,000 times is not automatically more valuable than one used 100 times.

Maybe those 100 uses saved a specialist from doing something manually that normally takes ages. Maybe the 5,000 uses were helpful, but only saved a few seconds each.

Both can matter. You just cannot tell from the usage number alone.

Time saved still needs a second question.

Microsoft has researched this. Early Copilot research found users were faster across common tasks such as searching, writing, and summarizing. Other research across thousands of workers found measurable improvements in document completion speed.

I love seeing numbers like that.

But they make me even more curious.

What happened with the time we got back?

Did we get more work done? Did the quality improve? Did someone finally have time for the thing that has been sitting on the to-do list since February? Did people spend a bit less time doing boring repetitive work?

Or did someone simply have a less hectic afternoon?

And actually, that last one sounds pretty valuable to me too.

The boring foundation matters too.

I know. Governance is not the exciting part of the demo.

Still, it matters.

If the data is messy, permissions are too broad, old documents are everywhere, and nobody knows which source is trusted, Copilot will feel worse than it should. Same for agents. Same for anything built on top of your content and processes.

So when we measure AI, I think we also have to measure the conditions around it.

Was the data ready? Were policies clear? Was sensitive information handled properly? Did people know when not to trust the answer?

Not glamorous. Very real.

What I want to explore

I am not trying to prove AI does not work. Good luck convincing me of that anyway.

I want to get better at showing where it works. And where I am wrong about it.

And when it works, what actually changed.

Using real use cases on the Microsoft AI stack:

  • Microsoft 365 Copilot
  • Copilot Studio
  • Microsoft Foundry
  • Power Platform
  • Azure AI

I want to test things. Measure things. Probably discover that some of the things we thought would be amazing are not.

And hopefully find some boring little use case nobody gets excited about that turns out to be ridiculously valuable.

That is the fun part.

Less: “We launched AI and 72% of people are using it.”

More: “Look what changed.”

Because if we can get better at proving where AI creates real value, we can make much better decisions about where to use it next.

And that gets me excited.

So the next time someone tells me, “our AI pilot was successful”, I will probably smile.

And then be that annoying person again:

Cool. According to what?