Microsoft AI is everywhere.

Let’s measure what actually changed.

A practical blog about Copilot, agents, Foundry, Power Platform, and the awkward question I keep coming back to: according to what?

This is for people who are excited about AI, but still want proof that survives a real meeting. I do not have the perfect method. I am trying to build one in public.

Start here

New essays

Essay 01

Your AI Pilot Was Successful. According To What?

The first post in the series. AI value, measured, and slightly annoying in the most useful way.

Open essay

Essay 02

Usage Is Not Value

Adoption matters, but the real question is whether the work changed.

Open essay

Essay 03

Measure Agents by Outcomes

A sharper way to measure agents than chats, sessions, and demo applause.

Open essay

Essay 04

Before You Measure AI, Fix the Work Around It

Some pilots are really readiness tests for data, ownership, policy, access, and review.

Open essay

Series

What this blog will test

01

Measure the work

Stop measuring “Copilot” as a vague object. Measure proposals, HR questions, document reviews, support handoffs, and customer prep.

02

Separate usage from value

Adoption matters. It still does not tell you whether quality improved, time turned into capacity, or risk went down.

03

Test boring use cases

The best AI value might hide in ordinary work nobody wants to put on a conference slide.

04

Measure agents properly

Conversation count is easy. Completed task, escalation rate, review time, and cost per outcome are more interesting.

05

Watch the review tax

AI can create drafts quickly and still move effort into checking, correcting, and explaining.

06

Follow the freed time

Saving time is not the same as creating capacity. Measure what people do with the time they get back.

07

Check the foundations

Data quality, permissions, policies, ownership, and governance can make or break the AI result.

Use cases

Examples with numbers, doubts, and next measurements.

Start with three practical cases: customer meeting briefs, an HR policy agent, and invoice triage. They are examples for now. Real evidence comes next.

Template preview

The AI Measurement Toolkit

A practical way to turn an AI pilot from “people used it” into stronger evidence about value, quality, time, cost, risk, and scale.

  1. 1Calculate monthly net value and break-even.
  2. 2Score whether the pilot is ready to measure.
  3. 3Generate a measurement plan.
  4. 4Create an experiment memo.
  5. 5Track quality, risk, cost, and adoption depth.
  6. 6Decide whether to scale, change, retest, or stop.
Open the measurement toolkit

Official publications

Useful research before making value claims.

Start with Microsoft WorkLab and Microsoft Research when you need real numbers or careful language about Copilot value.