AI at Work

The same spreadsheet, two assistants, a six-figure miss.

A published test on everyday Microsoft 365 work. Same file, same prompts, answers written down before either tool opens it.

Method

Stated first, because the result is only as good as the method.

  • The same file and the same prompts for each assistant.
  • The defects in the file and the correct answers are documented before either tool opens it.
  • Outputs are scored blind by a third model with no commercial stake in the result.
  • Every round is published, including the ones that go against the tool we would prefer.

Three rounds so far.

Round 1 · Outlook

Claude 4–1

Won on judgement.

Round 2 · Word

Copilot

Claude wrote the better output, 4–2, and lost the round on speed. Read this one first.

Round 3 · Excel

Claude 8.5–6.4

Claude ahead on five of six tasks.

−€531,000, in a table that looked fine.

The finding that travels.

In Round 3, Copilot returned a tidy, sorted regional table. The regional total was understated by €531,000 because the workbook spelt London three different ways, and the assistant counted them as three places.

Nothing about the output looked wrong. That is the point. An assistant working in a spreadsheet produces a plausible number, not a guaranteed one. Where a figure has to reconcile, calculate it with the deterministic tool and use the assistant for the work around it.

Your data is messier than a test workbook.

Next round

A fourth round is in preparation. If there is a task you would like to see tested, send it to hello@acuityai.co. The test is run by Ger Perdisatt; the earlier rounds were first published by Acuity AI Advisory.