Scorecard | Choosing AI

The same-job AI test

Benchmarks are somebody else's setup on somebody else's task. Give two AI tools the same messy job you actually have, track four things, and get an answer that is about your work.

Start here

One AI scored 62.7% and 99.9% on the same test

ARC Prize ran GPT-6 Astra on one of the hardest tests in AI. In a plain setup it scored 62.7%. With OpenAI's own memory and reasoning system wrapped around it, about 99.9%. Same model, both times.

The model is one part. Around it sits software that keeps notes, retries, and holds a long job together. Change that and the same AI behaves like a different one.

Which is why launch-week charts disagree with each other, and why none of them are about your work. This is the test that is.

One afternoon4 columnsAny two AI toolsNo technical setup
The test

Same job, two AIs, four numbers

Pick one real job you actually have this week. Not a puzzle, not a sample, not something you invented to be fair. The messiest recurring thing on your list is the best choice, because that is what you are actually buying an AI for.

1. Pick the job and write it down once

One paragraph describing what you want, what you are attaching, and what done looks like. You will paste this identical text into both tools. Writing it once is what makes the test fair.

2. Give it to both, same day, same files

Same wording, same attachments, same time of day. Do not help one more than the other, and do not rescue the one you are hoping wins.

3. Track four things while they run

Finished, rescues, time, cost. Definitions are below. Fill them in as you go, because you will not remember accurately afterwards.

4. Read the rescues column first

It is the one that predicts your week. A tool that gets 90% there but needs you three times is worse than one that gets 80% there alone.

The four columns

What each column means

Keep it on paper or in a note. Four numbers per tool. The whole point is that it takes ten minutes to read afterwards.

01
Finished

What to write down: how much of the job came back done. Not how good it looked, how much of it you did not have to do yourself. A fraction is fine: 3 of 4 sections, most of the sheet, all of it.

02
Rescues

What to write down: how many times you had to step in. Every correction, clarification, restart, and re-prompt counts as one. This is the column that matters most and the one people forget to keep.

Why this one wins

Finished tells you what it can do. Rescues tells you what it costs you. A tool you have to supervise is a tool you did not delegate to.

03
Time

What to write down: wall clock from when you sent it to when you had something usable. Include your rescue time. That is the number that reflects your actual afternoon.

04
Cost

What to write down: whatever the run consumed. On a paid plan that is usage against your limits. If you never came close to a limit, write down not close and move on.

Reading it

How to call the winner

Three rules, in order.

The thing this protects you from

Launch weeks are designed to make you feel behind. A benchmark chart is somebody else's setup on somebody else's task, and you saw at the top of this page how far apart two numbers for the same model can be.

Your four numbers are worth more than all of them, because they are about the job in front of you.

Bonus

3 jobs that separate AI tools fastest

If you do not have an obvious candidate, these three surface the differences quickest.