A working replica · blank by design

The AI Coworker Scorecard

The scoring instrument from the hiring framework described in We Fired an AI Coworker. Then We Hired It Back. Same categories, same weights, same decision bands we run internally. Score any AI initiative your team has deployed, or is about to.

Before you score

What would a human hire in this role be measured on?

This is the question the whole framework hangs on, and it gets answered before anything gets built, let alone scored. Name the two or three numbers a manager would put in a human's performance review for this exact job.

If you cannot answer this, stop here. In our framework the rule is: no KPIs, no build. The tool-version of that rule is gentler but the same shape: no measures, no score. Write the job description first.

Scoring is locked until the question above has an answer.
The four categories

Score each from 0 to 100

Score against evidence, not vibes. The anchors under each slider describe what the ends of the scale look like in practice.

Outcomes

Business KPIs · ×0.40contributes 0.0

Performance against the targets written into the job description before the build. Hours given back, cost saved, cycle time cut, errors prevented, revenue influenced. This is the category the others exist to serve, which is why it carries the most weight.

0
No measurable movement on any KPIAt or above every target, and trending up

Quality & Trust

×0.25contributes 0.0

User satisfaction, accuracy on sampled interactions, critical incidents, compliance. Trust is what earns a coworker its next, bigger job. It is also the thing that, once lost, quietly ends the whole deployment.

0
Users double-check everything it producesHigh sampled accuracy, zero critical incidents

Adoption & Engagement

×0.20contributes 0.0

Active users against the intended population, interactions per active user, share of work handled end-to-end. A capable coworker nobody uses is not delivering value; it is drawing a salary.

0
Novelty-shaped usage: a spike, then a fadeMost intended users return monthly, unprompted

Ops Readiness

Governance · ×0.15contributes 0.0

A named human manager, an escalation path, a standing QA cadence, access controls, stability. The least glamorous category, and the one whose absence explains most failures we have run the autopsy on.

0
Making it succeed is nobody's actual jobA named manager runs a standing review cadence
The verdict

Weighted score & recommendation

0 / 100
LockedAnswer the measures question above to unlock scoring.
InternApprenticeFull-Time

Where the score comes from · slot width = category weight

20406080
Outcomes0.0 / 40Quality & Trust0.0 / 25Adoption0.0 / 20Ops Readiness0.0 / 15

ScoreRecommendationWhat it means
80–100PromoteExpand the role: wider rollout, bigger scope, or the next stage.
60–79KeepOn track. Hold the cadence and keep managing.
40–59Improve planName the gap, set a re-score date, fix it deliberately.
20–39DemotePull it back a stage and narrow the scope while gaps get fixed.
0–19RetireLet it go. Document the lessons; revisit when conditions change.

The real instrument ends with a manual-override row, and so does this one, conceptually: the score recommends; a human decides. Bring the number to the review meeting as the start of the conversation, not the end of it.

What this deliberately does not do

  • It does not measure your KPIs for you. The score is exactly as honest as the evidence behind each slider.
  • It does not replace the stage-gate reviews. Each stage (Intern, Apprentice, Full-Time) has its own exit criteria; this is the cross-stage scoring layer that sits on top.
  • It does not know your context. A 55 with a clear improvement plan can be a better position than a 75 nobody owns.
  • It does not send your inputs anywhere. Nothing you type here leaves this page.

This instrument comes from a real program: we hire AI the way HR hires people, complete with job descriptions, probation, and the occasional termination. The full story, including the coworker we fired and rehired, is here.

Robin's Notebook

A new entry every couple of weeks. No promotion, no funnel, no manifesto. Just the unfiltered version.

Subscribing opens beehiiv.com in a new tab to confirm. Unsubscribe anytime; I won't share your email.