Microsoft Project Online retires September 30, 2026, migrate to a modern platform before it's too late.Start migration
Back to BlogShould AI Agents Mark Work Complete? The Real Test
PMO

Should AI Agents Mark Work Complete? The Real Test

Should AI agents mark work complete? Only when the result is objectively checkable, the check is automatable, and a wrong close is cheap to undo.

Onplana TeamAugust 17, 20264 min read

An agent that closes its own tasks removes the one checkpoint most PMOs still have left, and most teams grant that permission before deciding whether the specific task deserves it. The instinct is understandable: reviewing everything an agent does defeats the point of delegating it. But the fix is not "review everything" versus "review nothing." It is a test that tells you, task type by task type, which side of the line a given piece of work belongs on.

The direct answer: should AI agents mark work complete? Only when a task passes three checks: the result is objectively checkable (there is a fact, not an opinion, that determines whether it is done), the check itself can be automated (a human does not have to eyeball it every time), and the cost of a wrong close is recoverable (undoing a mistake is cheap, not a scramble). A task that fails any one of the three should be moved to review, not closed, which is the compromise that works for most teams: an agent may advance work to the point a human signs off, but never past it.

Should AI Agents Mark Work Complete? The Three-Part Test

Run every recurring task type through all three checks before deciding its default. A task only earns autonomous closing when it clears all three, not a majority.

Check The question Passes Fails
Objectively checkable Is there a fact, not a judgment call, that determines "done"? A test suite passes; a file was uploaded; a field matches a target value "The report reads well"; "the design feels right"; anything a reviewer would phrase as an opinion
Automatable Can the check itself run without a human watching it happen? A script confirms the output; a status field can be diffed against an expected value The check requires reading the output for tone, nuance, or stakeholder-specific context
Recoverable Is a wrong close cheap and fast to undo? Reopening a task and rerunning it costs minutes The task fed a decision already made, a report already sent, or money already moved

A task that clears all three, updating a status field from a source system, running a scripted data check, filing a routine intake request, is a reasonable candidate for autonomous closing. A task that fails even one, drafting a stakeholder-facing status report, reallocating budget, making a call that depends on reading the room, should stop at review regardless of how well the agent has performed on other work.

Why "Move to Review" Beats "Mark Done"

The compromise that holds up in practice is narrower than most autonomy debates suggest: an agent can move work all the way to the edge of done, drafted, evidence attached, ready for a decision, without ever making the decision itself. In Onplana, work an agent completes through Run with Agent lands in a review inbox as a draft with the evidence attached, not as a committed change, so a bad output costs a rejection click rather than a cleanup project. That single design choice is what makes the three-part test safe to apply liberally on the "review" side and conservatively on the "auto-close" side: the downside of getting a task wrong is bounded by design, not by how carefully anyone watched the agent work.

The diagram below shows how the three checks route a task to one of two outcomes.

Should AI agents mark work complete: the three-check decision tree A task finishes Objectively checkable? Check automatable? Wrong close recoverable? All three pass Agent may close it Any check fails Move to review, never past it

What Should Never Be Closed by an Agent

Some task types should stay on the review side of the line regardless of how clean an agent's track record gets, because the risk they carry is not the kind more data resolves. Three categories belong here by default: anything judgment-heavy, where "done" depends on reading a stakeholder's intent rather than checking a fact; anything expensive or slow to undo, a budget reallocation, a schedule change with downstream dependencies, a stakeholder-facing status report already sent; and anything where the check itself would require the same judgment call the task did, which makes the automatable test impossible to pass honestly rather than merely inconvenient to build. A task can fail this category even after months of a clean record, because the category is about what a wrong close costs, not about how often the agent has been right so far.

The Line Moves, and It Should

The three-part test is not a one-time classification. Onboarding an agent already treats scope as something that widens or narrows on evidence, and closing permission should follow the same pattern: a task type that starts on the review side of the line can earn autonomous closing once its review record shows a long run of clean, uncorrected outputs, the same evidence a day-five checkpoint already looks at. The reverse also has to be true. A task type that starts on the autonomous side loses that status the moment corrections start showing up, not after a fixed grace period. AI governance for PMOs covers writing this as an actual policy rather than an ad hoc call made once per task type, and multi-agent orchestration covers what changes once more than one agent is closing work on the same plan.

The task types worth writing down first are the ones your team disagrees about most often. If two PMs would give different answers to "should the agent have closed that," the task type fails the objectively-checkable check by definition, and belongs in review until someone rewrites the definition of done to remove the ambiguity.

should ai agents mark work completeagent closing taskstrusting agents with statusAI AgentsPMOOnplanaAutonomy

Ready to make the switch?

Start your free Onplana account and import your existing projects in minutes.