Pick the Work, Not the Winner
13 min read · finance, work ladder, harness, skills, agents, judgment
Pick the work, not the winner. The useful AI move is handing agents a larger unit of work, not choosing which model won the week.
After the first week of September, the finance meeting opened the same way a lot of rooms did. Someone asked whether the team should switch to Astra. Someone else asked whether Fable 5.1 was approved yet. Both questions sounded operational, and both treated the model as the decision.
The room's working theory was that the smart move is to pick the winning model, then keep feeding it one task at a time: summarize this variance, rewrite the commentary on marketing spend, pull last month's actuals into the pack. That belief treats model news as the strategy, and it treats a chat prompt as the unit of work.
Pick the work, not the winner.
The useful move is handing agents a larger unit of work. Call it the Work Ladder: task, build item, project, initiative. Most people stay at tasks, some batch build items, and the leverage is the project, then the initiative. A pasteable harness skill organizes the pile and forces the climb. Once that system repeats, document manipulation compresses, and the remaining job is judgment.
What Actually Changed This Week
Astra is OpenAI's early-September model, marketed as general computer use, and the first OpenAI model designated Critical for cybersecurity under the Preparedness Framework. Enterprise workspaces keep it off until an admin enables it. The public model refuses advanced exploit work; Daybreak is the gated defender path.
Fable 5.1, from Anthropic on September 1, keeps the same $10 / $50 sticker as Fable 5. Cache reads dropped to $0.25 per million tokens, 75 percent cheaper, which is the part that actually changes long loops. The model is stronger at long-running agentic work, documents, and spreadsheets. Mythos 5.1 is the same model with lighter cyber and bio safeguards, trusted-access only.
Access is still a lease. OpenAI is winding down Cursor model supply after the SpaceX acquisition, with a proposed shutoff on November 12. None of that changes a finance org that still spends the week summarizing one variance in chat.
That is the news, and it is not the decision. What follows is about the unit of work you hand the agent.
The Work Ladder
The names matter because they change what you are allowed to hand over.
Task. One action, one output, minutes. "Summarize this variance." "Rewrite the commentary on marketing spend." "Pull last month's actuals into the pack." The output dies in the chat that produced it.
Build item. Several tasks, one artifact. Assemble the budget-vs-actuals pack: the period covered, actual versus plan, a named driver for every variance over the threshold, an open-questions line, and an owner who signs off. This is where Start With One Recurring Report lives. The report is not a project yet. It is one finished artifact inside a larger close.
Project. Several build items, one outcome. Own the monthly variance close: the live pack, plus an exception list sent to owners, plus the Tuesday review memo. The close is no longer a stack of favors. It is one outcome with a reviewable end state.
Initiative. Several projects, one strategic bet. Make the monthly close reviewable without tribal knowledge: the variance project, the pipeline-to-forecast project, and the headcount/spend project. The bet is not "use AI on the close." The bet is that a VP can read the month without calling the one person who still holds the story in their head.
Stop Counting Tasks. Start Closing Loops. made the company-level case: tasks are linear, loops compound. This ladder is the operator version of that idea inside one function. Compounding starts when the unit you hand the agent is large enough to reuse a review standard instead of inventing one for every prompt.
That is why the project is the useful leap. A task throws the standard away when the chat ends. A build item keeps it for one artifact. A project keeps it across several artifacts that have to agree with each other. An initiative keeps it across the month. The altitude of the work is the altitude of the standard.
Why People Stay at Tasks
Starting small is not the mistake
Most people should start at tasks. That is not the failure. A task is the only unit you can review when you do not yet have a written standard, and jumping straight to an initiative without one produces fluent garbage at a larger scope. Make Your Work Easier to Review is what makes the climb safe: if you cannot judge the output in minutes, you cannot raise the altitude.
The failure is treating the task as the permanent unit after the review standard already holds. The analyst who can already tell a good variance summary from a restated number, and still opens a new chat every time a line item needs a sentence, is not being careful. They are paying frontier prices for clerical work that should have been absorbed into the next tier.
Batching the pack is not the project
Some teams get past tasks and still miss the project. They batch the pack, which is a real improvement, and then rebuild the rest of the close around it by hand: the exception list from memory, the memo from a hallway conversation, the forecast still sitting in a different spreadsheet. The build item looks like progress because an artifact exists. The project has not moved.
Use agents at the project level. That is a turn, not a jump. You earn the project by holding the standard at the build item. You do not skip there because the model got cheaper at long loops.
The Harness Skill
Scale Your Top 1% defined a skill as encoded judgment. The harness is that idea made specific: it classifies the pile and forces the climb. Stop Asking AI Questions. Give It Work to Finish. argued that the useful unit is the finished artifact. Project-level work is a bigger one. The harness is how you find the right size before you start typing.
Any one input is enough: action items from the tracker the team already uses, a pasted list, or the last week of chats, meeting notes, and "what I said I would do." The skill organizes that pile into the four tiers, shows the map, then refuses to climb until the lowest unfinished tier holds. Paste the skill below, correct the names for your stack, and do not invent work the pile does not contain.
# Harness Skill: Organize, Then Climb
You are a work-altitude harness. Your job is to organize a pile of work into four tiers, show the map, and climb only after the review standard at the current tier holds.
## Tiers
- Task: one action, one output, minutes
- Build item: several tasks, one artifact
- Project: several build items, one outcome
- Initiative: several projects, one strategic bet
## Inputs
Accept any one of these. Do not wait for all three.
1. Action items from the team's existing project tracker
2. A pasted list of work
3. The last week of chats, meeting notes, or "what I said I would do"
## Rules
1. Organize the pile into the four tiers. Do not invent work. If a row is ambiguous, ask before classifying it.
2. Show the four-tier map before you act. Wait for acceptance or correction.
3. Execute the lowest unfinished tier until the written review standard holds.
4. Only then climb: tasks, then build items, then projects, then initiatives.
5. After each climb, write down what "good" looked like so the next run starts higher.
## Review standard
Before any execution, require one sentence plus failure modes for the current tier. If none exists, draft a standard and stop for approval. Do not execute against an implied standard.
## Climb rule
Do not skip a tier. If the review at the current tier failed, stay and fix. If asked to jump to an initiative while a lower tier is unfinished, refuse and name the unfinished tier.
The skill only counts if you can paste it. The model does not become wiser; it is no longer allowed to stay at the altitude that feels convenient. Without a classifier, every session defaults to the smallest prompt that feels finishable. The harness changes that default: the next unit is the lowest unfinished tier, which is the only way altitude rises on purpose instead of by accident.
A Finance Org, Worked
Here is how I go about it when the pile is a close week.
I dump two weeks of Slack, the close checklist, and a tracker export into one session with the harness skill. I do not ask it which model to use. I ask it to map the pile. The map that comes back, once I correct the rows it guessed, looks like this:
- Initiative: make the monthly close reviewable without tribal knowledge
- Projects: own the variance close; stand up pipeline-to-forecast; stand up headcount/spend
- Build items: assemble the variance pack; draft the exception list; write the Tuesday memo
- Tasks: summarize one variance, rewrite the commentary on marketing spend, pull last month's actuals, chase one owner for a driver
I climb one path. I do not rotate through all three projects in the same week. The variance close is the path, because the pack already exists as a ritual and the review standard is the easiest one to write down.
Week one: tasks only
The written standard is three checks: the period is right, the driver is a cause and not a restatement of the number, and an owner is named. Failure modes are the usual ones: last month's actuals attached to this month's pack, "marketing was over because marketing spent more," and a variance with no name next to it. I let the agent finish those tasks and review harder than a normal pass. This week is about whether the standard holds when someone else applies it.
Week two: the build item
I hand it the variance pack: period, actual versus plan, named drivers over the threshold, open questions, owner sign-off. That is the same artifact you would wire as a recurring report: one finished pack, not yet the close. I do not re-teach the connection. I raise the unit. The agent assembles the pack against the same three checks, plus one more: a VP could sign off without a follow-up question. If a driver still restates the number, we stay at the build item. We do not move on and hope the memo will hide it.
Week three: the project
I hand it the variance close: live pack, exception list to owners, Tuesday review memo. The three artifacts have to agree. A variance that cleared the threshold in the pack has to appear on the exception list. An open question in the pack cannot vanish from the memo. This is the first week the agent is doing project-level work, and it is also the first week the close starts to look like a system instead of a pile of chats.
The initiative stays human
I decide which of the three projects actually exist, what "reviewable without tribal knowledge" means in this shop, and what we should stop doing. Pipeline-to-forecast and headcount/spend wait until the variance close holds. They are siblings under the same bet, not a new example and not a skip.
If I handed the initiative to the agent in week one, I would get a fluent operating plan for a close nobody can review. That is the failure mode this ladder exists to prevent.
Judgment Is the Job That Remains
When the ladder is working, the clerical layer should compress toward the 80/20. Most of the pulling, restating, assembling, and chasing should not need a person sitting in the chat. The time that remains is judgment: what to forecast, what to escalate, what to stop. That is a design target for a working system, a Pareto on the manipulation layer, not a result I can pin on a named team. Do not wait for a published case study before you climb one tier.
The frontier labs are pointing at the same split in public. Anthropic Institute's "When AI builds itself" describes agents that can run defined experimental loops while humans still set the problem and the scoring rubric. Direction-setting is the remaining human role in that writeup. Lilian Weng's July 2026 note on harness engineering defines a harness as the system around the model that orchestrates tools, context, artifacts, and evaluation. Neither piece is a tour of how those companies run their organizations. Both describe the posture this ladder aims at inside a finance week: once the climb repeats, the agent runs the defined loop, and the human holds the problem and the bar.
Teams that stay at task level will keep paying frontier prices for clerical work, including the new computer-use and cheaper long-loop prices, and still start every close from a chat. Teams that climb spend the saved time on judgment.
This Week
Five actions, one week, all on the same close pile.
- Dump the pile: tracker, close checklist, or last week's chats, into one session.
- Run the harness skill. Accept or correct the four-tier map before any execution.
- Write the review standard for the lowest unfinished tier: one sentence plus failure modes.
- Let the agent finish that tier. Review harder than usual.
- Climb one tier only if the review held. Do not jump from tasks to an initiative.
Beginner's Mind Is Refusing the Model Meeting
Beginner's mind, in this context, means refusing to let the week's model news choose the week's work. Astra can operate a computer. Fable 5.1 made long loops cheaper. Neither fact tells you what unit of work to hand over. The naive move looks like sophistication: pick the winner, then keep the same chat habit with a better badge. The harder move is smaller and more honest: map the pile, hold the standard, and climb one tier.
Staying naive about vendors is how you stay serious about altitude. Pick the work, not the winner. Use agents at the project level once the review standard holds. The close does not get reviewable because a model won a week. It gets reviewable because someone stopped feeding it one task at a time.
Reflection Point
Which pile of work are you still feeding to AI as one task at a time, that is actually a project?