The Question Economy
A company gives two product candidates the same problem: customer retention has fallen for three months.
Both return polished plans. The first recommends better onboarding, lifecycle emails, and a referral program. The second starts by asking whether the decline is concentrated in a customer segment, acquisition channel, product version, or stage of onboarding. Then the candidate names the evidence that would separate those explanations and proposes the cheapest test for the leading cause.
The slides may look equally competent. Only one plan makes the thinking visible.
Generative AI is making surface quality a weaker signal of understanding. A student can produce a fluent essay before class. An analyst can turn a thin brief into a market map before lunch. A manager can create a finished plan before proving that the problem has been framed correctly.
For most of modern education, the answer was the proof. That inference is starting to break. Schools and companies now need better ways to see who understands the work when competent output can appear before competence.
The question economy is not about asking AI more things. It is about deciding which problems deserve to be solved, what evidence should change the answer, and who owns the result.
The practical response is to preserve foundational knowledge while making question design, verification, and judgment visible in classrooms, hiring, meetings, and management. Those capabilities can be taught and evaluated when they are attached to real decisions.
In Stop Asking AI Questions. Give It Work to Finish., I argued that the useful unit of AI is becoming the finished artifact. That article focused on the workflow. This one focuses on the assessment signal: how schools and companies make judgment visible when polished output no longer reveals it.
Before a system can finish useful work, someone has to choose the objective, define the boundary, and decide what done means.
What Changed
The scale of this change will take years to measure. If curriculum, assessment, hiring, and management all adapt to the falling cost of answer production, the result may become education's largest shift since modern mass schooling expanded during the eighteenth and nineteenth centuries.
The comparison rests on a change in the underlying constraint. A polished answer no longer proves who holds the knowledge or understands how to use it.
Modern education did not emerge from a single factory model. Mass schooling preceded industrialization in several countries. Religion, state building, democratization, military competition, and labor markets shaped different systems at different times. Industrialization reinforced demand for literacy, numeracy, technical skill, and instruction for large populations, but it did not invent education or reduce learning to memorization.
That system created enormous value. A 2026 LSE and CEPR working paper, using records of school foundations, charitable bequests, and roughly 350,000 apprenticeship contracts from 1711 to 1805, estimates that expanded elementary schooling increased entry into trades that required literacy and mathematics, including engineering and machine building.
The lesson is not that schools taught the wrong things. The relative value of an activity changes when technology changes its cost.
Before cheap digital retrieval, internalized knowledge had greater immediate practical value. Books, colleagues, and institutions have always served as external memory, but access took time. Search engines sharply reduced the cost of locating indexed public documents. Large language models now reduce the time required to draft, compare, explain, and synthesize those documents into a usable first version.
The abundance is uneven. Proprietary data, tacit expertise, attention, and reliable evidence remain scarce. Generated text can be fluent and wrong. Verification can consume every minute that generation saved. Even so, a plausible answer is becoming cheaper across a growing range of knowledge work.
AI can also suggest questions and tests. It does not preserve a permanent human monopoly on problem framing. The narrower economic point is that generating another possibility is usually easier than deciding which possibility deserves resources, risk, and organizational commitment.
Knowledge Is the Base Layer
Curiosity does not work in a vacuum. A person cannot ask an incisive question about monetary policy without understanding inflation, interest rates, incentives, and institutions. A product leader cannot challenge a security design without recognizing common failure modes. A patient cannot evaluate competing medical claims without some grasp of physiology, risk, and probability.
Daniel Willingham's work makes the learning science clear. Critical thinking depends heavily on domain knowledge. General strategies, such as looking for contrary evidence, can help. They only become useful when a person knows enough to recognize relevant evidence, compare explanations, and detect what is missing.
Knowledge stored in long term memory also reduces demands on working memory. This is why AI retrieval is not a reason to empty the curriculum. We do not learn multiplication tables because calculators are difficult to find. We learn them because fluency creates room for more complex reasoning. Retrieval practice remains one of the best supported ways to make knowledge durable.
The curriculum question becomes more selective and more demanding. Designers must decide which ideas learners need to internalize to reason without assistance, which facts they can retrieve when needed, and when an assignment should test recall or application to an unfamiliar case.
Knowledge is how curiosity earns precision.
From Curiosity to a Decision
Curiosity, question design, and judgment are related, but they are not the same capability. Curiosity supplies the motivation to investigate. Question design turns uncertainty into a structured inquiry. Judgment decides what the evidence means and what happens next.
A better question is not more philosophical or more clever. It is closer to a consequential decision.
Consider the question, "How can we improve customer retention?" It invites a list of familiar tactics.
A useful version looks different:
Among customers who complete onboarding, which behaviors during the first 14 days predict voluntary churn, what rival explanations could produce that pattern, and what low cost intervention would help us distinguish correlation from cause?
The second question identifies a population, time horizon, measurable outcome, competing explanations, and a path to experimentation. It does not establish causality. It makes causal learning possible.
I use a five stage Question to Decision Loop:
- Decision: What decision will this inquiry change?
- Model: What mechanism do we think is producing the current result?
- Disconfirmation: What evidence would make us abandon that explanation?
- Test: What is the cheapest valid experiment that could separate the leading explanations?
- Owner: Who will act on the result, and when will the decision be reviewed?
The order matters. Teams often start with a broad prompt, collect a large answer, and then search for a decision to attach to it. The loop starts with the decision so the research has a boundary. It asks for disconfirming evidence before the team becomes invested in a preferred story. It ends with an owner because unanswered accountability turns good analysis into another document.
AI can contribute at every stage. It can propose mechanisms, generate rival hypotheses, search for contradictory evidence, and design candidate experiments. The human advantage does not come from being the only source of questions. It comes from connecting the question to context, consequence, and responsibility.
The Accountable Conductor
The orchestra conductor is a useful analogy if we keep its limits in view. AI can play more instruments, produce parts of the score, and sometimes automate pieces of direction. No metaphor guarantees a permanent division of labor.
In consequential work, however, someone still has to own the performance. The accountable conductor defines the objective, chooses the evidence threshold, identifies unacceptable risks, and decides whether the output is safe enough to use. Prompting is a small part of that responsibility.
The workplace evidence supports a bounded version of this shift. In a staggered rollout at one Fortune 500 software company, researchers studying 5,179 support agents estimated that a generative assistant increased issues resolved per hour by about 14 percent. Gains were concentrated among less experienced workers. The researchers found suggestive evidence that the system spread practices used by stronger performers.
In a preregistered experiment involving 758 Boston Consulting Group consultants, GPT 4 users completed more work, worked faster, and earned higher ratings on 18 tasks classified as within the model's capabilities. On one task deliberately selected outside that frontier, AI users were 19 percentage points less likely to reach the correct conclusion.
These studies do not prove that curiosity is already the economy's decisive skill. They show that AI can improve output while making misplaced trust expensive. As production gets cheaper, task selection, verification, and accountable judgment become more important parts of the job.
What This Changes Inside a Company
Imagine a support organization facing an 11 percent increase in weekly contacts. The obvious response is to use AI to draft replies faster. A better operating question asks which product defects account for the increase, which customer behaviors precede the contacts, and which intervention could prevent them without moving the cost to another team.
AI can cluster the tickets and surface candidate causes. The team still needs to verify the source data, select a pattern worth investigating, design the intervention, and define what result would justify a product change.
The greatest value may not come from answering tickets faster. It may come from removing the reason for the tickets.
Companies can build this discipline into four operating practices:
- Decision briefs: Every material initiative states the decision, current model, rival explanations, disconfirming evidence, and cheapest valid test.
- AI risk tiers: Low impact work receives spot checks. Medium impact work requires traceable sources and a reviewer. High impact work requires an accountable domain expert and independent validation.
- Decision reviews: Teams revisit forecasts, assumptions, and evidence. They examine whether contrary findings changed the decision and reward belief revision, not only confident advocacy.
- Work sample hiring: Give candidates an ambiguous problem from the actual job. Evaluate how they define it, request evidence, design a test, and update their view. Do not confuse the number of questions asked with the quality of inquiry.
This also changes management. A manager does not create value by carrying information that the team cannot access elsewhere. The manager creates value by improving how the team frames problems, makes tradeoffs, and learns from results.
Meetings should expose assumptions and sharpen decisions, not reward the performance of already knowing the answer.
The organization with the largest knowledge base will not automatically win. A knowledge base creates value when people can retrieve, challenge, and act on it well.
What Education Should Optimize For
If question design and verification become more important at work, education must make them visible and assessable. The sequence matters because novices cannot reliably evaluate an AI answer in a field where they lack the schemas to detect plausible mistakes.
Early instruction should still emphasize clear teaching, worked examples, guided practice, and retrieval. Developing learners can predict outcomes, explain their reasoning, and compare examples. Advanced learners can use AI to generate rival hypotheses, interrogate sources, and design tests.
Curiosity becomes productive after knowledge gives it somewhere to go.
Assessment should follow the same progression. Instead of grading only a final essay, an instructor might require the original question, an unaided hypothesis, a source ledger, selected AI interactions, an error log, the revised argument, and a short oral defense. A delayed transfer task without AI can test whether the student learned something that remains available after the tool is removed.
The final answer still matters. It is simply no longer the only evidence of learning.
There is a real risk of cognitive offloading. The OECD's 2026 Digital Education Outlook concludes from emerging evidence that AI used as a shortcut can displace the effort required for deep learning.
A 2025 survey of 319 knowledge workers found a correlational, self reported pattern: greater confidence in generative AI predicted less reported critical thinking, while greater task specific self confidence predicted more. It did not show that AI caused a lasting loss of reasoning ability.
The design principle is practical: preserve intellectual resistance when the resistance serves learning. Ask students to predict, attempt, explain, or compare before the system reveals an answer.
The goal is not maximum convenience. It is durable capability.
The Unresolved Boundary
The source of an original thought remains a deeper boundary, not an answer to the business argument.
We understand a great deal about neurons, synapses, electrical signaling, memory, and the neural correlates of conscious states. We still lack a consensus theory connecting brain activity to subjective experience or fully explaining why one thought emerges instead of another.
Consciousness, free will, emergence, and determinism remain open scientific and philosophical problems. Readers interested in that boundary can explore philosophy of mind further.
None of this proves that humans possess a mystical faculty machines can never acquire. The practical distinction today is responsibility: people authorize consequential objectives and remain accountable for what humans and AI do in pursuit of them.
A Weekly Practice
You do not respond to AI by learning less. You learn with greater purpose.
Once a week, choose one real decision and run it through the Question to Decision Loop:
- Record your current belief and confidence.
- Ask AI for three rival explanations.
- Identify the evidence that would distinguish them.
- Verify consequential claims using primary sources.
- Run the cheapest valid test available.
- Record whether the evidence changed your mind.
This practice develops more than prompting. It builds domain knowledge, calibration, decomposition, source judgment, and belief revision. Over time, it teaches disciplined curiosity: the habit of pursuing questions whose answers could change a decision.
As information becomes more abundant and execution more automated, judgment becomes scarcer and direction more strategic. Knowledge remains necessary. Curiosity multiplies its value when it becomes disciplined inquiry tied to a decision.
Before the next important decision, do not ask only who produced the best answer. Ask each owner to name the decision, explain the model, show what would change their mind, and define the test.
That is how you make thinking visible before you automate the work.
Reflection Point
If good answers are becoming abundant, which better question can you now afford to pursue?