About the Ladder

For colleagues weighing whether this is worth their students’ time. The short version: it teaches one skill, judgment, and it is assessed on the student’s reasoning rather than on anything the AI produced.

The problem it is built against

Most AI instruction fails in one of two directions. It teaches rules — what is permitted, what counts as misconduct — which produces students who know the policy and have no idea what the tool actually does. Or it teaches features — prompt techniques, plug-ins, the tool of the month — which produces students who are fluent and credulous, and whose training expires with the next release.

Both leave out the thing employers and faculty actually want, which is a student who can look at a confident, fluent, well-formatted answer and decide how much of it to believe.

The design claim

Judgment about AI cannot be taught before fluency with AI. A student who has never used these systems has no basis for knowing where they fail. Their caution is borrowed — repeated from an instructor, not earned from experience — and it collapses the first time the output looks good.

Why the rungs are in this order

The Ladder is not fifteen skills. It is one skill exercised on progressively harder targets. Each rung raises what is at stake in the judgment being asked for.

What each rung is actually for.
RungWhat it is forWhy it sits here
1 · PlayRemoving the fear that stops a student looking closely.Nothing is at stake, so nothing is being protected. This is the rung faculty most often want to cut, and the one the rest depends on.
2 · DirectDiscovering that the answer depends on them.Students arrive believing AI has an answer. Editing one prompt and watching the output move is the fastest correction to that belief.
3 · ApplyMeeting the limits on work they know well.Thinning is invisible on an unfamiliar topic. On their own material they can see exactly where the answers stop being real.
4 · InterrogateSeparating the machine’s recommendation from their own call.Only possible once they have found a limit themselves. Exercise 12 — marking where the AI stopped and they started — is the hinge of the whole sequence.
5 · VerifyIndependent checking, including of the check.This is ordinary information literacy, arrived at by students who now have a concrete reason to want it.

Reverse the order and it does not work. Verification taught first is a lecture about sources; taught fifth, it answers a question the student has already run into.

How it is assessed

Every exercise ends with a judgment, not an artifact. The log records two things: what the student kept, and what they judged. The second is the assessed one.

This is the answer to “how do you grade work a machine helped produce?” You do not grade the output. You grade the call the student made about it — which the machine cannot supply, because it is a judgment about the machine.

Exercise 14 is the clearest case. Students trace a claim to its original source, and “I could not find a source” is a passing answer. An honest null result is the correct outcome when the source does not exist, and saying so is the skill.

Ambassador applicants produce three different kinds of evidence, which is more than a completion certificate can carry:

Developmental

The log — fifteen dated judgments showing the change as it happened, not recalled afterwards.

Metacognitive

The calibration memo — one argument about how their judgment changed, including what they still cannot do well.

Performance

The walk-through — sitting with another student, hands off the keyboard, and reporting what that person was afraid of.

Objections worth taking seriously

“This teaches students to hand work to a machine.” Every exercise requires a judgment the machine cannot produce, and the last rung is spent checking it. A student optimising for effort would find this a poor bargain: the Ladder is more work than doing the assignment honestly, and none of it produces submittable text.

“Play is not rigour.” Agreed, and Rung 1 is not where the rigour is. It is where the fluency is, and rigour applied to a tool the student cannot operate is theatre. The stakes rise every rung after it.

“Four to five hours cannot matter.” The claim is narrow on purpose. This is a calibration, not a course. It does not make anyone expert; it makes them harder to fool.

“It will be obsolete in a year.” Nothing in it names a vendor, a model, or a feature. Exercise 10 asks students to compare two AIs precisely so that no single system’s behaviour is mistaken for how AI works. The exercises survive a model release; a prompt library does not.

“Students will just have AI write the log.” Some will, and the memo is where that shows. A student who never climbed cannot say which exercise changed their mind or name the boundary of their own competence — and a memo that claims mastery reads, to the person marking it, as the weaker submission.

Why this matters in a business college

The case for teaching AI judgment is usually made about employers. For a business school the sharper case is about firms — because the published data says adoption is already high, capability is not, and the gap is training.

What the surveys actually report.
FindingSource
76% of small businesses use AI. 84% name efficiency and productivity as the main benefit — and 73% say they would benefit from more training and implementation support. Adoption is not the bottleneck. Knowing how to use it well is. Goldman Sachs 10,000 Small Businesses Voices, survey of 1,256 owners by Babson College and David Binder Research, January–February 2026.
60% of small businesses use AI platforms, up from 23% in 2023. Generative AI use rose from 40% to 58% in a single year. The tools arrived faster than anyone was trained on them. U.S. Chamber of Commerce, Empowering Small Business: The Impact of Technology on U.S. Small Business, August 2025.
The smallest firms are falling behind. 37% of firms with 250+ employees use AI, against under 20% of firms with fewer than five. Between December 2025 and May 2026 use rose among firms with at least 20 employees and did not move for the smallest ones. U.S. Census Bureau, Business Trends and Outlook Survey, December 2025–May 2026.

That is the entrepreneurial case. A graduate who can use AI with judgment is useful to an employer. A graduate who can also teach it is useful to a small firm that has the tools and no one who knows how to use them — which, by the numbers above, is most of them. It is also the position a founder is in on day one.

Where this comes from

The Ladder is not itself a research finding, and this page does not claim it is. It is a design, and the design rests on results that are well established elsewhere. The relevant ones, and what each is doing here:

The evidence behind each design decision. None of it is about AI specifically; all of it is about people.
The decisionWhat supports it
Verification last, not first Expert fact-checkers evaluate sources by leaving the page and reading laterally — a strategy neither professional historians nor Stanford undergraduates used, despite both groups being able to read critically. The skill is procedural and has to be practised, not described.
Wineburg, S., & McGrew, S. (2019). Lateral reading and the nature of expertise. Teachers College Record, 121(11).
Students must find the limits themselves Learners who struggle with a problem before being taught the method outperform those taught the method first, even though their initial attempts fail. Being told where a tool breaks is not the same as meeting the break.
Kapur, M. (2008). Productive failure. Cognition and Instruction, 26(3), 379–424.
Difficulty rises rung by rung Conditions that slow performance during practice — spacing, variation, effortful retrieval — reliably improve what is retained afterwards. Easy practice produces confident performance and poor durability.
Bjork, E. L., & Bjork, R. A. (2023). Introducing desirable difficulties into practice and instruction. In In Their Own Words (pp. 19–30). APA Division 2.
The memo asks what they still cannot do Students who judge their own understanding inaccurately study the wrong things and learn less; accuracy of self-assessment predicts achievement independently of ability. Naming the edge of your competence is the trainable part.
Dunlosky, J., & Rawson, K. A. (2012). Overconfidence produces underachievement. Learning and Instruction, 22(4), 271–280.
Rung 4 separates the machine’s answer from the student’s call People systematically over-trust automated aids that are usually right, and stop checking them — a failure mode documented long before generative AI, and one that operator training, not better automation, is what mitigates.
Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230–253.
The problem is current, not hypothetical The same over-reliance pattern is now documented specifically for AI assistance: people accept confident wrong answers at rates that rise with the system’s general accuracy.
Passi, S., & Vorvoreanu, M. (2022). Overreliance on AI: Literature review. Microsoft Research, Aether.

What is not claimed. None of this is evidence that this Ladder works. It has not been evaluated, it has no control group, and the first cohort is the first data. What the citations establish is that the design is not arbitrary — each choice is the application of a result that holds elsewhere. Whether it holds here is an open question, and one worth answering properly.

What it does not do

Stated plainly, because the limits are part of the design.

It does not teach prompt engineering as a craft. It does not cover discipline-specific tools, and it is not a substitute for anything taught in a major. It does not assess writing. It is not a policy and it does not tell anyone what is permitted in any course — that stays with the instructor. And it does not certify competence: finishing means fifteen exercises were done and fifteen judgments were recorded, nothing more.

What it costs

Nothing to run. No licence, no platform, no account, no data collected — the site stores nothing, and the log lives in the student’s own document. It works with whatever AI a student already has, including the free tiers. Any instructor can point students at it without adopting anything, and any student can climb it without asking permission.

If you want to try one rung before deciding, Rung 1 takes about twenty minutes.