
Artificial intelligence didn’t break education. It exposed a question we’ve been avoiding for two hundred years: what are we actually measuring when we say someone has learned?
I remember the books just fine. Hamlet. Othello. The problem was never memory.
Here’s the confession, offered now that the statute of limitations has safely expired. Throughout high school, in honors and Advanced Placement English, whenever a book arrived with an essay attached, I would almost always reach for the little yellow-and-black booklet. If you went to school before the internet, you know the one. CliffsNotes occupied a strange shelf in the bookstore, right out in the open, sold legally to children, containing everything a teacher hoped you would discover on your own: the themes, the symbols, the character arcs, the interpretations that would earn approving comments in the margins. The booklet didn’t feel like contraband. It felt like a study aid, which is exactly what the cover said it was.
Now, before you picture a slacker, understand the scene. This was AP English. We were nerds, and mostly proud of it, and most of the time we genuinely made the effort to comprehend what we were reading, booklet or no booklet. The notes were often a supplement to understanding rather than a substitute for it. But not always. That’s the part that matters. When a book grabbed me personally, it got the full treatment: the real reading, the actual wrestling, the opinions I’d defend at lunch. When it didn’t grab me, or when the deadline was closing in, and I confess Shakespeare landed in this category more often than I’d like to admit, I’d let the booklet do the comprehending for me. I’d skim the summary, note which ideas kept repeating, pull a few lines from the play so I could quote something with page numbers, and assemble an essay that sounded like a person who had wrestled with literature.
And here’s the detail that took me thirty years to find disturbing: my grades could not tell the difference. The essays born of genuine engagement and the essays assembled from the booklet came back with the same marks and the same encouraging comments in the same red ink. Same student, same output, two completely different realities inside my head, and the instrument measuring me registered no difference whatsoever.
Well. Almost no difference. Let me give you my favorite exhibit.

Freshman year, first semester, the book was Romeo and Juliet, and I’ll be honest with you now in a way I wasn’t with anyone then: I didn’t read it. Not skimmed it, not started it and drifted off. What I read instead, closely, almost scholastically, was the CliffsNotes. I studied that booklet like it was the assigned text, absorbed its arguments, and then went back through the play itself hunting for quotes that would support the positions I’d borrowed. It was, in its way, a genuine research process, rigorous, methodical, and aimed entirely at the wrong document. The teacher held my essay up as the best in the class. And in the economy of that classroom, best in class carried a jackpot: an automatic A for the entire first half of the year. At fourteen, I had produced award-winning literary analysis of a play I had never read, which, in retrospect, should have been somebody’s clue.
The clue arrived at the end of the year, wearing an F. Our final exam circled back to Romeo and Juliet, and this time the questions weren’t essay-shaped. They asked what actually happened in the play, and I sat in that room discovering, in real time, that I remembered essentially nothing, because there had never been anything to remember. The booklet’s arguments had passed through me the way a script passes through an actor after the show closes. So the same institution measured me twice on the same play and returned two opposite verdicts: the essay declared me the finest literary mind in the room, and the final revealed I had learned nothing at all. The second verdict was the accurate one. And the lesson I should have drawn, though it took me three more decades, is that understanding leaves a residue and performance evaporates. The final wasn’t a better instrument in any sophisticated sense. It just showed up months later, which is long enough for a counterfeit to completely dissolve.
For most of my life, I filed all of this under mild adolescent guilt, somewhere between copying a friend’s homework and telling my parents the arcade was a library. Everybody cut corners. It wasn’t interesting.
Then, a few years ago, the world started arguing about ChatGPT in classrooms, and the memory came back with a different meaning attached.
I should say up front that I was, by the system’s official accounting, a good student, and also never a natural one, and I’ve never pretended otherwise. I struggled with basic multiplication. My grammar remains a work in progress, as my editors will confirm under oath. What I had instead was a kind of relentless practical curiosity, mostly aimed at video games. By ten I was doing summer work at hole-in-the-wall computer shops. By twelve I was play-testing games for SNK and Interplay. By fifteen I was behind the counter at Babbage’s, and every job since has been, if I’m honest, an elaborate scheme to stay close to gaming. School was something I could do well when I chose to, and the choosing depended, like everything else in my life, on whether the thing in front of me had earned my curiosity.
Which means I was precisely the kind of student the current AI panic is not supposed to be about. The panic has a character in mind: the lazy kid, the corner-cutter, the one who would learn nothing if you handed him a machine that writes essays. But I was the honor student. The AP kid. And the honor students were gaming the system too, not constantly, not carelessly, but selectively and rationally, deploying real comprehension where it felt worth the cost and manufacturing the appearance of it where it didn’t. The system as it existed in the early 1990s was already perfectly gameable by a kid with three dollars and a bookstore, and the kids gaming it were often the ones wearing the academic letter jackets. That gaming wasn’t a malfunction we discovered. It was, I’ve come to believe, the system working exactly as designed, revealing exactly what it was designed to reward.
The interesting question was never why students look for shortcuts.
The interesting question is why the shortcuts work.
The Thing Nobody Has Ever Seen
Here is a claim that sounds absurd until you sit with it: no teacher has ever observed a student’s understanding directly. Not the best teacher you ever had, not Socrates, not anyone.
They can’t. Understanding has no physical form. It can’t be weighed, photographed, or extracted for inspection. It lives entirely inside another person’s head, and the only access anyone has ever had to it is through the traces it leaves behind: the things a person says, writes, builds, and does. An essay is a trace. A test score is a trace. A diploma is a trace of traces. Each one is evidence from which we infer something invisible, the way a detective infers a burglar from footprints without ever meeting him.
A thermometer is not heat. A map is not the territory. And an essay is not understanding. It’s a measurement instrument pointed at understanding, and like all instruments, it can be fooled.
I didn’t fully grasp this distinction until I started sitting on the other side of the table, hiring people.
I’ve spent more than twenty-five years in branding and marketing, most of it in the gaming industry, and a large part of that time has involved hiring creative people: designers, artists, writers, video editors, marketers. If you’ve done a lot of hiring, you eventually make peace with a humbling fact. The résumé tells you almost nothing. The degree tells you almost nothing. Even the portfolio, which ought to be the gold standard for creative work, tells you less than you’d hope, because portfolios can be curated, coached, borrowed from group projects, and polished by friends. I have interviewed candidates with immaculate portfolios who could not explain a single decision in their own work, and I have hired people with thin, scrappy portfolios who turned out to be the best creative minds I’ve ever worked with.

What actually tells you something is conversation. Ask a designer why she chose that typeface and watch what happens. The candidate who understands her own work lights up and starts arguing with herself, telling you what she’d change, where the client pushed back, what she’d do differently with another week. The candidate who performed the work, or had it performed for him, gives you back adjectives. “It felt clean.” “It matched the brand.” The portfolio was the essay. The conversation is where the understanding either shows up or doesn’t.
Employers rely on résumés for the same reason teachers rely on essays: nobody has the time to do the thing that would actually work. If a company had unlimited time, it would never read a résumé again. It would spend six months working alongside every applicant, watching how they take criticism, what they do when a project collapses on a Thursday afternoon, whether they admit mistakes or relocate blame. Those qualities determine whether someone succeeds in a job, and none of them appears on paper. The résumé exists because those qualities only reveal themselves over time, and time is the one thing no hiring manager has.
So the degree becomes shorthand for knowledge. The previous job title becomes shorthand for competence. The credit score becomes shorthand for whether you’ll repay the loan, which is the only thing the bank ever cared about. The quarterly earnings report becomes shorthand for the long-term health of a business. Civilization runs on proxies because the realities we care about are invisible, slow, or both.
There’s nothing wrong with this. Proxies are one of humanity’s great inventions. The problem starts later, and it starts quietly, when we forget that the proxy is a proxy.
The Performance of Understanding
No institution wakes up one morning and decides that the measurement matters more than the thing being measured. The drift is gradual, and it happens because measurements have qualities that realities don’t. Measurements can be counted, compared, standardized, archived, reported to administrators, and graphed over time. Understanding resists every one of those conveniences. Given a choice between something manageable and something elusive, institutions will drift toward the manageable every single time, without a villain anywhere in the story.
Picture a teacher who assigns an essay because thoughtful writing genuinely reflects thoughtful reading, which it does. At first the assignment works beautifully. Students who engage deeply write better essays than students who don’t. The essay is functioning exactly as intended: not as learning itself, but as evidence of learning.
Then the essay becomes important. Grades depend on it. College admissions depend on grades. Scholarships depend on admissions. Parents ask how to earn higher scores. Test-prep companies publish model essays. Teachers, trying to be fair, hand out rubrics specifying exactly what earns an A, and the rubric becomes a blueprint. Students trade reliable thesis structures the way my generation traded game cheat codes. Every step in this chain is reasonable. Every actor is responding sensibly to incentives they didn’t create. And the cumulative result is that students get progressively better at producing the appearance of understanding, which sometimes reflects the real thing and sometimes doesn’t, and from the outside the two become harder and harder to tell apart.
Kids figure this out astonishingly early, and adults consistently misread what they’re seeing. Long before children can do algebra, they’ve learned to ask whether there’s partial credit. Long before they can appreciate a novel, they’ve learned to ask whether it will be on the test. We hear those questions and diagnose laziness. I think we’re witnessing something closer to brilliance. Children are extraordinary students of institutions. Nobody teaches them this; they simply observe, with the unclouded eyes of anthropologists, that every school contains two curricula. The official curriculum says understand Shakespeare. The hidden curriculum explains how Shakespeare will be graded. The official curriculum celebrates curiosity. The hidden curriculum pays out on performance. When the two drift apart, students don’t cause the drift. They’re just the first to notice it, and the first to act on what they see.
Economists gave this pattern a name decades ago, though the name usually gets misattributed. In 1975, the British economist Charles Goodhart, writing about monetary policy, observed that any statistical regularity tends to collapse once pressure is placed on it for control purposes. The catchier version everyone quotes, that when a measure becomes a target it ceases to be a good measure, was actually coined later by the anthropologist Marilyn Strathern, summarizing him. Around the same time as Goodhart, the social scientist Donald Campbell arrived at nearly the same conclusion from the direction of education, warning that the more any quantitative indicator is used for decision-making, the more it will distort and corrupt the very process it was meant to monitor. Campbell discussed education among the applications of this problem in the 1970s, well before today’s chatbots.
Neither man was describing dishonesty. They were describing physics. People optimize for whatever the system rewards; that is close to a law of nature. Companies optimize quarterly earnings. Hospitals optimize wait-time metrics. Police departments optimize crime statistics. Social platforms optimize engagement. Students optimize grades. The remarkable thing is not that people do this. The remarkable thing is how reliably institutions mistake the optimization for corruption, as if human beings should be expected to ignore the scoreboard hanging directly over their heads.
Teenage me, with the yellow-and-black booklet, was not corrupt. Teenage me had simply read the scoreboard correctly. The essay assignment asked, on its surface, “do you understand this book?” But the grading, the rubric, and the entire apparatus around it asked a different question: “can you produce the artifact we associate with understanding?” I answered the question that was actually being asked. Students always do.
Which brings us, finally, to the machines. But to understand why artificial intelligence detonated inside education the way it did, you first have to understand what schools were built to do. And for that we have to be fair to the schools, which almost nobody in this debate is willing to be.
In Defense of the Factory School
There’s a critique of education so common it has become furniture. You’ve heard it: schools were designed on a factory model, in rows, with bells, to produce obedient industrial workers, and that’s why they’re broken. The critique contains real truth, and it’s also deeply unfair, in the way that judging your grandparents by problems they couldn’t have imagined is unfair. Because the system everyone loves to condemn solved one of the greatest bottlenecks in human history, and it solved it so completely that we’ve forgotten the bottleneck ever existed.
For nearly all of civilization, knowledge was heartbreakingly fragile. A farmer might understand his local soil and weather better than any living person, and that understanding died with him. A master stonemason spent fifty years perfecting techniques an apprentice could inherit only through a decade of watching. Physicians accumulated experience one patient at a time and carried it to the grave. Books, even after Gutenberg, were precious objects rather than everyday possessions, and most people couldn’t read them anyway. Ideas traveled at the speed of ships and horses. A brilliant insight in one city could take a generation to reach another, and entire populations lived and died without encountering the most important discoveries of their own era.
Against that backdrop, the challenge facing education was almost the opposite of ours: how do you transfer as much knowledge as possible into as many heads as possible, using nothing but voices, chalk, and paper?
Before fast communication and searchable digital references, access to knowledge depended far more heavily on memory, books, local colleagues and apprenticeship. Memory mattered because retrieving information was slower and less reliable. It still matters today. The change is in how much information we can reach, and how quickly we can reach it, rather than a clean historical switch from memorization to intelligence.
Public school systems developed through different reforms in different places. Their history cannot be reduced to one factory blueprint or one military defeat. What interests me is the operating problem they share: how to make education available to many people while preparing teachers, organizing a curriculum and assessing progress. Standardization helped make that scale possible, even as it introduced compromises.
It’s worth pausing on how radical that was. For thousands of years, knowledge had been the private property of clergy, aristocrats, and guilds. Mann and the reformers proposed to give it away, to everyone, at public expense. The standardized features we now find soulless were the price of that ambition. Standardized curricula meant a child in one town learned roughly what a child in another town learned. Exams verified the knowledge had actually landed. Essays forced students to synthesize rather than merely recognize. Diplomas certified a minimum threshold to strangers who would never meet your teachers. None of it was perfect, and perfection was never the goal. Scale was. Socrates could spend an afternoon interrogating a single student because he was teaching a handful of Athenians; a nation trying to educate every child confronts arithmetic Socrates never faced.
And the compromise worked. Literacy went from privilege to default across Europe and North America in roughly a century. The system helped produce the scientists, engineers, physicians, and yes, the marketers, who built the modern world. If today’s education system deserves criticism, it also deserves credit for most of the things critics use to publish their criticism.
I insist on this defense for a reason beyond fairness. It guards against the laziest explanation of institutional failure: that institutions persist because people are too stubborn to change them. That’s rarely true. Institutions persist because they keep solving the problem they were designed to solve. The trouble begins when the problem itself quietly changes, and the institution, still functioning perfectly, keeps solving the old one. Education optimized for scarcity. We now live in abundance.
Sometime in the last half century, the problem changed. There was no headline. But humanity crossed a threshold that I’d argue matters more than the invention of any particular technology, including the one currently terrifying every school board in America.
When Knowledge Outgrew the Knower
For most of history, becoming an expert meant accumulating enough information to stand among the most knowledgeable people in a field, and the ideal, however difficult, remained conceivable. A physician could plausibly aspire to know the important medicine of his generation. A lawyer could hold the relevant law in his head. Even physics, for a while, fit inside a person: a well-trained physicist in 1900 could be genuinely conversant with most of what physics knew.
Then the twentieth century happened to knowledge. Fields didn’t just grow; they split, and the splits split. Physics alone birthed quantum mechanics, then particle physics, cosmology, condensed matter, string theory, each a paradigm that would have consumed an entire nineteenth-century career just to enter. Discoveries began arriving faster than textbooks could be revised. The twentieth century didn’t merely produce more knowledge than previous centuries. At some point, it produced more knowledge than any individual could possess, and that’s a different kind of event.
The scale is difficult to absorb. PubMed describes a collection of tens of millions of biomedical citations and abstracts. That is a discovery system, not a reading list any individual can finish. Specialists need ways to find relevant evidence, assess its quality and connect it to the problem in front of them. Looking something up does not make an expert a fraud. Knowing where to look, what to trust and when to seek help is part of the expertise.
Here’s a version of this problem I’ve never seen in an education op-ed, but it’s the one that made it click for me. Think about how history is taught. American students study American history, reasonably enough, and it spans about 250 years. A student in China faces thousands of years of continuous recorded civilization. A student in Iran or Egypt, older still. No Chinese student, however brilliant, actually learns Chinese history in any complete sense. They learn a curated sliver, the dynastic highlight reel, and we’d never call them ignorant for it, because the alternative is impossible. Now notice that every field is becoming China. Medicine is China. Law is China. Computer science became China in about a decade and a half. The territory has grown so vast that every expert is now working from a highlight reel, and the honest ones know it. The people we call experts today are, in the old sense of the word, superficial experts, not because they got lazier but because civilization got bigger than any one mind. Experts didn’t become less knowledgeable. Knowledge became larger than expertise.
That sentence is, I think, the defining condition of the twenty-first century, and I watched it happen inside my own industry in real time. When I started in gaming, a dedicated person could genuinely keep up: the platforms, the studios, the scene. Somewhere in the 2000s that stopped being possible by any amount of personal diligence. The industry fragmented into more games, genres, platforms, communities, and micro-cultures than anyone could track, and my job, which depends entirely on staying current, quietly transformed from a memory problem into a navigation problem. My solution was unorthodox: I learned to treat my own kid as a discovery system. When my son started watching an unknown streamer named Lirik on an obscure site called Twitch in 2013, that signal was worth more than any industry report I could have memorized, and the VCs who dismissed Twitch as kid stuff learned its value the following year, when Amazon announced its approximately $970 million acquisition agreement. The lesson wasn’t that my son knew more than the VCs. It’s that in a world of infinite information, the scarce skill had become knowing where to look and what mattered, not what to retain.
For thousands of years, knowledge was something individuals accumulated. It has become something individuals navigate. Those are profoundly different activities, requiring different skills, deserving different measurements.
And yet our schools, built with love and genius for the world of accumulation, kept right on measuring retention and reproduction, because the instruments still worked, more or less, and nothing had come along to prove otherwise. What does it mean to be educated in an age when knowledge is infinite? For fifty years, that question sat in the room, unasked, because no one was forcing the issue.
Then something came along.
The Exposure
The speed of what happened next is worth stating plainly, because I don’t think most adults have absorbed it. In 2023, 13 percent of American teens said they used ChatGPT for schoolwork. A year later, it was 26 percent. By late 2025, more than half of all teens were using AI chatbots for schoolwork, and 59 percent said that cheating with AI had become a regular feature of life at their school. That’s not adoption. That’s a dam breaking.
And here’s the detail from the survey data that the panic coverage always skips, because it ruins the “lazy kids” narrative. The teens have a moral framework, and it’s more coherent than the adults’. In the Pew research, a majority of teens said using ChatGPT to research new topics was acceptable, while only 18 percent said the same about using it to write essays. Sit with that. The survey responses distinguished between using the machine to access knowledge and using it to fake the evidence of knowledge. They drew the line exactly where a philosopher would draw it, between the tool and the fraud, and they drew it while the adults in the room were still arguing about whether the tool should exist. Students don’t lack an ethics of AI. They lack an assessment system whose ethics make sense to them.
The institutional response included an arms race in detection software. OpenAI withdrew its own classifier in July 2023 because of low accuracy. On its published English challenge set, it identified 26 percent of AI-written text as likely AI-written and incorrectly labeled 9 percent of human text. Those are results for that detector and evaluation, not a universal accuracy rate. A 2023 study in Patterns also found bias against non-native English writing in the detectors it tested. Detection may provide a signal, but it cannot establish authorship or comprehension on its own. The more useful question remains what the student can explain and apply.
Because think about what the panic is actually conceding. If a machine-generated essay is indistinguishable from a B-plus student essay, then the B-plus student essay was never demonstrating what we told ourselves it demonstrated. The essay was measuring the ability to produce essay-shaped text: organized, on-topic, gesturing at themes. That ability used to correlate with understanding well enough to be a useful proxy, the way a résumé used to correlate with competence. ChatGPT didn’t sever the correlation. It revealed how much of what we were grading was the shape, not the substance. My CliffsNotes essays proved the same thing in miniature, decades earlier, at much lower bandwidth. AI just industrialized the demonstration.
This is why I’ve become convinced that the technology isn’t the story. Artificial intelligence didn’t create this problem. It merely exposed it, and the exposure raises a far more uncomfortable question than anything about cheating: if a machine can produce the work we’ve always accepted as evidence of learning, were we ever measuring learning correctly in the first place? We spent two centuries perfecting instruments for measuring the reproduction of knowledge, in a world where reproduction was the bottleneck, and then the bottleneck moved, and we kept using the instruments. AI began producing convincing responses to many familiar assignments, and in doing so asked the question the instruments were built to avoid: fine, but what did the student comprehend?
The Last Closed-Book Room
Before deciding what schools should do about any of this, it’s worth noticing something about the world students graduate into, because that world settled the question years ago and nobody sent the schools a memo.
Every serious profession has already gone open-book. Your surgeon does not perform from memory alone; she operates inside a lattice of checklists, imaging, protocols, and reference systems, and the profession considers this a triumph, not a scandal, because checklists can support reliable practice alongside professional judgment. I could ask a working surgeon a specific technical question about a procedure adjacent to her specialty, and she might tell me, comfortably, that she’d need to look it up before answering, and no one in the room would think less of her. What we’d judge her on is whether she looked in the right place, understood what she found, and acted on it correctly. Pilots, the most trusted professionals on earth, are famous precisely for refusing to rely on memory; the pre-flight checklist exists because aviation learned, in blood, that recall under pressure is exactly the wrong thing to bet lives on. Nobody boards a plane hoping the captain memorized the manual and is winging it from recall. Lawyers practice law with databases open. Engineers design with references at their elbow. Programmers, and I say this with affection for every programmer I’ve ever employed, have spent two decades openly joking that their profession is professional-grade searching, and their industry runs the world.
In other words, the entire adult economy quietly converted to the model where intelligence means finding, evaluating, and applying, and where the shame is not in looking things up but in failing to understand what you found. Yet many classroom examinations still ask students to work alone, from memory, without references. We built a closed-book examination system to prepare people for a closed-book world, the world of the eighteenth-century physician, and then the world went open-book and we kept the exam. Students can see this. They are, remember, extraordinary students of institutions, and they can observe with their own eyes that no working adult in their lives operates the way a final exam demands. When students sense that an assessment measures a skill the real world visibly doesn’t use, the assessment loses its moral authority long before it loses its official one, and technologies like ChatGPT don’t create that erosion. They just give it a tool.
None of this means memory doesn’t matter; that’s the strawman version of the argument, and it deserves to be knocked down. You cannot evaluate information in a field you know nothing about. You cannot connect ideas you’ve never internalized. A head full of well-organized knowledge is still the substrate that judgment runs on, the way a gamer’s ten thousand hours are the substrate that lets him read a brand-new map in seconds. The point is narrower and sharper: memory is the foundation of the building, and we’ve been grading buildings exclusively by their foundations while the world pays for what happens on the upper floors. Intelligence in the twenty-first century is measured by judgment rather than storage, and judgment is an upper-floor activity.
What We Should Be Measuring
Here’s where I’ll say the thing that gets me looks at dinner parties. I think students should be allowed to use AI. To research, to draft, to write, even to write the whole thing. What I would change, radically, is what happens next: test them, relentlessly, on comprehension. Reading isn’t the goal. Writing isn’t the goal. Comprehension is the goal, and it always was; reading and writing were simply the best comprehension-building and comprehension-revealing technologies available at the time, so long ago and so successfully that we forgot they were the means and started worshipping them as the end.
Because here’s what my CliffsNotes teachers never once did. They never asked me what the book meant, in my own words, to my own face. They never asked why it mattered, or what I would have done in the protagonist’s place, or what the author got wrong. They tested whether I knew what to write, which is a fundamentally different skill from understanding, and one I had gotten very good at. A single five-minute conversation would have sorted my essays into two piles instantly: the ones where I understood and the ones where I performed. My teachers never had a reason to ask, because both piles earned the same grade. A student can perform understanding on paper. It’s remarkably difficult to perform understanding in a conversation with someone who genuinely knows the material, as anyone who has bluffed through a job interview with an actual expert can testify.
And I’d go further, into genuinely strange territory: reading something an AI wrote, to the point of real comprehension, proves more about a person’s intellect than writing something they never understood. That inverts everything we assume about which activity is virtuous. But test it against reality. An engineer who reads AI-generated code, understands it deeply, catches its subtle mistakes, and knows when to trust it is displaying more intelligence than one who types original code by hand without grasping why it works. The comprehension is the intelligence. The typing never was.
The Information Age did not eliminate the need for intelligence. It changed the kind of intelligence that matters. What does intelligence look like when knowledge is infinite and retrieval is instant? It looks like finding: knowing what to look for and where. Evaluating: telling the credible from the confident, which becomes the whole ballgame when machines generate wrong answers as fluently as right ones. Connecting: seeing that a pattern from one domain explains a problem in another. Synthesizing, explaining, judging. The threshold between someone intelligent and someone merely equipped is no longer who possesses the answer. Both can find an answer in eight seconds. The difference is who can discern the right answer from the wrong one, and explain why, without notes, to a skeptical human being.
I know what unfakeable learning looks like because the most meaningful learning of my life happened entirely outside the grading system. When I decided to teach myself illustration, I was so intimidated by faces that for my first two years, every person I drew was faceless, a gallery of well-dressed mannequins. By the third year I risked rough features. By the fourth I could draw realistic faces entirely in vector. No one graded this. There was no rubric to reverse-engineer, no CliffsNotes for it, and no way to fake it, because the evidence of my understanding was the understanding: either the face looked human or it didn’t, and when it didn’t, I could tell you precisely why, which is how I knew I was actually learning. Every gamer knows this feeling too. Nobody has ever cheated their way to being genuinely good at a fighting game; the skill is verified sixty times a second, in conversation with an opponent, and the community can smell a carried player instantly. The most honest assessment systems humans have ever built are the ones where performance and understanding physically cannot be separated. School assessments sit at the opposite pole, and everything wrong with them flows from that separation.
Educators will recognize that I have walked up an old staircase. Bloom and his colleagues published their taxonomy in 1956. The familiar sequence of remembering, understanding, applying, analyzing, evaluating and creating belongs to the 2001 revision. The distinction matters: these frameworks already give teachers a vocabulary for goals beyond recall. My concern is that assessment can reward the easiest evidence to collect while overlooking the harder work of judgment. AI makes that imbalance harder to ignore; it does not make foundational knowledge unnecessary.
Education already uses dialogue to examine understanding. Doctoral defenses are one familiar example: a written dissertation is accompanied by questions about the reasoning behind it. The format varies across institutions, and an oral examination can introduce its own biases. It is not infallible. But it illustrates a principle that matters beyond graduate school: a document becomes stronger evidence when its author can explain, defend and revise the thinking inside it.
The objection writes itself: conversation doesn’t scale. One teacher, thirty-five students, and I’ve just prescribed the pedagogical equivalent of a private audience with each one. For two hundred years that objection ended the discussion, and rightly so; scale, remember, was the whole point of the system. But notice the strange symmetry of our moment. The same technology that broke the essay is the first technology in history that makes dialogue scalable. A machine that can converse with every student simultaneously, probe their reasoning, ask why three different ways, and surface to the teacher which students actually understand and which are performing, is a machine that dissolves the very constraint that forced us to invent the essay in the first place. I’m not naive about the execution; educational technology has a long history of overpromising. But the constraint that made proxies necessary is, for the first time since Prussia, negotiable.
The Return of Socrates
Everyone assumes AI drags education into the future. I think it may drag education back, about twenty-four centuries back, and that this would be the best thing to happen to it.

The oldest method we have is a teacher asking a student questions, face to face, following the answers wherever they lead, until the student’s actual understanding, or the absence of it, stands in the open air where both can see it. The Socrates we encounter in Plato is known for this kind of questioning. He asked, and the asking was the assessment, and no CliffsNotes has ever been printed that survives contact with a good follow-up question. Athens, it should be noted, eventually sentenced the man to death, which I suppose proves that disruptive assessment methodologies have never been popular with the administration. But the method outlived the hemlock. Dialogue is not impossible to game; students can bluff, rehearse, and charm, and examiners can be fooled like anyone else. But it is far more difficult to counterfeit once an informed questioner begins following the answers, because the dialogue goes wherever your last answer just went. There’s no rubric to reverse-engineer. There’s only what you understand, unspooling in real time, and every follow-up question raises the price of the bluff.
And there’s a delicious irony buried in Plato for our exact moment. In the Phaedrus, Socrates worries at length about a disruptive new information technology that he fears will destroy education. Students who rely on it, he warns, will stop exercising memory, will absorb the appearance of wisdom without the reality, and will seem knowledgeable while understanding nothing. The technology he was describing was writing. The parallel is striking, although the passage concerns writing broadly, not the modern school essay. Every generation since has met its version: the printing press, the calculator, Google, and now this. Socrates was wrong that writing would end education. But look closely at what he actually predicted: students absorbing the appearance of wisdom without the reality, and tell me the old man didn’t see the CliffsNotes rack coming.
The deepest educational technology was never the essay, the exam, or the chalkboard. It was the question, asked by someone who can hear the difference between an answer and an act. Everything else was scaffolding we built because questions didn’t scale. The scaffolding served us honorably. It may be time to thank it and climb down.
Every Business Runs a School
If you’ve read this far and you don’t have kids in school, you may be wondering why a branding guy is this worked up about essays. Here’s why: everything I’ve described is happening in your business, right now, wearing different clothes. Every organization is an education system. It has things it actually wants (customers who love the brand, employees who do great work, marketing that drives growth) and it has proxies it can measure. And Goodhart’s Law is as undefeated in commerce as it is in classrooms.
Marketing might be the most Goodhart-diseased discipline on earth. We wanted brand love, so we measured followers, and an entire economy of purchased followers appeared, some of whom were even human. We wanted attention, so we measured impressions, and got fraud sophisticated enough to require its own detection industry. We wanted engagement, so we measured clicks and comments, and got rage bait and engagement pods. Every vanity metric in your dashboard is a five-paragraph essay: an artifact that once correlated with something real, until the moment we started grading it, at which point an industry sprang up to produce the artifact without the reality. Your dashboard isn’t lying to you, exactly. It’s answering the question you asked. The problem is that, like my English teachers, you asked the wrong question, and your market, like any student body, is optimizing for exactly what you measure.
I’ve watched this cycle from the inside for a quarter century: circulation, then traffic, then followers, then engagement rate, each proxy honest for a honeymoon and manufactured ever after. The best marketing decision I ever made, going all-in on Anime Expo while our competitors kept buying booths at the traditional gaming shows, would have looked indefensible on every dashboard we had, because no spreadsheet contained a column for the thing I’d actually observed: people at that show crying over their first real piece of gaming gear. The dashboard measures the essay. The show floor is the conversation.
Hiring offers another example. Harvard Business School and the Burning Glass Institute examined whether removing degree requirements changed who employers hired. Their 2024 research found that the incremental opportunity created by those changes amounted to fewer than one in 700 hires across the overall labor market. That does not mean only one in 700 people is hired without a degree. It means the measured additional opportunity from these reforms was small. Some employers changed their practices meaningfully; others changed the wording of job advertisements without comparable changes in hiring. Deleting a proxy is not enough. An organization needs a credible way to evaluate the skills it says it values.
So here’s the audit I’d suggest, whether you run a classroom, a brand, or a team of nine people. Take your most important metric and ask the CliffsNotes question: could someone hit this number without the underlying reality existing at all? Could an agency deliver this engagement without anyone caring about your brand? Could a student produce this essay without understanding the book? If the answer is yes, and it’s almost always yes, then your measure has become a target, and somewhere out there, someone rational is optimizing it. Not because they’re corrupt. Because you told them to. A better detector cannot solve this by itself. I would start with the conversation: the customer interview instead of the sentiment score, the working session instead of the portfolio review, the follow-up question instead of the rubric. Proxies for scale, dialogue for truth, and never confuse which one you’re looking at.
The Essay I Finally Understand
Which brings me back to a teenager, a bookstore, and three dollars. To the A and the F, earned by the same kid on the same play, months apart.
I used to think that kid got away with something. Then for a while I thought he’d been failed by the system, which is the fashionable interpretation. I’ve landed somewhere less comfortable than either. That kid was rational. He looked at an institution, correctly identified what it measured, and delivered it efficiently, exactly the way a company delivers on its KPIs, exactly the way half of America’s teenagers are, by their own account, delivering essays right now. And here’s the part that still gets me: the books I genuinely understood in high school, I understood because something in them caught me, not because an essay was due. The system neither caused that learning nor ever once detected the difference between it and the counterfeit. Students have never optimized for learning or against it. They optimize for whatever we measure, because that’s what humans do, in every institution, at every age, often enough that the incentives deserve close attention. If schools measure essays, students will optimize essays. If schools measure memorization, students will optimize memorization. And if schools ever find the nerve to measure comprehension, students will optimize comprehension, and for once the hidden curriculum and the official one will be pointing at the same thing.
The machine that writes essays isn’t the end of education. It’s the end of pretending the essay was ever the point. Perhaps intelligence was never about possessing information. Perhaps it has always been about knowing what to do with it. If that’s true, then the arrival of AI doesn’t simply change the tools students use. It changes the very thing education should be trying to cultivate, and what comes next depends on whether we have the courage to measure the invisible thing: the understanding, the comprehension, the discernment, that the essay was only ever standing in for.
Because education, like every institution, always gets the behavior it measures.
It’s about time we measured what we actually want.
Sources
- PubMed Central (PMC), peer-reviewed analysis of annual biomedical publication volume (1.5M+ papers per year), 2024. https://pmc.ncbi.nlm.nih.gov/articles/PMC11240179/
- arXiv preprint, computational analysis of the PubMed corpus (36M+ indexed articles, 1M+ added annually), 2024. https://arxiv.org/pdf/2402.03484
- Pew Research Center, “About a Quarter of U.S. Teens Have Used ChatGPT for Schoolwork, Double the Share in 2023,” January 2025. https://www.pewresearch.org/short-reads/2025/01/15/about-a-quarter-of-us-teens-have-used-chatgpt-for-schoolwork-double-the-share-in-2023/
- Pew Research Center, “How Teens Use and View AI,” February 2026. https://www.pewresearch.org/internet/2026/02/24/how-teens-use-and-view-ai/
- OpenAI, “New AI Classifier for Indicating AI-Written Text,” updated July 20, 2023. https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/
- Liang, Weixin, et al. “GPT Detectors Are Biased Against Non-Native English Writers.” Patterns, 2023. https://www.cell.com/patterns/fulltext/S2666-3899%2823%2900130-7
- Fuller, J., et al., Harvard Business School and the Burning Glass Institute, “The Emerging Degree Reset” and follow-up research on skills-based hiring, 2022-2024. https://www.hbs.edu/bigs/joseph-fuller-college-degree-gap