AI and our measure of intelligence in education
| |

The Measure of Intelligence

Artificial intelligence didn’t break education. It exposed a question we’ve been avoiding for two hundred years: what are we actually measuring when we say someone has learned?

I remember the books just fine. Hamlet. Othello. The problem was never memory.

Here’s the confession, offered now that the statute of limitations has safely expired. Throughout high school, in honors and Advanced Placement English, whenever a book arrived with an essay attached, I would almost always reach for the little yellow-and-black booklet. If you went to school before the internet, you know the one. CliffsNotes occupied a strange shelf in the bookstore, right out in the open, sold legally to children, containing everything a teacher hoped you would discover on your own: the themes, the symbols, the character arcs, the interpretations that would earn approving comments in the margins. The booklet didn’t feel like contraband. It felt like a study aid, which is exactly what the cover said it was.

Now, before you picture a slacker, understand the scene. This was AP English. We were nerds, and mostly proud of it, and most of the time we genuinely made the effort to comprehend what we were reading, booklet or no booklet. The notes were often a supplement to understanding rather than a substitute for it. But not always. That’s the part that matters. When a book grabbed me personally, it got the full treatment: the real reading, the actual wrestling, the opinions I’d defend at lunch. When it didn’t grab me, or when the deadline was closing in, and I confess Shakespeare landed in this category more often than I’d like to admit, I’d let the booklet do the comprehending for me. I’d skim the summary, note which ideas kept repeating, pull a few lines from the play so I could quote something with page numbers, and assemble an essay that sounded like a person who had wrestled with literature.

And here’s the detail that took me thirty years to find disturbing: my grades could not tell the difference. The essays born of genuine engagement and the essays assembled from the booklet came back with the same marks and the same encouraging comments in the same red ink. Same student, same output, two completely different realities inside my head, and the instrument measuring me registered no difference whatsoever.

Well. Almost no difference. Let me give you my favorite exhibit.

Cliffnotes and smarknotes

Freshman year, first semester, the book was Romeo and Juliet, and I’ll be honest with you now in a way I wasn’t with anyone then: I didn’t read it. Not skimmed it, not started it and drifted off. What I read instead, closely, almost scholastically, was the CliffsNotes. I studied that booklet like it was the assigned text, absorbed its arguments, and then went back through the play itself hunting for quotes that would support the positions I’d borrowed. It was, in its way, a genuine research process, rigorous, methodical, and aimed entirely at the wrong document. The teacher held my essay up as the best in the class. And in the economy of that classroom, best in class carried a jackpot: an automatic A for the entire first half of the year. At fourteen, I had produced award-winning literary analysis of a play I had never read, which, in retrospect, should have been somebody’s clue.

The clue arrived at the end of the year, wearing an F. Our final exam circled back to Romeo and Juliet, and this time the questions weren’t essay-shaped. They asked what actually happened in the play, and I sat in that room discovering, in real time, that I remembered essentially nothing, because there had never been anything to remember. The booklet’s arguments had passed through me the way a script passes through an actor after the show closes. So the same institution measured me twice on the same play and returned two opposite verdicts: the essay declared me the finest literary mind in the room, and the final revealed I had learned nothing at all. The second verdict was the accurate one. And the lesson I should have drawn, though it took me three more decades, is that understanding leaves a residue and performance evaporates. The final wasn’t a better instrument in any sophisticated sense. It just showed up months later, which is long enough for a counterfeit to completely dissolve.

For most of my life, I filed all of this under mild adolescent guilt, somewhere between copying a friend’s homework and telling my parents the arcade was a library. Everybody cut corners. It wasn’t interesting.

Then, a few years ago, the world started arguing about ChatGPT in classrooms, and the memory came back with a different meaning attached.

I should say up front that I was, by the system’s official accounting, a good student, and also never a natural one, and I’ve never pretended otherwise. I struggled with basic multiplication. My grammar remains a work in progress, as my editors will confirm under oath. What I had instead was a kind of relentless practical curiosity, mostly aimed at video games. By ten I was doing summer work at hole-in-the-wall computer shops. By twelve I was play-testing games for SNK and Interplay. By fifteen I was behind the counter at Babbage’s, and every job since has been, if I’m honest, an elaborate scheme to stay close to gaming. School was something I could do well when I chose to, and the choosing depended, like everything else in my life, on whether the thing in front of me had earned my curiosity.

Which means I was precisely the kind of student the current AI panic is not supposed to be about. The panic has a character in mind: the lazy kid, the corner-cutter, the one who would learn nothing if you handed him a machine that writes essays. But I was the honor student. The AP kid. And the honor students were gaming the system too, not constantly, not carelessly, but selectively and rationally, deploying real comprehension where it felt worth the cost and manufacturing the appearance of it where it didn’t. The system as it existed in the early 1990s was already perfectly gameable by a kid with three dollars and a bookstore, and the kids gaming it were often the ones wearing the academic letter jackets. That gaming wasn’t a malfunction we discovered. It was, I’ve come to believe, the system working exactly as designed, revealing exactly what it was designed to reward.

The interesting question was never why students look for shortcuts.

The interesting question is why the shortcuts work.

The Thing Nobody Has Ever Seen

Here is a claim that sounds absurd until you sit with it: no teacher has ever observed a student’s understanding directly. Not the best teacher you ever had, not Socrates, not anyone.

They can’t. Understanding has no physical form. It can’t be weighed, photographed, or extracted for inspection. It lives entirely inside another person’s head, and the only access anyone has ever had to it is through the traces it leaves behind: the things a person says, writes, builds, and does. An essay is a trace. A test score is a trace. A diploma is a trace of traces. Each one is evidence from which we infer something invisible, the way a detective infers a burglar from footprints without ever meeting him.

A thermometer is not heat. A map is not the territory. And an essay is not understanding. It’s a measurement instrument pointed at understanding, and like all instruments, it can be fooled.

I didn’t fully grasp this distinction until I started sitting on the other side of the table, hiring people.

I’ve spent more than twenty-five years in branding and marketing, most of it in the gaming industry, and a large part of that time has involved hiring creative people: designers, artists, writers, video editors, marketers. If you’ve done a lot of hiring, you eventually make peace with a humbling fact. The résumé tells you almost nothing. The degree tells you almost nothing. Even the portfolio, which ought to be the gold standard for creative work, tells you less than you’d hope, because portfolios can be curated, coached, borrowed from group projects, and polished by friends. I have interviewed candidates with immaculate portfolios who could not explain a single decision in their own work, and I have hired people with thin, scrappy portfolios who turned out to be the best creative minds I’ve ever worked with.

c226a407 b83a 4b25 b97b 91c87c150d46

What actually tells you something is conversation. Ask a designer why she chose that typeface and watch what happens. The candidate who understands her own work lights up and starts arguing with herself, telling you what she’d change, where the client pushed back, what she’d do differently with another week. The candidate who performed the work, or had it performed for him, gives you back adjectives. “It felt clean.” “It matched the brand.” The portfolio was the essay. The conversation is where the understanding either shows up or doesn’t.

Employers rely on résumés for the same reason teachers rely on essays: nobody has the time to do the thing that would actually work. If a company had unlimited time, it would never read a résumé again. It would spend six months working alongside every applicant, watching how they take criticism, what they do when a project collapses on a Thursday afternoon, whether they admit mistakes or relocate blame. Those qualities determine whether someone succeeds in a job, and none of them appears on paper. The résumé exists because those qualities only reveal themselves over time, and time is the one thing no hiring manager has.

So the degree becomes shorthand for knowledge. The previous job title becomes shorthand for competence. The credit score becomes shorthand for whether you’ll repay the loan, which is the only thing the bank ever cared about. The quarterly earnings report becomes shorthand for the long-term health of a business. Civilization runs on proxies because the realities we care about are invisible, slow, or both.

There’s nothing wrong with this. Proxies are one of humanity’s great inventions. The problem starts later, and it starts quietly, when we forget that the proxy is a proxy.

The Performance of Understanding

No institution wakes up one morning and decides that the measurement matters more than the thing being measured. The drift is gradual, and it happens because measurements have qualities that realities don’t. Measurements can be counted, compared, standardized, archived, reported to administrators, and graphed over time. Understanding resists every one of those conveniences. Given a choice between something manageable and something elusive, institutions will drift toward the manageable every single time, without a villain anywhere in the story.

Picture a teacher who assigns an essay because thoughtful writing genuinely reflects thoughtful reading, which it does. At first the assignment works beautifully. Students who engage deeply write better essays than students who don’t. The essay is functioning exactly as intended: not as learning itself, but as evidence of learning.

Then the essay becomes important. Grades depend on it. College admissions depend on grades. Scholarships depend on admissions. Parents ask how to earn higher scores. Test-prep companies publish model essays. Teachers, trying to be fair, hand out rubrics specifying exactly what earns an A, and the rubric becomes a blueprint. Students trade reliable thesis structures the way my generation traded game cheat codes. Every step in this chain is reasonable. Every actor is responding sensibly to incentives they didn’t create. And the cumulative result is that students get progressively better at producing the appearance of understanding, which sometimes reflects the real thing and sometimes doesn’t, and from the outside the two become harder and harder to tell apart.

Kids figure this out astonishingly early, and adults consistently misread what they’re seeing. Long before children can do algebra, they’ve learned to ask whether there’s partial credit. Long before they can appreciate a novel, they’ve learned to ask whether it will be on the test. We hear those questions and diagnose laziness. I think we’re witnessing something closer to brilliance. Children are extraordinary students of institutions. Nobody teaches them this; they simply observe, with the unclouded eyes of anthropologists, that every school contains two curricula. The official curriculum says understand Shakespeare. The hidden curriculum explains how Shakespeare will be graded. The official curriculum celebrates curiosity. The hidden curriculum pays out on performance. When the two drift apart, students don’t cause the drift. They’re just the first to notice it, and the first to act on what they see.

Economists gave this pattern a name decades ago, though the name usually gets misattributed. In 1975, the British economist Charles Goodhart, writing about monetary policy, observed that any statistical regularity tends to collapse once pressure is placed on it for control purposes. The catchier version everyone quotes, that when a measure becomes a target it ceases to be a good measure, was actually coined later by the anthropologist Marilyn Strathern, summarizing him. Around the same time as Goodhart, the social scientist Donald Campbell arrived at nearly the same conclusion from the direction of education, warning that the more any quantitative indicator is used for decision-making, the more it will distort and corrupt the very process it was meant to monitor. Campbell was writing specifically about standardized testing in American schools in 1976, which means the diagnosis predates the personal computer, never mind the chatbot.

Neither man was describing dishonesty. They were describing physics. People optimize for whatever the system rewards; that is close to a law of nature. Companies optimize quarterly earnings. Hospitals optimize wait-time metrics. Police departments optimize crime statistics. Social platforms optimize engagement. Students optimize grades. The remarkable thing is not that people do this. The remarkable thing is how reliably institutions mistake the optimization for corruption, as if human beings should be expected to ignore the scoreboard hanging directly over their heads.

Teenage me, with the yellow-and-black booklet, was not corrupt. Teenage me had simply read the scoreboard correctly. The essay assignment asked, on its surface, “do you understand this book?” But the grading, the rubric, and the entire apparatus around it asked a different question: “can you produce the artifact we associate with understanding?” I answered the question that was actually being asked. Students always do.

Which brings us, finally, to the machines. But to understand why artificial intelligence detonated inside education the way it did, you first have to understand what schools were built to do. And for that we have to be fair to the schools, which almost nobody in this debate is willing to be.

In Defense of the Factory School

There’s a critique of education so common it has become furniture. You’ve heard it: schools were designed on a factory model, in rows, with bells, to produce obedient industrial workers, and that’s why they’re broken. The critique contains real truth, and it’s also deeply unfair, in the way that judging your grandparents by problems they couldn’t have imagined is unfair. Because the system everyone loves to condemn solved one of the greatest bottlenecks in human history, and it solved it so completely that we’ve forgotten the bottleneck ever existed.

For nearly all of civilization, knowledge was heartbreakingly fragile. A farmer might understand his local soil and weather better than any living person, and that understanding died with him. A master stonemason spent fifty years perfecting techniques an apprentice could inherit only through a decade of watching. Physicians accumulated experience one patient at a time and carried it to the grave. Books, even after Gutenberg, were precious objects rather than everyday possessions, and most people couldn’t read them anyway. Ideas traveled at the speed of ships and horses. A brilliant insight in one city could take a generation to reach another, and entire populations lived and died without encountering the most important discoveries of their own era.

Against that backdrop, the challenge facing education was almost the opposite of ours: how do you transfer as much knowledge as possible into as many heads as possible, using nothing but voices, chalk, and paper?

Consider what an expert meant under those conditions. An eighteenth-century physician standing over a patient had no database, no reference app, no colleague reachable in under a week. The era’s version of cloud backup was an older doctor who hadn’t died yet. The knowledge available to him consisted entirely of what he had managed to store inside his own skull. Every disease he could recognize, every treatment he could recall, every anatomical relationship he had drilled into memory might determine whether the person in front of him lived. Knowledge that wasn’t memorized effectively didn’t exist. The same was true for the lawyer arguing without archives, the navigator crossing an ocean with charts and stars, the engineer building bridges centuries before simulation software. Memorization wasn’t the enemy of intelligence, the way we casually treat it today. For most of human history, memorization was intelligence, or at least the closest usable approximation of it.

Schools evolved to serve that world, and the much-maligned Prussians got there first. After Napoleon humiliated Prussia at Jena in 1806, its reformers concluded that national survival required an educated population, not just educated elites, and it turns out nothing accelerates education reform quite like losing a war to France. They built the first serious attempt at universal, state-run schooling: trained teachers, standardized curricula, compulsory attendance, and yes, discipline and uniformity that modern educators rightly wince at. Leave it to the Prussians to respond to a military catastrophe by inventing compulsory homework. Horace Mann toured those schools in the 1840s and came home to Massachusetts convinced that America needed its own version, not to manufacture obedient workers but because he believed education was the great equalizer of the human condition, the one machine capable of making knowledge a public inheritance rather than an aristocratic one.

It’s worth pausing on how radical that was. For thousands of years, knowledge had been the private property of clergy, aristocrats, and guilds. Mann and the reformers proposed to give it away, to everyone, at public expense. The standardized features we now find soulless were the price of that ambition. Standardized curricula meant a child in one town learned roughly what a child in another town learned. Exams verified the knowledge had actually landed. Essays forced students to synthesize rather than merely recognize. Diplomas certified a minimum threshold to strangers who would never meet your teachers. None of it was perfect, and perfection was never the goal. Scale was. Socrates could spend an afternoon interrogating a single student because he was teaching a handful of Athenians; a nation trying to educate every child confronts arithmetic Socrates never faced.

And the compromise worked. Literacy went from privilege to default across Europe and North America in roughly a century. The system helped produce the scientists, engineers, physicians, and yes, the marketers, who built the modern world. If today’s education system deserves criticism, it also deserves credit for most of the things critics use to publish their criticism.

I insist on this defense for a reason beyond fairness. It guards against the laziest explanation of institutional failure: that institutions persist because people are too stubborn to change them. That’s rarely true. Institutions persist because they keep solving the problem they were designed to solve. The trouble begins when the problem itself quietly changes, and the institution, still functioning perfectly, keeps solving the old one. Education optimized for scarcity. We now live in abundance.

Sometime in the last half century, the problem changed. There was no headline. But humanity crossed a threshold that I’d argue matters more than the invention of any particular technology, including the one currently terrifying every school board in America.

When Knowledge Outgrew the Knower

For most of history, becoming an expert meant accumulating enough information to stand among the most knowledgeable people in a field, and the ideal, however difficult, remained conceivable. A physician could plausibly aspire to know the important medicine of his generation. A lawyer could hold the relevant law in his head. Even physics, for a while, fit inside a person: a well-trained physicist in 1900 could be genuinely conversant with most of what physics knew.

Then the twentieth century happened to knowledge. Fields didn’t just grow; they split, and the splits split. Physics alone birthed quantum mechanics, then particle physics, cosmology, condensed matter, string theory, each a paradigm that would have consumed an entire nineteenth-century career just to enter. Discoveries began arriving faster than textbooks could be revised. The twentieth century didn’t merely produce more knowledge than previous centuries. At some point, it produced more knowledge than any individual could possess, and that’s a different kind of event.

The numbers today are almost comic. Biomedical research alone now generates over 1.5 million new papers every year, and PubMed, the database that tries to track it all, indexes more than 36 million articles and adds over a million annually. A physician who read one paper every ten minutes, around the clock, skipping sleep, meals, and, inconveniently, all of her actual patients, would still fall further behind every single day. No physician alive can read a meaningful fraction of the literature in her own specialty, let alone medicine broadly; if she read one paper a minute, around the clock, she would fall further behind every single day, which is a soothing thing to contemplate in a waiting room. Software engineers work inside ecosystems of millions of libraries that mutate weekly. And the doctor you trust with your life? You could ask her a specific question about a procedure, and she might tell you, without embarrassment, that she’d have to look it up. We don’t consider her a fraud for that. We consider her a professional, because we’ve quietly accepted, in practice if not in our schools, that expertise stopped meaning “contains the information” a while ago.

Here’s a version of this problem I’ve never seen in an education op-ed, but it’s the one that made it click for me. Think about how history is taught. American students study American history, reasonably enough, and it spans about 250 years. A student in China faces thousands of years of continuous recorded civilization. A student in Iran or Egypt, older still. No Chinese student, however brilliant, actually learns Chinese history in any complete sense. They learn a curated sliver, the dynastic highlight reel, and we’d never call them ignorant for it, because the alternative is impossible. Now notice that every field is becoming China. Medicine is China. Law is China. Computer science became China in about a decade and a half. The territory has grown so vast that every expert is now working from a highlight reel, and the honest ones know it. The people we call experts today are, in the old sense of the word, superficial experts, not because they got lazier but because civilization got bigger than any one mind. Experts didn’t become less knowledgeable. Knowledge became larger than expertise.

That sentence is, I think, the defining condition of the twenty-first century, and I watched it happen inside my own industry in real time. When I started in gaming, a dedicated person could genuinely keep up: the platforms, the studios, the scene. Somewhere in the 2000s that stopped being possible by any amount of personal diligence. The industry fragmented into more games, genres, platforms, communities, and micro-cultures than anyone could track, and my job, which depends entirely on staying current, quietly transformed from a memory problem into a navigation problem. My solution was unorthodox: I learned to treat my own kid as a discovery system. When my son started watching an unknown streamer named Lirik on an obscure site called Twitch in 2013, that signal was worth more than any industry report I could have memorized, and the VCs who dismissed Twitch as kid stuff learned its value the following year, when Amazon paid a billion dollars for it. The lesson wasn’t that my son knew more than the VCs. It’s that in a world of infinite information, the scarce skill had become knowing where to look and what mattered, not what to retain.

For thousands of years, knowledge was something individuals accumulated. It has become something individuals navigate. Those are profoundly different activities, requiring different skills, deserving different measurements.

And yet our schools, built with love and genius for the world of accumulation, kept right on measuring retention and reproduction, because the instruments still worked, more or less, and nothing had come along to prove otherwise. What does it mean to be educated in an age when knowledge is infinite? For fifty years, that question sat in the room, unasked, because no one was forcing the issue.

Then something came along.

The Exposure

The speed of what happened next is worth stating plainly, because I don’t think most adults have absorbed it. In 2023, 13 percent of American teens said they used ChatGPT for schoolwork. A year later, it was 26 percent. By late 2025, more than half of all teens were using AI chatbots for schoolwork, and 59 percent said that cheating with AI had become a regular feature of life at their school. That’s not adoption. That’s a dam breaking.

And here’s the detail from the survey data that the panic coverage always skips, because it ruins the “lazy kids” narrative. The teens have a moral framework, and it’s more coherent than the adults’. In the Pew research, a majority of teens said using ChatGPT to research new topics was acceptable, while only 18 percent said the same about using it to write essays. Sit with that. The students, unprompted, distinguished between using the machine to access knowledge and using it to fake the evidence of knowledge. They drew the line exactly where a philosopher would draw it, between the tool and the fraud, and they drew it while the adults in the room were still arguing about whether the tool should exist. Students don’t lack an ethics of AI. They lack an assessment system whose ethics make sense to them.

The institutional response was an arms race, and the arms race went badly. Schools bought detection software. Teachers ran essays through classifiers. And the classifiers turned out to be roughly as reliable as a mood ring. OpenAI, the company that built ChatGPT and presumably understands its output better than anyone, released its own AI-text detector and then shut it down within months for its low rate of accuracy; at the time of its launch, it correctly caught only about a quarter of AI-written text while falsely accusing human writers 9 percent of the time. Commercial detectors performed variably, with false positives that fell disproportionately on non-native English speakers, meaning the students most likely to be accused of not writing their own words were the ones for whom writing English words was hardest. The detectors couldn’t reliably do it because there is nothing there to detect. Competent prose is competent prose. Which is precisely the problem, and it’s a much older problem than anyone wants to admit.

Because think about what the panic is actually conceding. If a machine-generated essay is indistinguishable from a B-plus student essay, then the B-plus student essay was never demonstrating what we told ourselves it demonstrated. The essay was measuring the ability to produce essay-shaped text: organized, on-topic, gesturing at themes. That ability used to correlate with understanding well enough to be a useful proxy, the way a résumé used to correlate with competence. ChatGPT didn’t sever the correlation. It revealed how much of what we were grading was the shape, not the substance. My CliffsNotes essays proved the same thing in miniature, decades earlier, at much lower bandwidth. AI just industrialized the demonstration.

This is why I’ve become convinced that the technology isn’t the story. Artificial intelligence didn’t create this problem. It merely exposed it, and the exposure raises a far more uncomfortable question than anything about cheating: if a machine can produce the work we’ve always accepted as evidence of learning, were we ever measuring learning correctly in the first place? We spent two centuries perfecting instruments for measuring the reproduction of knowledge, in a world where reproduction was the bottleneck, and then the bottleneck moved, and we kept using the instruments. AI walked in, aced every one of them without understanding anything, and in doing so asked the question the instruments were built to avoid: fine, but what did the student comprehend?

The Last Closed-Book Room

Before deciding what schools should do about any of this, it’s worth noticing something about the world students graduate into, because that world settled the question years ago and nobody sent the schools a memo.

Every serious profession has already gone open-book. Your surgeon does not perform from memory alone; she operates inside a lattice of checklists, imaging, protocols, and reference systems, and the profession considers this a triumph, not a scandal, because the checklist-heavy era of medicine is dramatically safer than the memory-heavy one it replaced. I could ask a working surgeon a specific technical question about a procedure adjacent to her specialty, and she might tell me, comfortably, that she’d need to look it up before answering, and no one in the room would think less of her. What we’d judge her on is whether she looked in the right place, understood what she found, and acted on it correctly. Pilots, the most trusted professionals on earth, are famous precisely for refusing to rely on memory; the pre-flight checklist exists because aviation learned, in blood, that recall under pressure is exactly the wrong thing to bet lives on. Nobody boards a plane hoping the captain memorized the manual and is winging it from recall. Lawyers practice law with databases open. Engineers design with references at their elbow. Programmers, and I say this with affection for every programmer I’ve ever employed, have spent two decades openly joking that their profession is professional-grade searching, and their industry runs the world.

In other words, the entire adult economy quietly converted to the model where intelligence means finding, evaluating, and applying, and where the shame is not in looking things up but in failing to understand what you found. There is exactly one place left where a person’s worth is still determined by what they can produce alone, from memory, with their references confiscated: the classroom. We built a closed-book examination system to prepare people for a closed-book world, the world of the eighteenth-century physician, and then the world went open-book and we kept the exam. Students can see this. They are, remember, extraordinary students of institutions, and they can observe with their own eyes that no working adult in their lives operates the way a final exam demands. When students sense that an assessment measures a skill the real world visibly doesn’t use, the assessment loses its moral authority long before it loses its official one, and technologies like ChatGPT don’t create that erosion. They just give it a tool.

None of this means memory doesn’t matter; that’s the strawman version of the argument, and it deserves to be knocked down. You cannot evaluate information in a field you know nothing about. You cannot connect ideas you’ve never internalized. A head full of well-organized knowledge is still the substrate that judgment runs on, the way a gamer’s ten thousand hours are the substrate that lets him read a brand-new map in seconds. The point is narrower and sharper: memory is the foundation of the building, and we’ve been grading buildings exclusively by their foundations while the world pays for what happens on the upper floors. Intelligence in the twenty-first century is measured by judgment rather than storage, and judgment is an upper-floor activity.

What We Should Be Measuring

Here’s where I’ll say the thing that gets me looks at dinner parties. I think students should be allowed to use AI. To research, to draft, to write, even to write the whole thing. What I would change, radically, is what happens next: test them, relentlessly, on comprehension. Reading isn’t the goal. Writing isn’t the goal. Comprehension is the goal, and it always was; reading and writing were simply the best comprehension-building and comprehension-revealing technologies available at the time, so long ago and so successfully that we forgot they were the means and started worshipping them as the end.

Because here’s what my CliffsNotes teachers never once did. They never asked me what the book meant, in my own words, to my own face. They never asked why it mattered, or what I would have done in the protagonist’s place, or what the author got wrong. They tested whether I knew what to write, which is a fundamentally different skill from understanding, and one I had gotten very good at. A single five-minute conversation would have sorted my essays into two piles instantly: the ones where I understood and the ones where I performed. My teachers never had a reason to ask, because both piles earned the same grade. A student can perform understanding on paper. It’s remarkably difficult to perform understanding in a conversation with someone who genuinely knows the material, as anyone who has bluffed through a job interview with an actual expert can testify.

And I’d go further, into genuinely strange territory: reading something an AI wrote, to the point of real comprehension, proves more about a person’s intellect than writing something they never understood. That inverts everything we assume about which activity is virtuous. But test it against reality. An engineer who reads AI-generated code, understands it deeply, catches its subtle mistakes, and knows when to trust it is displaying more intelligence than one who types original code by hand without grasping why it works. The comprehension is the intelligence. The typing never was.

The Information Age did not eliminate the need for intelligence. It changed the kind of intelligence that matters. What does intelligence look like when knowledge is infinite and retrieval is instant? It looks like finding: knowing what to look for and where. Evaluating: telling the credible from the confident, which becomes the whole ballgame when machines generate wrong answers as fluently as right ones. Connecting: seeing that a pattern from one domain explains a problem in another. Synthesizing, explaining, judging. The threshold between someone intelligent and someone merely equipped is no longer who possesses the answer. Both can find an answer in eight seconds. The difference is who can discern the right answer from the wrong one, and explain why, without notes, to a skeptical human being.

I know what unfakeable learning looks like because the most meaningful learning of my life happened entirely outside the grading system. When I decided to teach myself illustration, I was so intimidated by faces that for my first two years, every person I drew was faceless, a gallery of well-dressed mannequins. By the third year I risked rough features. By the fourth I could draw realistic faces entirely in vector. No one graded this. There was no rubric to reverse-engineer, no CliffsNotes for it, and no way to fake it, because the evidence of my understanding was the understanding: either the face looked human or it didn’t, and when it didn’t, I could tell you precisely why, which is how I knew I was actually learning. Every gamer knows this feeling too. Nobody has ever cheated their way to being genuinely good at a fighting game; the skill is verified sixty times a second, in conversation with an opponent, and the community can smell a carried player instantly. The most honest assessment systems humans have ever built are the ones where performance and understanding physically cannot be separated. School assessments sit at the opposite pole, and everything wrong with them flows from that separation.

Educators will recognize that I’ve just walked up a very old staircase. In 1956, Benjamin Bloom and his colleagues published their taxonomy of educational objectives, ranking cognitive skills from remembering at the bottom, up through understanding, applying, analyzing, evaluating, and creating. Teachers have had this map for seventy years. The honest reading of our situation is that we built our entire industrial-scale measurement apparatus around the bottom two floors because they were cheap to test at scale, and then AI came along and automated the bottom two floors. Everything schools are struggling to measure right now lives upstairs, exactly where Bloom said the good stuff was all along. Comprehension, judgment, and discernment are replacing recall as the defining intellectual skills, and the map for teaching them has been hanging in the faculty lounge for seventy years.

And here’s the detail that makes the whole argument feel less radical than it sounds: education already knows this. At the very top of the system, where the stakes are highest, we never trusted the document alone. A PhD, the highest credential the academy awards, is not conferred when the dissertation is submitted. It’s conferred after the defense, an oral examination in which a committee of experts interrogates the candidate about the work, live, for hours, precisely because a dissertation could in principle be produced by someone who doesn’t understand it, and a defense cannot be survived by someone who doesn’t. Medical students face oral boards. Trial lawyers face judges. The military war-games its officers with questions, not essays. Everywhere that an institution absolutely cannot afford to be fooled by a well-shaped artifact, the institution reverts to conversation. We reserved the honest measurement for the elite end of education and gave everyone else the proxy, not out of malice but out of arithmetic. Which means the question in front of us was never whether dialogue is the superior assessment. Everyone already agrees it is. The question was only ever whether we could afford it.

The objection writes itself: conversation doesn’t scale. One teacher, thirty-five students, and I’ve just prescribed the pedagogical equivalent of a private audience with each one. For two hundred years that objection ended the discussion, and rightly so; scale, remember, was the whole point of the system. But notice the strange symmetry of our moment. The same technology that broke the essay is the first technology in history that makes dialogue scalable. A machine that can converse with every student simultaneously, probe their reasoning, ask why three different ways, and surface to the teacher which students actually understand and which are performing, is a machine that dissolves the very constraint that forced us to invent the essay in the first place. I’m not naive about the execution; educational technology has a long history of overpromising. But the constraint that made proxies necessary is, for the first time since Prussia, negotiable.

The Return of Socrates

Everyone assumes AI drags education into the future. I think it may drag education back, about twenty-four centuries back, and that this would be the best thing to happen to it.

Socrates top

The oldest method we have is a teacher asking a student questions, face to face, following the answers wherever they lead, until the student’s actual understanding, or the absence of it, stands in the open air where both can see it. Socrates taught this way exclusively. He wrote nothing, tested nothing, assigned nothing. He asked, and the asking was the assessment, and no CliffsNotes has ever been printed that survives contact with a good follow-up question. Athens, it should be noted, eventually sentenced the man to death, which I suppose proves that disruptive assessment methodologies have never been popular with the administration. But the method outlived the hemlock. Dialogue is not impossible to game; students can bluff, rehearse, and charm, and examiners can be fooled like anyone else. But it is far more difficult to counterfeit once an informed questioner begins following the answers, because the dialogue goes wherever your last answer just went. There’s no rubric to reverse-engineer. There’s only what you understand, unspooling in real time, and every follow-up question raises the price of the bluff.

And there’s a delicious irony buried in Plato for our exact moment. In the Phaedrus, Socrates worries at length about a disruptive new information technology that he fears will destroy education. Students who rely on it, he warns, will stop exercising memory, will absorb the appearance of wisdom without the reality, and will seem knowledgeable while understanding nothing. The technology he was describing was writing. The original AI panic, delivered twenty-four hundred years early, was about the essay itself. Every generation since has met its version: the printing press, the calculator, Google, and now this. Socrates was wrong that writing would end education. But look closely at what he actually predicted: students absorbing the appearance of wisdom without the reality, and tell me the old man didn’t see the CliffsNotes rack coming.

The deepest educational technology was never the essay, the exam, or the chalkboard. It was the question, asked by someone who can hear the difference between an answer and an act. Everything else was scaffolding we built because questions didn’t scale. The scaffolding served us honorably. It may be time to thank it and climb down.

Every Business Runs a School

If you’ve read this far and you don’t have kids in school, you may be wondering why a branding guy is this worked up about essays. Here’s why: everything I’ve described is happening in your business, right now, wearing different clothes. Every organization is an education system. It has things it actually wants (customers who love the brand, employees who do great work, marketing that drives growth) and it has proxies it can measure. And Goodhart’s Law is as undefeated in commerce as it is in classrooms.

Marketing might be the most Goodhart-diseased discipline on earth. We wanted brand love, so we measured followers, and an entire economy of purchased followers appeared, some of whom were even human. We wanted attention, so we measured impressions, and got fraud sophisticated enough to require its own detection industry. We wanted engagement, so we measured clicks and comments, and got rage bait and engagement pods. Every vanity metric in your dashboard is a five-paragraph essay: an artifact that once correlated with something real, until the moment we started grading it, at which point an industry sprang up to produce the artifact without the reality. Your dashboard isn’t lying to you, exactly. It’s answering the question you asked. The problem is that, like my English teachers, you asked the wrong question, and your market, like any student body, is optimizing for exactly what you measure.

I’ve watched this cycle from the inside for a quarter century: circulation, then traffic, then followers, then engagement rate, each proxy honest for a honeymoon and manufactured ever after. The best marketing decision I ever made, going all-in on Anime Expo while our competitors kept buying booths at the traditional gaming shows, would have looked indefensible on every dashboard we had, because no spreadsheet contained a column for the thing I’d actually observed: people at that show crying over their first real piece of gaming gear. The dashboard measures the essay. The show floor is the conversation.

Hiring is even more damning, because we have receipts. Over the past decade, companies loudly embraced skills-based hiring; degree requirements dropped from about 51 percent of job postings in 2017 to roughly 44 percent by 2024. Then researchers at Harvard Business School and the Burning Glass Institute checked what actually changed, and found that at firms that had publicly dropped degree requirements, fewer than 1 in 700 new hires was actually affected. Companies deleted the proxy from the job posting and kept grading the proxy in the screening process, because they had built no instrument for measuring the real thing. It’s the corporate version of banning ChatGPT while keeping the take-home essay. Meanwhile, at the few firms that genuinely changed how they evaluate people, non-degree hires stuck around at higher rates than their credentialed colleagues. Measuring the real thing is expensive and slow. It also works.

So here’s the audit I’d suggest, whether you run a classroom, a brand, or a team of nine people. Take your most important metric and ask the CliffsNotes question: could someone hit this number without the underlying reality existing at all? Could an agency deliver this engagement without anyone caring about your brand? Could a student produce this essay without understanding the book? If the answer is yes, and it’s almost always yes, then your measure has become a target, and somewhere out there, someone rational is optimizing it. Not because they’re corrupt. Because you told them to. The fix is never a better detector. The fix is the conversation: the customer interview instead of the sentiment score, the working session instead of the portfolio review, the follow-up question instead of the rubric. Proxies for scale, dialogue for truth, and never confuse which one you’re looking at.

The Essay I Finally Understand

Which brings me back to a teenager, a bookstore, and three dollars. To the A and the F, earned by the same kid on the same play, months apart.

I used to think that kid got away with something. Then for a while I thought he’d been failed by the system, which is the fashionable interpretation. I’ve landed somewhere less comfortable than either. That kid was rational. He looked at an institution, correctly identified what it measured, and delivered it efficiently, exactly the way a company delivers on its KPIs, exactly the way half of America’s teenagers are, by their own account, delivering essays right now. And here’s the part that still gets me: the books I genuinely understood in high school, I understood because something in them caught me, not because an essay was due. The system neither caused that learning nor ever once detected the difference between it and the counterfeit. Students have never optimized for learning or against it. They optimize for whatever we measure, because that’s what humans do, in every institution, at every age, without exception. If schools measure essays, students will optimize essays. If schools measure memorization, students will optimize memorization. And if schools ever find the nerve to measure comprehension, students will optimize comprehension, and for once the hidden curriculum and the official one will be pointing at the same thing.

The machine that writes essays isn’t the end of education. It’s the end of pretending the essay was ever the point. Perhaps intelligence was never about possessing information. Perhaps it has always been about knowing what to do with it. If that’s true, then the arrival of AI doesn’t simply change the tools students use. It changes the very thing education should be trying to cultivate, and what comes next depends on whether we have the courage to measure the invisible thing: the understanding, the comprehension, the discernment, that the essay was only ever standing in for.

Because education, like every institution, always gets the behavior it measures.

It’s about time we measured what we actually want.

Sources

  1. PubMed Central (PMC), peer-reviewed analysis of annual biomedical publication volume (1.5M+ papers per year), 2024. https://pmc.ncbi.nlm.nih.gov/articles/PMC11240179/
  2. arXiv preprint, computational analysis of the PubMed corpus (36M+ indexed articles, 1M+ added annually), 2024. https://arxiv.org/pdf/2402.03484
  3. Pew Research Center, “About a Quarter of U.S. Teens Have Used ChatGPT for Schoolwork, Double the Share in 2023,” January 2025. https://www.pewresearch.org/short-reads/2025/01/15/about-a-quarter-of-us-teens-have-used-chatgpt-for-schoolwork-double-the-share-in-2023/
  4. Pew Research Center, “How Teens Use and View AI,” February 2026. https://www.pewresearch.org/internet/2026/02/24/how-teens-use-and-view-ai/
  5. K-12 Dive, “Double the Teens Using ChatGPT for Schoolwork,” 2025. https://www.k12dive.com/news/double-the-teens-using-chatgpt-for-schoolwork/739616/
  6. TechCrunch, “OpenAI Scuttles AI-Written Text Detector Over Low Rate of Accuracy,” July 2023. https://techcrunch.com/2023/07/25/openai-scuttles-ai-written-text-detector-over-low-rate-of-accuracy/
  7. EdCafe AI, “Goodbye, AI Detection Tools,” on detector false positives and non-native English speakers, 2024. https://www.edcafe.ai/blog/goodbye-ai-detection-tools
  8. MajorMatch, “Employers Hiring Without Degrees,” degree-requirement trends in U.S. job postings 2017-2024, 2026. https://majormatch.us/blog/employers-hiring-without-degrees-2026/
  9. Fuller, J., et al., Harvard Business School and the Burning Glass Institute, “The Emerging Degree Reset” and follow-up research on skills-based hiring, 2022-2024. https://www.hbs.edu/bigs/joseph-fuller-college-degree-gap

Share This Article:

Similar Posts

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.