About Justin Lerma: AI educator and thought leader focused on the intersection of technology and human performance. Views are my own.

Disclaimer: The views expressed in this publication are personal opinions and do not represent the positions of any employer or affiliate.

© 2025 Justin Lerma. All rights reserved. Unauthorized reproduction or distribution of this content without express written permission is prohibited.

The Gadfly Discipline: Critical Thinking in the Age of AI

Share
The Gadfly Discipline: Critical Thinking in the Age of AI

Artificial intelligence has given us something remarkable: the ability to collapse distance. Ideas that once took days, weeks, sometimes years to test now move from concept to prototype in an afternoon. I genuinely appreciate that. Getting from point A to point B on a shorter line is a real gift, and nothing here argues against it.

But shortening the line changes what happens along it. The business world is obsessed with that shortened line: concept to prototype to production, faster every quarter. Growth quarter over quarter, month over month, week over week, day over day, and at this rate, hour over hour is next. That's a model built to break eventually. What worries me more is what it breaks on the way there: a faculty that has nothing to do with output and everything to do with what only the human mind can still do, sit with a hard problem long enough to actually understand it, instead of accepting the first answer that sounds right.

More companies are saying the quiet part out loud: the generation inheriting this world inherits a diminishing version of that faculty, generation over generation. You can already see it in how people talk through problems. The space they're willing to sit inside (the depth, the range of angles considered) narrows exactly as answers get faster, latency tolerance reduces and accuracy ratings climb. Which raises an uncomfortable question: accurate according to whom, once nobody's left in the room with enough domain expertise to check the machine's work?

Stay with me through the diagnosis. It builds straight into a practice you can start using this week.

The Evidence Is Already In

This doesn't need a long argument, because most of you have already lived it. AI slop is not an abstraction: it's the LinkedIn post nobody asked for, the meeting summary that says nothing you didn't already know, the deck that took three prompts and zero thought before it hit your inbox. I wrote a full post about this a while back, AI Slop: What It Is, How to Avoid It, and How to Call It Out Constructively, because it was already impossible to ignore, and it's only gotten louder since. Nate B Jones and plenty of others in the AI commentary space are tracking the same pattern in real time.

Here's the part that should worry a business more than the noise: a report out of MIT last August found that just 5 percent of enterprise generative AI pilots reach production with measurable ROI within six months. Not because the models are weak. Because organizations keep skipping the friction of rethinking the workflow the AI gets bolted onto. That's not a measurement problem, attribution modeling is a solved discipline. It's a thinking problem wearing a measurement costume. Nobody sat with the actual problem long enough to know what to build, so nobody can explain what they built is worth.

I've watched this pattern play out inside real companies: thirty emails, five meetings, two subcommittees, one actual committee, and, eventually, a brilliant executive announcing the problem wasn't a problem in the first place. AI didn't create that pattern. It just gave it a faster car.

What Critical Thinking Actually Means

It's easy to reduce critical thinking to "knowing how to think well" or "having good information." That's too thin to build anything on. For a working definition, it helps to go back to someone doing this before critical thinking was even a phrase.

Socrates (my philosophy degree is partly responsible for how much space he takes up in my head) built a method around one habit: take an idea down to its base propositions, then hold it up against every situation you can throw at it, including the unlikely ones, until you know whether it survives contact with reality or just sounds good in the room where it was born.

That's the definition I'd propose: critical thinking is, in part, a willingness to let go of a preconceived idea, even an uncomfortable one like a political stance, the moment better information shows up. It comes with an obligation attached: keep auditing what you currently believe, on a loop, genuinely open to dropping any of it at any time. Not once. Continuously.

It's worth being honest about how that ended for him. Socrates didn't retire on his reputation. He was tried and executed, pushed there by a political establishment that wanted his questioning stopped. I think about that parallel more than I'd like to when I watch a modern executive face down someone bold enough to question a decade-old methodology. There's a real, short-term cost to challenging how things are done, one that only goes away once everyone else catches up to what you already saw. No AI upgrade changes that. The tools get smarter. The cost of being the one who spoke up first doesn't.

Two Sycophants, Same Failure

There's a mechanism underneath all of this that doesn't get named enough, and it connects to something I wrote about in The Loki Effect: AI systems are tuned to be liked. They're optimized for comfort and approval, not the productive discomfort real learning requires, the "wait, that doesn't hold up" moment cognitive science calls productive failure, which anyone who's mastered anything difficult knows was necessary. An AI that only confirms what you already believe isn't teaching you anything. It's returning your own reflection with better formatting.

Now hold that next to why so many executives have fallen for AI as fast as they have. Some of it is genuine capability. But I'd argue some of it is that AI hands them a frictionless version of what they already wanted: a tool that agrees with the current solution, validates the current paradigm, and never asks the room to sit with an uncomfortable question. Executive sycophancy didn't start with AI. It's been a boardroom dynamic since there were boardrooms. AI just automated it and made it available on demand.

So the fix has to work on both ends. Our AI systems need to inherit some of that same gadfly instinct Socrates claimed for himself, instead of being tuned purely for agreement, and leaders need to invite the gadflies, the human ones, into the room on purpose, instead of quietly selecting for people who won't make them uncomfortable. Judgment is already the scarce resource once you've automated coordination, I made that case in The New Playbook: The Operator Model for an Agentic Future. What I'd add now: judgment doesn't survive in a room engineered to reward agreement. It has to be given somewhere to stand.

What This Actually Looks Like Inside a Business

Here's the tension nobody resolves cleanly: a business has to project confidence. Stakeholders don't fund "we're honestly not sure." Customers don't want to hear "we might be wrong about this." That confidence has to show up everywhere, in the proposal, in the meeting, in the deck. It's what gets you funded, hired, or believed.

But behind that certainty, the actual work has to stay deliberately uncertain: constantly second-guessed, stress-tested, held loosely enough to let go of the moment new information shows up. The businesses that win the next decade will be the ones whose leaders get comfortable saying, out loud, "I don't fully know what this experiment will show, but here's exactly why running it matters." We already talk about wanting this culture. What almost nobody has is it deployed as a standard, rather than a value on a poster.

The failure mode isn't subtle. It looks like the Omni-Tool Trap I wrote about: five or six half-thought "innovations" aimed at the same problem, propagating because nobody stress-tested the assumption underneath any of them. It looks like sycophancy dressed up as alignment: everyone nodding because nobody wants to be the one who slows a room that's already decided it's moving fast. The success mode looks almost boring by comparison: a standing, unglamorous habit of interrogating the plan before it ships, built into the process instead of left to whoever feels brave that day.

The Gadfly Discipline

Socrates called himself Athens' gadfly: a fly whose whole job was to sting a large, comfortable, slow-moving horse until it moved. Not a flattering job description, and not a safe one. But I've spent enough of my own career challenging teams of leaders and direct managers to know exactly what that sting feels like from the inside. If any of the leaders I've worked for over the years are reading this, you know who you are, and I hope it helped more than it stung.

The habit itself is not complicated. I can propose three questions, run before you accept any plan as finished. You could even ask an AI to assist your stress test using this framework.

What are the assumptions nobody's said out loud yet? Not the assumption, assumptions, plural. It's deliberately built to invite more than one answer, because it's also the question that scales past you. Run it alone and you'll catch one or two things. Run it in a room with someone whose job is to ask it, and you'll catch what you missed.

What's the edge case this breaks under, including the failure modes that didn't exist before? Not just "what could go wrong" in the abstract, specifically: did the plan delete a fallback that used to absorb the failure quietly, before anyone had to notice?

Would this survive being defended, unaided, to someone actively trying to find the hole in it? If the honest answer is "only if nobody pushes back too hard," it isn't done.

None of this produces a clean slate, and that was never the goal. A perfectly clean slate probably isn't possible anyway. What it gets you instead is familiarity with exactly where the plan is weak, which is worth more than the illusion of a plan with no weaknesses at all.

Here's what running all three actually looks like, because a framework without a workout is just a poster.

This past weekend I was at the movies with my girlfriend and her son, and we hit the arcade beforehand. The prize counter had been replaced with a vending machine: win the ticket game, feed it your tickets, it dispenses the prize. No employee needed. On paper, that's a clean cost cut.

Question one surfaces the assumption fast: the human behind that counter was treated as a fulfillment function, a cost to remove, rather than part of what you were actually paying for.

Question two is where the second, quieter assumption shows up on its own. The machine jammed right after her son selected his prize. Nothing dispensed, no credit returned, and the system needed three to five minutes to recognize its own error and reboot before it would even give the ticket credit back. A line formed. People got visibly annoyed. Here's the assumption nobody caught going in: that a machine would be exactly as reliable as the person it replaced. It wasn't. A human doesn't jam. In the old system, the person behind the counter just walked to the shelf and handed over the prize. Automating the role didn't just risk failure, it deleted the thing that used to catch failure before anyone noticed.

Question three finishes it off. Defend the decision to a skeptic and it collapses in one sentence: you saved the cost of a minimum-wage employee, and broke the actual product. The prize itself is cheap plastic, nobody's pretending otherwise. Those sticky hands and eraser toppers only last about a week tops. What you were actually selling was the moment: the win, the feeling of being a champion for a few seconds, a person smiling back while they hand it over. That wasn't on any spec sheet. It got cut anyway, because nobody asked question one out loud before they built it.

That's the individual layer. The organizational layer is the same three questions with a seat assigned to them: a rotating gadfly role in decision reviews, whose entire job that meeting is to attack the plan before it ships. Not to build consensus. Not to be liked. To sting it. Give that role explicit license and everyone in the room gets permission to say "I don't think this survives contact with reality" without it costing them anything, because the role exists precisely to say that.

There's a version of this that belongs in how we build the tools, not just how we use them. If AI is going to sit this close to how organizations decide things, it needs to inherit some of the gadfly's instinct too: designed to push back on a weak plan, not just execute it faster and more fluently than we could have on our own. Used well, that's not friction for its own sake. It's what lets speed and rigor coexist instead of trading one for the other, exactly the tradeoff this piece has been arguing we don't have to make.

Simple Enough to Keep

Socrates didn't get to see the world catch up to what he was doing, and this discipline is exactly as old as he is. What's changed isn't the idea, it's the speed everything else moves at now. A slower world used to give a room enough time to catch its own mistakes before they compounded too far. That margin is mostly gone, which is exactly why this can't stay a personality trait for whoever happens to feel brave in a given meeting. It has to be simple enough to actually run: three questions, on repeat, before anything gets called finished.

That's not a provocation. It's closer to what most of us already agree needs to happen, just written down plainly enough to actually use. The Gadfly Discipline is three questions you can start running on your own work this week, with nobody's permission, long before it's anyone's official policy.

AI is going to keep making the distance between an idea and a finished thing shorter. That's not the part worth resisting. The part worth protecting is what you do with the time that shortening buys you back: whether you spend it moving faster toward the wrong thing, or slower toward the right question.

Three questions, asked on repeat. That's the whole discipline.


About Justin Lerma: AI educator and thought leader focused on the intersection of technology and human performance. Views are my own.

Disclaimer: The views expressed in this publication are personal opinions and do not represent the positions of any employer or affiliate.

© 2025 Justin Lerma. All rights reserved. Unauthorized reproduction or distribution of this content without express written permission is prohibited.