What can the world’s worst bank robbers teach us about AI’s impact on organisational decision-making?
When Life Gives You Lemons

It started with a bank robbery. Two bank robberies to be precise. The first day of a glorious and unstoppable crime spree by two master criminals who’d cracked the code and made themselves invincible. First Pittsburgh, then the world!
You can imagine then, the shock when the “mastermind” of the enterprise, Clifton Earl Johnson, was arrested within a week. His accomplice, McArthur Wheeler, managed to evade justice for three months before he too was apprehended when police broadcast a surveillance photo.
When shown photos of himself from the banks’ security cameras, Wheeler was shocked, exclaiming, “But I wore the juice, I wore the juice!”

Upon further questioning, it emerged that the hapless pair had soaked their faces in lemon juice. Johnson had read somewhere that lemon juice is an ingredient in invisible ink and figured it could make anything invisible, including a face. What could possibly go wrong?
Wheeler had been sceptical. He tested the plan by covering himself in juice before taking a selfie with an old Polaroid camera he’d found in a drawer. The resulting picture seemed to prove Johnson’s theory: Wheeler wasn’t visible.
The story might have remained an obscure footnote in the annals of true crime, had it not come to the attention of Cornell Psychology Professor, David Dunning. He thought the story illustrated an interesting phenomenon:
“If Wheeler was too stupid to be a bank robber, perhaps he was also too stupid to know that he was too stupid to be a bank robber…”
David Dunning & Justin Kruger
Dunning and a graduate student, Justin Kruger, set out to compare individuals’ self-perception of competence with their actual competence. The resulting paper, Unskilled and Unaware of It: How Difficulties in Recognizing One’s Own Incompetence Lead to Inflated Self-Assessments, has, belatedly, become one of the best-known in Social Psychology.
Across four separate experiments, the pair investigated the relationship between capability and confidence. They consistently found that subjects scoring in the bottom quartile on tests of humour, grammar, and logic overestimated their performance.
The paper’s key finding has become known as the Dunning-Kruger Effect:
“…when people are incompetent in the strategies they adopt to achieve success and satisfaction, they suffer a dual burden: not only do they reach erroneous conclusions and make unfortunate choices, but their incompetence robs them of the ability to realise it. Instead, like Mr. Wheeler, they are left with the mistaken impression that they are doing just fine.”
David Dunning & Justin Kruger
Despite its status within academia, however, the concept was relatively unknown among the general public. That is, until 2016, when for reasons that no-one could possibly fathom, it suddenly shot to the attention of the world’s media.

Like many concepts which migrate from academia to the popular consciousness, the Dunning-Kruger Effect has been widely misunderstood. The popular interpretation provides a comforting illusion, exemplified by the image above, that someone else is the idiot, that stupid people are generally too stupid to know how stupid they are.
The real effect, however, is the much more nuanced claim that in domains where our level of expertise is low, we are all prone to overestimating our knowledge and ability, seemingly because we simply don’t know enough to know how incompetent we are in that particular field. Conversely, as our expertise in a domain increases, we become better able to accurately assess our capability.
The precise size and causes of the Dunning-Kruger Effect remain disputed, with some researchers arguing that the famous pattern is little more than a statistical artefact. But the underlying problem is real enough: we are often poorly calibrated judges of our own performance, especially when we lack the knowledge needed to recognise a good answer.
Confidence Tricks
The Dunning-Kruger Effect shines a light on the uneasy relationship between confidence and competence. It’s an uneasy relationship chiefly because humans tend to use confidence as a proxy for competence, particularly in organisational settings.
The core thesis of Behavioural Economics is that much human decision-making occurs automatically, with little or no conscious awareness. On this account, many of our decisions are made using mental shortcuts or rules-of-thumb. We rely on these heuristics because making decisions in complex, dynamic environments is difficult.
Rather than figuring out a hard problem, like “what should I do in this novel scenario?”, our brains often answer an easier question instead, like “what do I normally do?”, “what’s everyone else doing?”, or “what’s the easiest thing to do?”
One of the shortcuts we tend to use is known as the “confidence heuristic”. In domains we don’t know much about, figuring out if someone else knows what they’re talking about is nigh-on impossible. Much easier to ask: “do they speak eloquently?”, or “do they seem confident?”.
It’s not hard to see how this can go wrong. Life is full of people whose confidence far outweighs their competence; the sort of people who’ll say absolutely anything to get ahead, irrespective of whether it’s true.

It’s a tendency observed by one of history’s greatest thinkers more than a century before the hapless Pittsburgh bank robbers embarked on their ill-fated crime spree.
“Ignorance more frequently begets confidence than does knowledge”
Charles Darwin
On Bullshit
When I studied Philosophy at the LSE a quarter of a century ago, there was little doubt as to students’ favourite text. Not Plato’s Republic, the foundation for the whole of western thought. Not student radical bible, Marx and Engels’ Communist Manifesto. Not Karl Popper’s The Open Society and Its Enemies, the foundational text of our entire department.
No, the best-loved Philosopher among undergraduates was a man you’ve never heard of: Harry Frankfurt. His most famous work has a title perfectly devised to appeal to undergrad humour, On Bullshit.
Frankfurt’s beautifully transgressive thesis is that bullshit is crucially different from lying, because the liar deliberately seeks to hide the truth, whereas the bullshitter has no concern whatsoever for it, preferring instead to say something, anything, in order to appear credible. Frankfurt convincingly argues that bullshitters are considerably worse than liars, whose attempts to hide the truth mean they hold it in high regard. He also makes clear when and why bullshit occurs:
“Bullshit is unavoidable whenever circumstance requires someone to talk without knowing what he is talking about.”
Harry Frankfurt
Frankfurt’s 5-minute summation of the idea is well worth a watch.
What, you might be wondering, does this have to do with AI?
Here’s Princeton Professor of Computer Science, Arvind Narayanan and his graduate student Sayash Kapoor, co-authors of a book called AI Snake Oil and an excellent Substack titled AI As Normal Technology, and named as two of TIME Magazine’s 100 Most Influential People in AI.
“OpenAI’s new chatbot ChatGPT is the greatest bullshitter ever. Large Language Models (LLMs) are trained to produce plausible text, not true statements.”
Arvind Narayanan & Sayash Kapoor
Contrary to what this statement might suggest, the two are not anti-AI, merely sceptical of much of the wilder hype accompanying the rise of Large Language Models, and inclined to see them as an incremental technological improvement rather than a species-changing revolution. (A description which, in case you haven’t noticed yet, very much describes my own position.)
Here’s another version of the same sentiment, from a paper with a title also presumably destined to make it an undergrad favourite: ChatGPT is bullshit:
“Because these programs cannot themselves be concerned with truth, and because they are designed to produce text that looks truth-apt without any actual concern for truth, it seems appropriate to call their outputs bullshit.”
Michael Hicks, James Humphries & Joe Slater
The problem all these authors identify is very simple: LLMs aren’t designed to determine if statements are true, but to predict the next chunk of text that will sound plausible and convincing.
Humans tend to treat language as the expression of a mind trying to say something true, whereas LLMs generate language by optimising for plausible continuation. In the purely Frankfurtian sense, Claude and ChatGPT and Gemini are bullshitting.
The term “hallucination” is widely used to describe LLM errors. However, this is more marketing than reality; “hallucination” suggests the LLM is experiencing a delusion, a framing that contributes to the anthropomorphism that’s prevalent among AI users, and is almost certainly wrong.
More importantly though, “hallucination” suggests anomalous behaviour, implying that the hallucination is a one-off error, a bug in the code. Nothing could be further from the truth. Hallucination is a feature of LLMs, not a bug. All LLM outputs are generated in exactly the same way, strings of syllables combined probabilistically, based on the frequency of their occurrence in the training data. The same process produces both accurate outputs and inaccurate outputs.
What’s notable about LLMs is not that their probabilistically generated text strings are sometimes nonsense, it’s that most of the time, they’re not. The difficulty arises because the user must distinguish between accurate and inaccurate outputs, but the users most likely to seek AI assistance on a topic are the ones least equipped to make that judgement.
Dunning-Kruger-as-a-Service

The problem of course, is that confidence heuristic we saw earlier. Because the LLM output is coherent, eloquent, and well-structured, even a world-leading expert in a particular domain can sometimes fail to spot the errors. Perversely, what makes this worse is the infrequency of those errors; because most of the content is accurate, it’s easy for users to be lulled into a false sense of security. The blend of fact and fabrication is incredibly potent.
It’s eerily reminiscent of a line from William Peter Blatty’s ’70s horror classic The Exorcist. As they prepare for the climactic ritual, Max Von Sydow’s experienced exorcist Father Merrin warns his young protégé, Jason Miller’s Father Karras, that:
“The demon is a liar. He will lie to confuse us; but he will also mix lies with the truth to attack us. His attack is psychological, Damien. And powerful.”
Father Merrin (The Exorcist)
Of course LLMs aren’t demonic, no matter what Elon Musk or some weirder corners of the internet might say. To give credit where credit is due, LLMs are incredibly powerful, with a great many use cases, some of which will likely prove transformative. But like any powerful tool, whether a chainsaw or a spreadsheet, that usefulness relies on prudent usage by a careful, expert user.
Logically you would think that more experienced and more expert AI users would be more aware of the risks and better able to mitigate against them, but a recent study from Aalto University in Finland suggests the opposite.
Researchers found that using LLM-based chatbots changes the familiar Dunning-Kruger pattern, in an unexpected way. Their study showed that using AI to help with logical reasoning tasks improved almost every participant’s performance, but not as much as the participants believed.
More strikingly, while all users struggled to accurately rate their own performance when using chatbots, the more experienced they were with AI, the less accurate their assessments became.
In other words: familiarity with AI seems to increase trust in AI while reducing the ability to recognise its limitations.
That’s not the end of the bad news. OpenAI’s latest research, hot off the press in late July 2026, is entitled Work at the Frontier: How AI is expanding what people do at work. The report highlights a growing trend for professional users of ChatGPT to rely on the tool for tasks that traditionally would have been done by other teams or departments.
“AI may change not only how work gets done, but who does it. Our new research suggests that ChatGPT users frequently seek help with tasks historically associated with other occupations, suggesting AI may broaden roles and change how work is divided.”
OpenAI
Is this, as many will infer, a welcome and long overdue breaking down of traditional silos and democratisation of work, with everyone becoming a generalist?
Or is it the Dunning-Kruger Effect on steroids, with marketers vibe coding software, salespeople doing accountancy, and engineers writing their own legal documents, all unaware of what they don’t know?
“With AI, a salesperson who once handed a customer dataset to an analyst can now easily and quickly explore it themselves. A marketer who once waited for a software or web developer can now troubleshoot a website or write a simple script.”
OpenAI
Much executive excitement about LLMs stems from its potential to slash the time required to produce deliverables. Vast numbers of white-collar workers spend huge parts of their working lives writing reports, crafting presentations, documenting decisions, and communicating with each other about that work. LLMs offer the tantalising possibility of automating much of this: outputs that used to take weeks can now be produced in seconds.
This assumption is a crucial underpinning of the near-universal expectation that LLMs will lead to a white-collar job apocalypse. (It’s an assumption I have long argued is wrong, but that’s a different story…)
On the face of it, this narrative makes a lot of sense. After all, much of white collar work does consist of producing documents, spreadsheets and the like. While the term has lost its meaning since computers became ubiquitous, there’s a reason office work used to be disparaged as mere “paper pushing”.
However, if you look closer, all is not as it seems. Outcomes are not the same as outputs. The product of expertise is not the same as having, or acquiring, that expertise through time and experience. The ability to instantly and effortlessly generate a report comparable to what an expert might produce does not confer the user with the same expertise.
The Dunning-Kruger Effect suggests that the absence of expertise makes the user unable to accurately judge the output. Perhaps most crucially, it robs the user of the ability to know what’s missing.
The marketer using Claude Code might be able to make the code work, but will they know enough to notice the inherent security vulnerabilities?
The engineer using ChatGPT might be able to draft a decent-looking legal document, but will they know enough to recognise the crucial missing clause?
The tech bro using Grok to do quantum physics might…wait, I don’t know enough quantum physics to finish that sentence. Though at least I recognise that fact, unlike Uber founder Travis Kalanick:
As the unfortunate Kalanick illustrates, this isn’t breaking down silos, it’s Dunning-Kruger-as-a-service. We’re now at serious risk of every organisation rapidly filling with McArthur Wheelers, blissfully ignorant of their own ignorance, bathing in lemon juice and waving to the security cameras.
Decisions, Decisions
In a now famous podcast interview with Lex Fridman, Jeff Bezos outlined his unusual protocol for meetings at Amazon and rocket company Blue Origin. Rather than participants sitting through a PowerPoint presentation, Bezos insists that the first 30 minutes of a meeting be spent reading and carefully considering a detailed memo, before group discussion commences.
Whatever you think about Bezos, his reasoning for this provides a profound insight about organisational life:
“PowerPoint is really designed to persuade, it’s kind of a sales tool and internally the last thing you want to do is sell. You’re truth-seeking, you’re trying to find truth.”
Jeff Bezos
As I put it in BS At Work, it’s worth pondering for a moment whether Bezos’ statement is true of most organisational meetings. How many of the meetings you attend are really about truth-seeking as opposed to selling to each other?
Truth-seeking is extraordinarily difficult. Many organisational decisions end up shaped by personal ambition and social dynamics. The organisational core decision-making ritual, the meeting, is almost perfectly engineered not for objective truth-seeking, but as a stage for social dynamics, psychodrama, and politics.
The HIPPO – the HIghest Paid Person’s Opinion – is incredibly powerful. Once everyone knows what the boss thinks, their own views are irreversibly altered, often without them realising it. Fear of ostracism from the crowd, a core human trait with deep evolutionary roots, can lead to groupthink.
“Everybody else was doing it” isn’t an acceptable excuse for kids misbehaving in the playground, but for organisations, it’s generally considered a solid basis for decisions, as, with depressing frequency, is: “we’ve always done it this way”.
And how much of organisational life, how many organisational decisions can be explained with what classic British comedy Yes Prime Minister described as “Politician’s logic: Something must be done, this is something, therefore we must do it.”
Certainty and decisiveness are often valued over analysis. Organisations often promote the people most willing to speak with certainty about things nobody can know with certainty.
When we add AI into this already combustible mix, the confidence heuristic is magnified. Well-structured arguments, confidently and eloquently presented, so often win the day, irrespective of validity or veracity. This is precisely what LLMs produce.
As chatbot use proliferates, it won’t just be the creators of content who come to rely on it. LLMs will also increasingly be used to validate and verify the sales pitches we’re hearing in meetings, further diluting the role of true experts by democratising access to expert-seeming content.
Cost, risk, and practicality will drive most organisations to limit staff AI access to particular tools or models. This creates a serious but subtle risk of exacerbating groupthink. The wisdom of crowds relies on diversity of thought and independence of views, but if the same model, trained on the same data set, produces the proposal, the feedback, the executive summary, and all the counterarguments, we’re in an AI-powered echo chamber.
This brings us right back to where we started: a sceptical McArthur Wheeler testing the lemon juice theory of invisibility by taking a selfie with an old Polaroid. The true moral of the story is not that Wheeler believed a stupid theory, but that he tested it, obtained misleading feedback, and became more confident.
It’s a perfect description of how LLMs magnify and embed flawed beliefs in organisational life. Dunning-Kruger-as-a-Service.
Overcoming this will require a very deliberate step-change in how organisations approach decision-making. Organisations have always struggled to distinguish confidence from competence, a confusion industrialised by LLMs. The challenge is to ensure genuine expertise is properly valued, independent thought encouraged, and constructive challenge prioritised, particularly from those willing to contradict the apparently omniscient machines.
Discover more from The Behaviour Boutique
Subscribe to get the latest posts sent to your email.
