Companion piece: this essay has a looser Intuition sibling, Do I Have to Take Gary Marcus Seriously?.
|
The conversation behind this The actual author + AI conversation that produced this issue — who brought what. |
Audio companion: Listen to this essay as a narrated audio episode: The Gary Marcus Audit.
I want Gary Marcus to be wrong.
That is not the same as knowing that he is.
The easy dismissal fails quickly. Marcus is not simply anti-AI. He grants practical usefulness in some contexts while disputing the conversion of usefulness into reliability, understanding, or AGI. He argues for different AI, not no AI, and often for more machinery around current AI rather than for pure abstinence.
His objection is specific enough to survive my irritation: current LLM-centered systems can be genuinely useful while remaining unreliable in ways that matter. Companies and commentators can start acting as if "capable and helpful" means "trustworthy enough for authority, access, and institutional dependence."
The question here is which claims survive after I subtract irritation, tribe, and politics.
The provocation is not imaginary. Marcus has used the headline "Generative AI was a scam", then qualified the literal fraud claim in the body. On X, he called Meta's AI direction a "data-labeling sweatshop" (mirror/context). In a post about Yann LeCun and AI-bubble warnings, he used "sociopathic" and later added that "the only constant is his ego" (direct X lead; mirror/context).
Those examples matter because they explain why this reader sometimes experiences his style as jarring. Depending on your politics and where you stand in the AI debate, they might sound bracing and vindicating, or petty, unfair, and exhausting.
They do not settle the technical claims.
They also do not measure his overall influence.
That is the point.
The more serious hypothesis is not that Marcus is wrong because he is irritating. It is that he may be technically right about important failure modes, yet still fail to change the minds of readers already committed to the technology. Builders, adopters, investors, product leaders, and AI-adjacent readers have often moved from "Should we use this?" to "How far can this go?"
That is a reception hypothesis, not a measured cultural fact.
One part of the current context is adoption.
These tools are plainly useful in many contexts. People and organizations use them to draft, code, summarize, translate, search, plan, explain, automate, and avoid blank pages. The question of use moved faster than the question of trust.
Before adoption became routine, warnings about those failure modes could sound to some readers like: do not use these tools because they are too flawed.
Now that many people and organizations already use them every day, the same warning lands differently.
Not: do not use this.
But: do not confuse this with trust.
The second warning can remain live after the first decision has been made.
This is where the design of the system starts to matter.
Marcus often argues that expecting today's large language models to turn into AGI is a basic mistake. He says current systems are broad but shallow. They sound smooth and confident, yet they remain unreliable at sticking to the truth, applying what they know to new situations, and catching their own errors.
But "LLMs will become AGI" blends together several different claims.
One claim is about improving the language models themselves. Make the model bigger. Train it on better data, including data made by other AI systems. Use reinforcement learning. Let the model spend more compute while answering. The model may get more capable, but the story still centers on the core model: how it is trained, how it answers, and how it is tested.
A second claim is about adding things around the model. Give it search, code execution, calculators, databases, fact-checkers, outside memory, safety layers, and loops where the model decides what to do next. The overall system can now do things a plain chatbot cannot. But much of the added ability may come from the extra parts doing real work, not from the model becoming deeply intelligent on its own. If the system needs accurate math, it may call a calculator. If it needs to run code, it may call a code interpreter. An "agent" is often this kind of setup: a model repeatedly choosing and using tools. It is a deployment pattern, not a new theory of intelligence.
A third claim is about building AI with a genuinely different structure. This is closer to what Marcus usually has in mind. These approaches combine neural networks, which are good at pattern recognition, with more explicit forms of reasoning: rules, logical steps, planning, search, verification, or ordinary computer programs. Some try to build world models, internal representations of how reality changes, so the system can predict consequences and plan accordingly. These are not just bigger, smoother chatbots. They try to make reasoning less dependent on fluent text generation alone.
Once those ideas are separated, the argument shifts.
Marcus's target gets narrower. Major labs do not publicly describe their strategy as bare pretraining scale alone. They combine LLM-centered training and inference improvements with tools, verification, multimodality, robotics, orchestration, and safety work, even as some leaders retain strong confidence in scaling-driven gains.
From public descriptions, OpenAI's reasoning work remains largely LLM-centered. DeepMind's robotics combines foundation models with embodied action. LeCun's world-model program is a clearer architectural departure. Those are outside classifications, not accounts of proprietary internals.
The sharper Marcus claim is not that labs have done nothing beyond scaling. It is that confidence in LLM-centered systems may still outrun demonstrated reliability, grounding, security, and governance.
The distinction matters because labels like "agent," "reasoning," "tools," and "multimodal" can sound more explanatory than they are. An agent can simply be an LLM in a loop. A tool-using system can still be unreliable at deciding when to use the tool. A reasoning model can spend more compute and still produce confident falsehoods. A multimodal model can process images and text without developing a stable understanding of cause, consequence, and truth.
The real question is not whether AI labs have moved past simple 2022-style chatbots. They have.
The real question is whether we can safely work around the problems that still exist. In some areas, a workaround can be good enough if it is stable, easy to check, and limited in scope. The hard part is figuring out which domains actually give us those limits.
Coding is one of the easier cases. An AI can write code, and the normal software environment can compile it, run tests, check it with a linter, and show stack traces when something breaks. Those are fast, concrete signals that something went wrong. They do not guarantee the code is correct or secure, but they catch a lot of mistakes quickly and clearly.
Most fields do not have feedback that clean.
Medicine, law, education, management, security, government work, and personal advice are different. Many of their biggest mistakes show up late, depend heavily on context, involve other people, or stay hidden. The world also keeps changing. A policy shifts. A patient develops new symptoms. A legal precedent changes. Yesterday's facts become stale. A model can still sound fluent and confident while the situation it was describing has already moved.
In those areas, strong benchmark performance is related to real capability, but it is not the same thing as dependable judgment. A model can perform impressively across many tests and still fail in exactly the kind of situation where a person or institution might be tempted to trust it.
This tension shows up in the recent dispute between Anthropic and the U.S. government over Fable 5 and Mythos 5.
Anthropic described those models as especially strong at long-term planning, coding, science, visual tasks, and memory. On June 12, Anthropic said a U.S. export-control directive forced the company to disable access to both models. Anthropic pushed back, saying the government had not given a clear technical reason and that the issue was narrow enough to be handled with ordinary defense-in-depth safety practices. Axios reported a government-side account involving a jailbreak report and subsequent testing. The complete technical record is not public, so outsiders cannot fully judge who is right.
That creates two pressures at once.
On the capability side, the models' public capability claims and the government's response put pressure on blanket claims that frontier systems are trivial or strategically unimportant. Neither validates the strongest claims or reveals how much performance came from the core model rather than tools, memory, specialized environments, or human oversight.
On the governance side, the dispute shows how quickly strong capability claims can run into disagreements about safeguards, access, evidence, and acceptable residual risk.
Anthropic did not argue that the models were perfectly reliable. It argued that no current AI is fully immune to jailbreaks, that multiple layers of protection are the realistic standard, and that the specific behavior identified was not unique to its models. Its position was essentially: this risk exists across many systems, our safeguards are relatively strong, and the government's response was too broad.
That is the core tension.
A system can be treated as consequential enough to raise national-security concern while still not being secure, governed, or dependable enough for high-trust institutional use. In cybersecurity, even an inconsistent system can be dangerous because attackers can keep trying, ignore failures, and chain partial successes. In medicine, law, finance, government, or education, the tolerance for hidden error is much lower.
So the reliability concern has not gone away. It has changed.
A blanket "the models are too dumb to matter" dismissal now looks less adequate. The sharper concern is that a system can create real leverage while remaining too unreliable or insecure for the authority people want to give it.
That is a stronger version of Marcus than the one I wanted to dismiss.
It is also harder for Marcus's own position to carry.
Marcus has supplied meaningful AGI criteria: flexibility, generality, resourcefulness, depth, reliability, and the ability to generalize and check answers. If "LLMs are not the path to AGI" means "scaling alone will not solve reliability, grounding, and generalization," the claim is serious and increasingly specific.
If it means "no system built around current LLM-style approaches can ever reach transformative general capability," a hard low ceiling, it needs clearer update conditions.
By update conditions, I mean something simple: tell me what the world would have to look like in two to five years for you to change your mind. If you cannot or will not specify that, your position becomes less useful as a guide to reality.
What sustained real-world performance would count against that ceiling? What kind of reliability outside benchmarks would matter? When does a successful system that combines models, tools, verification, and human oversight count as support for the current path, and when does it count as the hybrid approach critics expected all along?
Those are not trick questions. They are the questions a useful critique has to keep answering as the technology changes.
So where does the audit land?
Not with endorsement.
I still do not want Marcus as my oracle for AI. His rhetoric often makes this reader do avoidable cleanup.
Part of that is the "confirms what he already knew" problem. New failures can arrive in Marcus's writing less like fresh evidence to be examined and more like vindication of a story already in place. The tone can carry an implicit "I told you so." The analysis can feel like scorekeeping rather than inquiry. That does not make the underlying critique wrong, but it gives opponents an easy way to dismiss it as preloaded.
There is another complication. Provocation may also be part of the strategy. On X, sharper language may travel farther than careful, constrained technical criticism. It can engage supporters, irritate opponents, and turn even people who want Marcus to be wrong into distributors of his claims. What lands to me as a persuasion defect may also function as a distribution engine.
That leaves an awkward split verdict.
Some of Marcus's claims are substantive. Current AI can be useful and still unreliable; agents make those failures more consequential; serious AI risk does not require AGI; and Marcus has a longstanding hybrid-AI alternative behind his criticism.
But some of his position remains on shakier ground.
His public criteria do not yet give me a clear rule for how he would update if future mixed systems succeed. That leaves room, fairly or unfairly, for the goalposts to keep moving. His rhetorical style can make new evidence sound like confirmation of what he already knew. And his ability to change the minds of readers already committed to the technology remains an unmeasured reception question.
My own conclusion is narrower than endorsement.
Usefulness has been overconverted into reliability in some public and product rhetoric. In other words: people are acting as if "this AI is capable and helpful" means "this AI is trustworthy enough to receive real authority and access." That conversion is not justified by usefulness alone.
Agents make reliability failures matter more because the system can act. Serious AI risk does not require superintelligence. Powerful but unreliable AI with access can create institutional harm before anything like settled AGI arrives.
Marcus is on shakier ground when the claim becomes that hard low ceiling. The public capability claims and the dispute make it harder to state casually, even though they do not validate the claims, reveal the architecture, or settle the question.
But that does not collapse the governance critique. It clarifies it.
A warning can be live even if its messenger is jarring. A technology can be useful, powerful, and institutionally consequential before its reliability and governance are mature. A culture can accept a tool before it understands the risk it is accepting.
The difficulty is holding the full tension without letting one part collapse into the other. Preventing that collapse requires actively stress-testing our own conclusions. Self-critique sounds simple:
Write the sentence you want to be true, then name what would embarrass your view if it were real. In practice, motivated reasoning feels like discernment while it is happening.
An AI-assisted adversarial critique can help because it can be made to work against preference. Its objections are not evidence; the model can invent, flatten, and misread, and its output still has to be checked against the record. The value is narrower: it can expose a missing distinction or convenient certainty and make avoidance harder.
I began with a wish: I want Gary Marcus to be wrong.
The audit gives me something less convenient.
Marcus may fail to change the minds of readers already committed to the technology.
His claims may be weakest where his public criteria do not clearly say what would count against him.
His rhetoric can be hard to distinguish from confirming what he already knew.
And he may still be right about enough that dismissing him says more about my filters than his claims.
There is an inescapable irony here. While part of me would love to watch his ideas fade into irrelevance, the very act of writing this piece helps spread them. If provocation is Marcus's strategy, it appears more effective than I would like to admit.
What this is: Field Notes testing an irritating but potentially useful AI critic against primary-source claims, public architecture descriptions, and a dated governance dispute. It is not a Gary Marcus profile, a proprietary architecture analysis, policy advice, investment advice, or proof that an editorial process makes the conclusions true.
Review date: June 20, 2026. The Fable/Mythos access dispute and company descriptions are time-sensitive.
Confidence: Medium on the Marcus claim map; medium on the public-source architecture taxonomy; low on proprietary internals and the government's complete technical rationale; medium-low on Marcus's ability to change minds among adoption-entangled readers, which remains an unmeasured reception hypothesis.
What would change this view over the next 12-24 months: sustained, inspectable evidence that LLM-centered systems operate reliably in changing high-trust contexts; clearer evidence about how Marcus would update his low-ceiling claim when mixed systems succeed; or measured evidence that his warnings materially alter, or fail to alter, deployment and governance decisions.
Process transparency: AI tools assisted with drafting and adversarial critique. The human author selected the frame, checked the sources, made the judgments, and owns the published claims and errors. Process review is not evidence that the claims are true.
Sources and anchors
Marcus's positions
Gary Marcus, "AGI versus 'broad, shallow intelligence'": AGI criteria, broad/shallow capability, reliability, generalization.
Gary Marcus, "Why DO large language models hallucinate?": hallucination, factuality, statistical learning, world-model criticism.
Gary Marcus, "AI Agents have, so far, mostly been a dud": agent reliability concerns.
Gary Marcus, "AI risk != AGI risk": near-term AI risk distinguished from AGI and superintelligence.
Marcus, Davis et al., "The Next Decade in AI": hybrid, knowledge- and reasoning-based alternatives.
Gary Marcus, "Turns out Generative AI was a scam": useful-context concession plus overhype/reliability critique.
Current public architecture framing
OpenAI, "Learning to reason with LLMs": reinforcement learning and train/test-time compute as LLM-centered reasoning work.
OpenAI, "Introducing GPT-5.5": current tool use, agentic execution, multimodality, and test-time-computation framing.
Anthropic, "Measuring agent autonomy": agents as tool-using loops and autonomy as a deployment property.
Google DeepMind, "Gemini Robotics 1.5 brings AI agents into the physical world": embodied vision-language and vision-language-action systems built on a Gemini core.
Meta, "V-JEPA 2" and Reuters on LeCun / AMI: the distinct world-model bet and its current institutional home.
Meta, "Muse Spark": Meta's current LLM-scaling and subagent framing.
Dario Amodei, "Policy on the AI Exponential": continuing strong scaling-law confidence.
Fable/Mythos
Anthropic, "Claude Fable 5 and Claude Mythos 5": Anthropic's own capability and safeguard claims.
Anthropic, "Statement on the US government directive to suspend access to Fable 5 and Mythos 5": Anthropic's account of the directive and defense-in-depth argument.
Rhetoric examples
Marcus's "Generative AI was a scam" headline/body contrast.
Meta "data-labeling sweatshop" direct X lead (
https://x.com/i/status/2067375496346894404) plus mirror/context.LeCun "sociopathic" / "ego" direct X lead (
https://x.com/i/status/2067592852587401567) plus mirror/context.
Source limitations
Public material does not reveal complete proprietary architecture. The taxonomy is an editorial synthesis of disclosed system features.
Anthropic and government-side accounts do not provide a complete public technical record.
X is unstable as a reference surface. Mirrors are included for inspectability, not as stronger evidence than original posts.
Marcus's ability to change minds among adoption-entangled readers remains a reception hypothesis, not an empirical measurement.
