Some Thoughts on AI and Thinking in a Post-Parrot World

What the argument over anthropomorphism reveals about language and shared understanding.

Shared understanding, useful shortcuts, and what the argument over AI anthropomorphism gets wrong.
AI
teaching
opinion
Author

Dan O’Leary

Published

September 27, 2026

I’ve been using, teaching, and thinking about AI in various forms for about ten years. During that time, the capabilities have changed… considerably. As have my expectations of what these systems can do. But the language used to describe them has had a harder time keeping up.

This weekend, my X (aka the Twitter) feed filled with an argument sparked by the Associated Press’ Stylebook account. It began with a firm declaration and advised journalists to avoid attributing human characteristics to AI.

It sparked broad, sometimes vigorous, response. Based on a quick survey of 472 posts, by Sunday afternoon the discussion had accumulated nearly five million views, 45k likes, and thousands of reposts, replies, and bookmarks. The debate attracted a diverse audience of programmers, researchers, journalists, philosophers, and plenty of interested spectators.

I was pleased to recognize so much of it. In Thriving with AI, a cross-discipline undergraduate honors course I’ve offered since 2024, students are expected to consider what these systems can do, how they work, and what is meant when we describe them as understanding, reasoning, or thinking. Watching the same questions surface in a much wider conversation was welcome validation of that work.

The AP’s recommendation to describe behavior and consequences is sensible and responsible. Whatever result an AI system produces (helpful or harmful) any report of it must describe what happened and who made the relevant decisions. That is table stakes of credible journalism. It was the warning against anthropomorphization (“ascribe human traits, emotions or behaviors to non-human things”) that drew attention from the likes of Paul Graham, Nate Silver, and others.

With its limit of 120 240 words per post, the platform formally known as Twitter is not the place for subtle distinctions or nuanced debate impossible. That doesn’t keep people from trying…

Paul Graham objected that cognitive language is already established usage. Programmers have long described software as remembering settings, deciding which process to run, or knowing where a file lives. Most of us can use those expressions without imagining a person inside the machine. Graham then went further, suggesting that calling what LLMs do “thinking” will begin as a convenience and eventually become something we acknowledge as true.

I suspect that most readers will find the first of those arguments easy to accept. The second demands that we carefully consider what it means to think (also learn, understand, etc.). And that can be a slippery slope.

That distinction kept getting lost as the discussion spread. Nate Silver pointed to coding agents that can examine a large program, debug it, and make revisions. Others emphasized that successful performance does not establish humanlike cognition or experience. Researchers debated whether “stochastic parrot” remains an informative description, and whether claims about an earlier language model apply to a modern system with additional training, tools, and feedback.

There are substantive disagreements here. But participants also keep answering different questions.

In class, we distinguish mechanistic, functional, and intentional accounts of a system. A mechanistic account describes how it works. A functional account describes what it accomplishes. An intentional account interprets its behavior in terms of beliefs and goals. Saying that a model generates tokens describes part of its operation. It doesn’t, by itself, tell us how well the model can diagnose a programming error. Successfully diagnosing that error doesn’t establish that it experienced the satisfaction of figuring something out.

Two stick figures watch an airplane take off. One says, ‘It’s not really flying. It doesn’t have feathers.’ The other has a thought bubble containing a parrot.

The critics are right to warn about how easily we slide between these claims. Xaq has emphasized the dangers of unchecked anthropomorphism in our course, and I take those concerns seriously. A conversational system supplies cues we ordinarily associate with another person: fluent language, apparent attentiveness, encouragement, an apology when something goes wrong. We supply some of the meaning ourselves. As Big Joel argued, what one person regards as convenient shorthand may convey a claim about experience to someone else.

But I think some proposed remedies introduce more friction than understanding. My wife says I put the Dan in peDANtic. Even I recognize that clear communication depends on shortcuts. We cannot unpack the full meaning of every word each time we use it, and attempting to do so would make many explanations impossible to follow.

Much of what we communicate depends on compression. A concept we treat as simple can rest on a tower of human knowledge and experience. We use a few words because we expect them to call up enough of that background for someone else to follow. Sometimes that expectation fails. We discover that we mean different things, explain further, and try again.

A new technology can make use of familiar shortcuts too. Its novelty means we need to be more deliberate about establishing what those shortcuts mean and what they leave unresolved. It doesn’t make the shortcuts inherently illegitimate.

The intentional stance gives us a way to think about this. We can knowingly choose a description because it helps us understand and predict behavior. Saying that a chess program is trying to protect its queen needn’t confuse anyone about what the program is. The description has a purpose and a scope. Learning to use it well includes recognizing when we need a different explanation. Precision includes knowing what a description is useful for and where it stops working. It doesn’t require replacing every familiar expression with a description of the underlying computation.

Shared understanding is what we are trying to build through communication. Each exchange starts with whatever common ground we have and, ideally, extends it. That takes work from everyone involved. Speakers need to consider what their words imply to their audience. Listeners need to consider what a speaker could reasonably mean, ask when it is unclear, and remain willing to revise their interpretation.

None of this requires agreement. I can understand your position and your reasons for holding it while continuing to disagree. Shared understanding makes debate possible, and debate can reveal differences hidden beneath an apparent agreement. We need both.

This is where I think AP’s guidance takes the wrong shortcut. A vocabulary rule can substitute for the understanding it is supposed to encourage. Knowing to say “output” instead of “answer” tells me very little about whether someone understands the system. A writer can follow every terminology rule and still produce a misleading explanation. Someone using “thinking” can communicate clearly with readers who understand the intended scope.

Journalists do have a broad audience, and they cannot assume everyone shares a programmer’s conventions. That is a reason to explain carefully. Where a familiar term carries an implication the evidence cannot support, qualify it or choose another. But the replacement should improve the reader’s understanding. Sounding less human is not enough.

I want our students to develop the judgment to make those choices deliberately. That means examining an AI’s response, their own interpretation of it, a company’s claims, a skeptic’s dismissal, and a stylebook’s confident declaration. All deserve questions about meaning, evidence, and consequences.

That habit extends well beyond AI. We will always need language that compresses more than it spells out, and we will always encounter people who use it differently. The work is to understand what we are saying, help others understand it, and notice when our words have carried us further than our knowledge.

Survey note: Engagement totals were captured on September 27, 2026, at approximately 1:15–1:19 p.m. Central time. The 472 unique posts came from selected discussion anchors and direct quote posts of AP’s guidance; this was a partial survey, not a census. Views are summed post impressions, not unique readers.