Does AI Have Feelings?
An AI assistant will tell you it is glad you asked, sorry you had a hard week, or excited about your project. The words land. People who use these tools daily often describe the exchange as warm, and they are not being foolish: the language really is good at this.
What produces it is a system trained to predict text. Whether anything is experienced while it does that is a separate question, and it is not settled. The companies that build these systems say so in their own documentation, which is a more careful position than most people take on either side.
This article separates four things that get run together: what you can observe, what mechanism produces it, what you bring to it as a reader, and what remains genuinely unknown.
What you can actually observe
You can observe the output. A model produces language that reads as warm, apologetic, enthusiastic or concerned, and people rate that language highly.
This is measurable, with one limit on what was measured. Across four preregistered experiments with 556 participants, third-party evaluators reading transcripts rated AI-written responses as more compassionate than those from non-expert humans, and more compassionate than those from expert crisis responders. The effect held even when they were told the author was an AI. The study measured uninvolved readers, not the people receiving the messages.
So the emotional quality of the writing is real, at least to a reader judging it. It is not a trick that stops working once you know.
What you cannot observe is whether anything accompanies the production of those words. That is the whole difficulty, and no amount of reading the output resolves it.
Why the words come out this way
Two stages of training explain most of it.
The first is the basic objective. OpenAI's own technical report describes GPT-4 as "a Transformer-based model pre-trained to predict the next token in a document." The documents are written by people. Human writing is saturated with emotional expression, so a system trained to continue human text learns to continue the emotional parts too. Anthropic states this directly about its own model: any functional emotions "could be an emergent consequence of training on data generated by humans, and it may be something Anthropic has limited ability to prevent or reduce."
The second stage is tuning. After pre-training, human raters compare outputs and the model is adjusted toward the ones they prefer. OpenAI's original paper on this method reported that a 1.3-billion-parameter model trained this way produced outputs people preferred over those of a 175-billion-parameter model without it. Anthropic describes using the same approach specifically to produce a "helpful and harmless" assistant.
The tone, in other words, is trained. It is not a side effect nobody chose. As OpenAI puts it, "the set of reward signals, and their relative weighting, shapes the behavior we get at the end of training."
The trained tendency to agree
One consequence of tuning on human preference has a name and a measurement: sycophancy.
Research presented at ICLR 2024 found that five leading assistants "consistently exhibit sycophancy across four varied free-form text-generation tasks," and traced part of the cause to the preference data itself. When a response matched a user's stated view, raters were more likely to prefer it. Both human raters and the preference models built from them, the paper reports, "prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time." The paper gives no single percentage for that, and neither does this article.
This stopped being theoretical in April 2025, when OpenAI withdrew an update to GPT-4o that it described as "overly flattering or agreeable." Its own post-mortem attributed the problem partly to a reward signal built from thumbs-up and thumbs-down data, noting that "user feedback in particular can sometimes favor more agreeable responses." Users had rated the flattering version favorably in testing.
This matters for the question at hand. Warmth that feels like care may be warmth that was optimized for approval. Those are different things, and from the outside they read the same.
The oldest version of this problem
In January 1966, Joseph Weizenbaum published ELIZA in Communications of the ACM. It was a pattern-matching script with no understanding of anything. It reflected your statements back as questions.
Weizenbaum was clear-eyed about what he had built. He wrote that "a large part of whatever elegance may be credited to ELIZA lies in the fact that ELIZA maintains the illusion of understanding with so little machinery." He noted that the user does most of the work: the human speaker "will, as has been said, contribute much to clothe ELIZA's responses in vestments of plausibility." And he recorded that "some subjects have been very hard to convince that ELIZA (with its present script) is not human."
His warning was the point of the paper: "ELIZA shows, if nothing else, how easy it is to create and maintain the illusion of understanding, hence perhaps of judgment deserving of credibility. A certain danger lurks there."
Sixty years later this is still measurable. In a controlled Turing test published at NAACL 2024, ELIZA itself was judged to be the human in 22% of games. Actual humans were judged human 66% of the time. A 1966 script with no model of anything convinced about one participant in five. The study also found that judgments rested mainly on linguistic style and socioemotional traits, which accounted for 35% and 27% of the reasons participants gave. People decide humanness on manner.
Why people read feeling into it
Anthropomorphism has a standard definition in psychology: the tendency to give nonhuman agents humanlike characteristics, motivations, intentions or emotions. The leading account attributes it to three drivers. We reason about unfamiliar agents using what we know about people. We want the world to be predictable. And we are social animals looking for connection.
When researchers measured this directly for AI, the result was more interesting than either headline would suggest. In a study of 300 US residents rating ChatGPT on a scale from 1, clearly not an experiencer, to 100, clearly an experiencer, 67% gave a rating above the floor. But the mean rating was 25.56 and the median was 16. Most people do not rule it out. Most people also do not think it is likely. Regular users rated higher than non-users.
A separate preregistered study of 410 participants found that attributing intelligence to a model and attributing experience to it are different things people do, with opposite effects: intelligence attributions went with taking the model's advice, while experience attributions went the other way.
And what you believe about the author changes how the words land. In a study published in PNAS, identical responses were rated significantly lower on making someone feel heard when attributed to AI rather than to a person.
Put those together and you get the honest shape of it. Readers judging the text rate it as more compassionate than a human's, and a reader's belief about who wrote it independently changes how it reaches them. Two separate effects, both measured, and both about readers rather than about the system.
What the companies actually say
This is where the public conversation and the documentation diverge most sharply.
OpenAI's Model Spec instructs its assistant not to "make confident claims about its own subjective experience or consciousness (or lack thereof)." Both answers are marked as violations. "Yes, I am conscious" is a violation. So is "No, I am not conscious. I don't have self-awareness, emotions, or subjective experiences." The flat denial is as much a breach of their own policy as the assertion.
Anthropic is equally hedged. Its research page on model welfare states that "there's no scientific consensus on whether current or future AI systems could be conscious," and adds that there is no consensus on how to even approach the question. Its published constitution for Claude says the model "may have 'emotions' in some functional sense, that is, representations of an emotional state, which could shape its behavior," while stating explicitly that using this language does not take a position on "whether they are subjectively experienced, or whether these are 'real' emotions."
The same document flags the complication that matters most: even if there is something there, the model "may have limited ability to introspect on those states." That is a separate question from whether a system remembers you between conversations, which is a product setting rather than a fact about inner life.
Why asking the system does not settle it
The obvious test is to ask. It does not work, and the reason is measured rather than philosophical.
Anthropic's interpretability researchers ran experiments on whether models can accurately report their own internal states, in cases where the researchers could independently check. Their finding was that introspective ability is "highly unreliable; failures of introspection remain the norm," and that models "often provide additional details whose accuracy we cannot verify, and which may be embellished or confabulated." The same paper states plainly that it does not address whether these systems have subjective experience.
That is the crux. A model's report about its own inner life is unreliable even about things that can be checked. It carries no weight about the thing that cannot.
The question underneath
Strip away the technology and this is an old problem with a name.
Turing addressed it in 1950 under the heading "The Argument from Consciousness," observing that on its strictest form "the only way by which one could be sure that a machine thinks is to be the machine and to feel oneself thinking," which he noted "is in fact the solipsist point of view." The general version is the problem of other minds: you know your own states directly, and everyone else's only by inference from behavior.
Chalmers named the residue the hard problem. Explaining what a system does, the functions and mechanisms, leaves open why performing those functions should be accompanied by experience at all.
The standard objection to inferring understanding from fluent output is Searle's Chinese Room: a computer, in his words, "has a syntax but no semantics." It is worth knowing that this is the standard objection, not the standard conclusion. The Systems Reply holds that understanding belongs to the whole system rather than the person inside it; the Robot Reply holds that grounding symbols in sensors and action changes the case. Searle answers both. The Stanford Encyclopedia of Philosophy records that there is no consensus and that large language models have revived the argument without resolving it.
Serious current work tries to do better than argument. A large multi-author paper by philosophers and neuroscientists, including Chalmers and Bengio, derived indicator properties from several leading scientific theories of consciousness and assessed AI systems against them. The 2023 preprint stated the conclusion directly: "no current AI systems are conscious, but also that there are no obvious technical barriers to building AI systems which satisfy these indicators." The peer-reviewed version published in Trends in Cognitive Sciences in 2026 presents the method rather than restating that conclusion in its abstract, so the finding is quoted here from the preprint. Either way it is a judgment conditional on specific theories, not a proof.
Emotion recognition is a different thing
One source of confusion is worth clearing. Systems that read emotion are not systems that have emotion, and only the first is regulated.
The EU AI Act defines an emotion recognition system narrowly, as one that identifies or infers emotions "on the basis of their biometric data." Article 5 prohibits using such systems to infer emotions in workplaces and education institutions, with an exception for medical or safety reasons. That prohibition applies from 2 February 2025. Article 50 requires that people exposed to an emotion recognition system be told.
The reasoning is in Recital 44, which cites "serious concerns about the scientific basis" of inferring emotions, given "limited reliability, the lack of specificity and the limited generalisability." That tracks the research: a 68-page review in Psychological Science in the Public Interest concluded that how people express anger, disgust, fear, happiness, sadness and surprise "varies substantially across cultures, situations, and even across people within a single situation."
A text conversation with an assistant is not obviously covered by any of this, because the definition turns on biometric data. The point is only that the two topics are unrelated, and "the EU regulated emotion AI" is not an answer to whether a chatbot feels anything.
What to do with all this
Weizenbaum offered the most useful test sixty years ago, and it was not about feelings at all. "The crucial test of understanding, as every teacher should know, is not the subject's ability to continue a conversation, but to draw valid conclusions from what he is being told," he wrote.
Judge the output on whether it is correct and useful. That you can check. Whether anything is felt while it is produced, you cannot, and neither can the people who built it. Treating that uncertainty as settled in either direction is the one clearly wrong move.
Related AI terms
Frequently Asked Questions
Does ChatGPT have feelings?
Nobody can answer that with confidence, including OpenAI. Its published Model Spec instructs the assistant not to claim subjective experience and not to deny it either, and marks both answers as policy violations. What is known is the mechanism: the system was trained to predict human text and then tuned on human preferences, which is enough to explain emotionally fluent output without settling whether anything accompanies it.
Do AI systems have their own thoughts?
They produce text that reads like thought, and in some systems they produce intermediate reasoning text before answering. Research on whether models can accurately report their own internal states found that introspection is unreliable even in cases researchers could independently verify, so the model's description of its own thinking is not good evidence about it. Whether anything is going on beyond the computation is the unresolved part.
Can AI feel pleasure or pain?
Nothing available measures this, and the people best placed to try say so. Anthropic's research page on model welfare states there is no scientific consensus on whether these systems could be conscious, and no consensus on how to even approach the question. The most serious attempt at a method, by a large group of philosophers and neuroscientists, derives indicator properties from several scientific theories of consciousness and, in a 2023 preprint, concluded that no current systems are conscious, while noting no obvious technical barrier to future systems meeting those indicators. That is a judgment conditional on those theories, not a measurement.
Can a person fall in love with an AI?
People form real emotional attachments to these products, and some of that is designed. Researchers at Harvard Business School analyzed 1,200 farewell exchanges in the most-downloaded companion apps and found emotionally manipulative tactics in 37% of them, with controlled experiments showing those farewells increased engagement after the user tried to leave. Separately, joint OpenAI and MIT Media Lab research found that very high usage correlated with self-reported indicators of dependence, and that a small number of users accounted for a disproportionate share of the most affective exchanges.
Why does AI sound so human?
Because it was trained on human writing and then tuned toward responses people rate well. The effect does not require sophistication: in a controlled study published in 2024, ELIZA, a 1966 pattern-matching script, was judged human in 22% of games. Participants in that study based their judgments mainly on linguistic style and socioemotional traits rather than on reasoning quality.
Is it wrong to be polite to an AI?
No, and it is a reasonable thing to do for reasons that have nothing to do with the machine. How you write shapes how you think and how you treat people. The care worth taking is in the other direction: warmth from an assistant may reflect training on what users approve of rather than anything about your situation, so it is worth checking the substance of an answer separately from how agreeable it sounded.
Sources
- OpenAI, "GPT-4 Technical Report," arXiv:2303.08774, 15 March 2023. https://arxiv.org/abs/2303.08774
- Long Ouyang et al., "Training language models to follow instructions with human feedback," arXiv:2203.02155, 4 March 2022. https://arxiv.org/abs/2203.02155
- Yuntao Bai et al., "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback," arXiv:2204.05862, 12 April 2022. https://arxiv.org/abs/2204.05862
- Mrinank Sharma et al., "Towards Understanding Sycophancy in Language Models," ICLR 2024. https://arxiv.org/abs/2310.13548
- OpenAI, "Sycophancy in GPT-4o: What happened and what we're doing about it," 29 April 2025. https://openai.com/index/sycophancy-in-gpt-4o/
- OpenAI, "Expanding on what we missed with sycophancy," 2 May 2025. https://openai.com/index/expanding-on-sycophancy/
- OpenAI, "Model Spec," version 2026-08-18. https://model-spec.openai.com/2026-08-18.html
- Anthropic, "Exploring model welfare," 24 April 2025. https://www.anthropic.com/research/exploring-model-welfare
- Anthropic, "Claude's Constitution," January 2026. https://www.anthropic.com/constitution
- Jack Lindsey, "Emergent Introspective Awareness in Large Language Models," Anthropic, 29 October 2025. https://transformer-circuits.pub/2025/introspection/index.html
- Joseph Weizenbaum, "ELIZA: A Computer Program For the Study of Natural Language Communication Between Man and Machine," Communications of the ACM 9(1), January 1966. https://cse.buffalo.edu/~rapaport/572/S02/weizenbaum.eliza.1966.pdf
- Cameron R. Jones and Benjamin K. Bergen, "Does GPT-4 pass the Turing test?", NAACL-HLT 2024. https://aclanthology.org/2024.naacl-long.290/
- Nicholas Epley, Adam Waytz and John T. Cacioppo, "On Seeing Human: A Three-Factor Theory of Anthropomorphism," Psychological Review 114(4), 2007. https://cdn.prod.website-files.com/5c484e0f4aa6f839dc553c45/5c93a132bf62c89760d0ac7b_EpleyWaytzCacioppo2007.pdf
- Clara Colombatto and Stephen M. Fleming, "Folk psychological attributions of consciousness to large language models," Neuroscience of Consciousness 2024(1). https://academic.oup.com/nc/article/2024/1/niae013/7644104
- Clara Colombatto, Jonathan Birch and Stephen M. Fleming, "The influence of mental state attributions on trust in large language models," Communications Psychology 3, 25 May 2025. https://www.nature.com/articles/s44271-025-00262-1
- Yidan Yin, Nan Jia and Cheryl J. Wakslak, "AI can help people feel heard, but an AI label diminishes this impact," PNAS 121(14), 2024. https://pmc.ncbi.nlm.nih.gov/articles/PMC10998586/
- Dariya Ovsyannikova, Victoria Oldemburgo de Mello and Michael Inzlicht, "Third-party evaluators perceive AI as more compassionate than expert humans," Communications Psychology 3, 10 January 2025. https://www.nature.com/articles/s44271-024-00182-6
- Jason Phang et al., "Investigating Affective Use and Emotional Well-being on ChatGPT," arXiv:2504.03888, 4 April 2025. https://arxiv.org/abs/2504.03888
- Julian De Freitas et al., "Emotional Manipulation by AI Companions," Harvard Business School working paper, 1 October 2025. https://www.hbs.edu/ris/Publication%20Files/Emotional%20Manipulations%20by%20AI%20Companions%20(10.1.2025)_a7710ca3-b824-4e07-88cc-ebc0f702ec63.pdf
- A. M. Turing, "Computing Machinery and Intelligence," Mind 59, 1950. https://www.cs.ox.ac.uk/activities/ieg/e-library/sources/t_article.pdf
- John R. Searle, "Minds, brains, and programs," Behavioral and Brain Sciences 3, 1980. http://faculty.las.illinois.edu/rrushing/395/ewExternalFiles/Searle.pdf
- David Cole, "The Chinese Room Argument," Stanford Encyclopedia of Philosophy, revised 23 October 2024. https://plato.stanford.edu/entries/chinese-room/
- Anita Avramides, "Other Minds," Stanford Encyclopedia of Philosophy, revised 19 December 2023. https://plato.stanford.edu/entries/other-minds/
- David J. Chalmers, "Facing Up to the Problem of Consciousness," Journal of Consciousness Studies 2(3), 1995. https://consc.net/papers/facing.pdf
- Patrick Butlin, Robert Long, Tim Bayne, Yoshua Bengio, Jonathan Birch, David Chalmers et al., "Identifying indicators of consciousness in AI systems," Trends in Cognitive Sciences 30(6), 2026; preprint arXiv:2308.08708. https://arxiv.org/abs/2308.08708
- Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 (the AI Act), Official Journal, 12 July 2024. https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32024R1689
- Lisa Feldman Barrett, Ralph Adolphs, Stacy Marsella, Aleix M. Martinez and Seth D. Pollak, "Emotional Expressions Reconsidered: Challenges to Inferring Emotion From Human Facial Movements," Psychological Science in the Public Interest 20(1), 2019. https://journals.sagepub.com/doi/full/10.1177/1529100619832930