THE ARTICLE · 8 MIN
A lot of what circulates about AI chatbots comes from headlines and social posts rather than from the companies’ documentation or the research papers. Below are ten common beliefs, each checked against the original source. Some are wrong, some are half right, and a few are genuinely unsettled. This is a snapshot as of September 2026.
How to read the verdicts: MYTH means the sources contradict the belief. PARTLY TRUE means there is something to it, but the usual version overstates it. UNVERIFIED means no primary source settles it. DISPUTED means credible sources disagree.
1. “AI chatbots learn from your conversation in real time”
Partly true The model itself does not change while you talk to it. Its knowledge comes from training that happened before you started chatting, and that takes a long time: Anthropic wrote in August 2025 that “models released today began development 18 to 24 months ago”.
But many chat apps now have a memory feature that carries details from your chats into later conversations. OpenAI says memory, when enabled, “helps ChatGPT automatically remember useful context from your chats, files, and connected apps to personalize your experience”. Anthropic says Claude “can also remember context from your chats and carry it into new conversations”, and that “Memory is on by default for Free, Pro, and Max plans”. That is saved context, not retraining.
Separately, your chats may be used to train later models, depending on the company and your settings.
- Anthropic (consumer accounts, August 2025): “We will train new models using data from Free, Pro, and Max accounts when this setting is on”. It said the change does not apply to services under its Commercial Terms, including API use.
- OpenAI (policy page updated March 2026): “ChatGPT, for instance, improves by further training on the conversations people have with it, unless you opt out”. For business products it says: “By default, we do not train on any inputs or outputs from our products for business users, including ChatGPT Team, ChatGPT Enterprise, and the API.”
Check the data settings of the app you use.
2. “GPT-4 has 1.8 trillion parameters”
Unverified OpenAI never published GPT-4’s size. Its technical report says it “contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar”. We found no primary source for the 1.8 trillion figure.
3. “OpenAI invented the Transformer”
Myth The Transformer — the design behind today’s large language models — was introduced in the 2017 paper “Attention Is All You Need”. Its authors list affiliations with Google Brain and Google Research, and one with the University of Toronto. OpenAI built on it: its GPT-4 report describes GPT-4 as “a Transformer-based model”.
4. “AI chatbots look things up on the web when they answer”
Partly true It depends on the product. OpenAI’s help centre says: “ChatGPT may search the web automatically when your question would benefit from current information.” In developer APIs, search is a tool that has to be switched on. OpenAI’s API guide says: “To enable this, use the web search tool in the Responses API or, in some cases, Chat Completions.” Anthropic describes its version as a tool that “gives Claude direct access to real-time web content, allowing it to answer questions with up-to-date information beyond its knowledge cutoff.”
When no search happens, the answer comes from what the model learned in training and from what is already in the conversation.
5. “AI is conscious”
Unverified There is no established evidence that current AI systems are conscious, and no agreed test. A 2023 report by a group of researchers assessed AI systems against “indicator properties” drawn from scientific theories of consciousness. It concluded: “Our analysis suggests that no current AI systems are conscious”. It also said “there are no obvious technical barriers to building AI systems which satisfy these indicators”.
That assessment covered systems available in 2023. A 2025 follow-up led by the same first author noted that for tests of consciousness “it remains unclear how they should (or even could) be validated”.
6. “AI detectors can reliably tell if text was written by AI”
Partly true OpenAI released its own AI-text classifier in January 2023 and then withdrew it: “As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy.” In its own evaluation, it “correctly identifies 26% of AI-written text (true positives) as “likely AI-written,” while incorrectly labeling human-written text as AI-written 9% of the time”.
A 2023 study tested “seven widely-used GPT detectors on a corpus of 91 human-authored TOEFL essays” from a Chinese educational forum — essays written by people taking an English test. The detectors “consistently misclassify non-native English writing samples as AI-generated”, with an average false positive rate of 61.22%, and “89 of the 91 TOEFL essays (97.80%) are flagged as AI-generated by at least one detector”.
Newer detectors do much better. Nature reported in August 2026 that current leading tools “correctly flag solely human-written content as human almost all the time”, but that “they still sometimes make mistakes, so the tools can be used only as starting points for investigation”, and that their results are less illuminating for AI-edited writing. A detector’s verdict alone is still not proof that someone used AI.
7. “The Turing test has been passed”
Partly true In a study published in PNAS in May 2026 (first posted as a preprint in March 2025), two researchers ran a three-party version of the test. “Participants had 5 min conversations simultaneously with another human participant and one of these systems”. “When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time”, while the baseline systems, ELIZA and GPT-4o, were judged to be the human only 23% and 21% of the time. The published paper concludes: “The results constitute empirical evidence that artificial systems can pass a standard three-party Turing test.”
The limits matter. Without the persona prompt, the same models “performed significantly worse (38% and 36%)”. A third, 15-minute study, using LLaMa-3.1 and GPT-5 because GPT-4.5 had been retired, found both persona-prompted models still passed (56% and 59%). And passing shows only that a system can be mistaken for a person in conversation. The authors themselves say the results bear on debates about “what kind of intelligence is exhibited by large language models” — they do not claim the question is settled.
8. “Every AI prompt uses a lot of water”
Disputed Published estimates differ by more than a hundredfold, because they measure different things:
| Estimate | What it covers | Figure |
|---|---|---|
| Academic study (Li et al.) | GPT-3 in Microsoft data centres in the US and other countries (estimate); includes water used to generate the electricity | “a 500ml bottle of water for roughly 10 – 50 medium-length responses, depending on when and where it is deployed” |
| Google (own measurement, 2025) | Median text prompt in its Gemini app; water used for cooling in its data centres | “the equivalent of five drops of water (0.26 mL)” |
| Mistral AI (own life-cycle analysis, 2025) | A 400-token response from its Le Chat assistant | “45 mL of water”, including “upstream emissions” such as manufacturing servers |
The models, the years and the boundaries all differ, and the company figures are self-reported. In short, water use per prompt depends heavily on what is counted, and these numbers cannot be compared directly.
9. “Bigger AI models are always better”
Myth Size helps, but it is not the whole story. OpenAI’s InstructGPT paper stated: “Making language models bigger does not inherently make them better at following a user’s intent.” In its human evaluations, “outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters”.
A 2022 study found that “current large language models are significantly undertrained”, and its smaller, better-trained model, Chinchilla, “uniformly and significantly outperforms Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B)”. The authors concluded that, for compute-optimal training, model size and the amount of training data should be scaled equally.
10. “Being polite to a chatbot gets you better answers”
Disputed The studies point in different directions:
- A 2024 study across English, Chinese and Japanese found “impolite prompts often result in poor performance, but overly polite language does not guarantee better outcomes” and that “The best politeness level is different according to the language.”
- A 2025 report found “sometimes being polite to the LLM helps performance, and sometimes it lowers performance”.
- A short 2025 paper using one model and 50 base questions found “impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts”.
Tone can change results on some tasks, but the effect is inconsistent across studies, models and languages.
What these have in common
- The product is not the model. Apps add search, settings and data policies on top of the model, so “the AI does X” often depends on which app and which settings.
- Dates matter. Several of these answers describe 2023–2025 systems and may change with new models.
- The original source is usually more careful than the headline. Most of the beliefs above are a simplified version of something a paper or policy page says with conditions attached.
Sources
- Anthropic, “Updates to Consumer Terms and Privacy Policy” (28 August 2025).
- OpenAI, “How your data is used to improve model performance” (updated 13 March 2026).
- OpenAI, “GPT-4 Technical Report”, arXiv:2303.08774.
- Vaswani et al., “Attention Is All You Need”, arXiv:1706.03762.
- OpenAI Help Center, “Searching the web with ChatGPT” and “Memory FAQ”; Claude Help Center, chat search and memory; OpenAI API docs, “Web search”; Anthropic docs, “Web search tool”.
- Butlin, Long et al., “Consciousness in Artificial Intelligence: Insights from the Science of Consciousness”, arXiv:2308.08708; Butlin et al., “Identifying indicators of consciousness in AI systems”, Trends in Cognitive Sciences (2025).
- OpenAI, “New AI classifier for indicating AI-written text” (31 January 2023, with July 2023 update).
- Liang et al., “GPT detectors are biased against non-native English writers”, arXiv:2304.02819; Nature, “AI-detection tools have made huge leaps forward — how good are they?” (25 August 2026).
- Jones and Bergen, “Large language models pass a standard three-party Turing test”, PNAS 123(21), e2524472123 (2026); preprint arXiv:2503.23674.
- Li et al., “Making AI Less ‘Thirsty’”, arXiv:2304.03271; Elsworth et al., “Measuring the environmental impact of delivering AI at Google Scale”, arXiv:2508.15734; Mistral AI, “Our contribution to a global environmental standard for AI” (22 July 2025).
- Ouyang et al., “Training language models to follow instructions with human feedback”, arXiv:2203.02155; Hoffmann et al., “Training Compute-Optimal Large Language Models”, arXiv:2203.15556.
- Yin et al., “Should We Respect LLMs?”, arXiv:2402.14531; Meincke et al., “Prompting Science Report 1”, arXiv:2503.04818; Dobariya and Kumar, “Mind Your Tone”, arXiv:2510.04950.
Checked September 2026.
Related: How large language models work · AI benchmarks explained · How to spot a fake quote or an invented source
- artificial intelligence
- ai models
- myths
- fact check
