The short answer
You usually cannot tell for certain, and you should be deeply suspicious of anyone, or any tool, that claims otherwise. AI detection software does not work reliably: it produces false accusations against real human writing, especially from people writing in a second language or in a plain, formal style, and it can be defeated by light editing. No detector should ever be the basis for accusing someone of anything.
What you can do is read for signals and, far more importantly, shift the question. Instead of asking "did a machine write this", ask "is this true, is it sourced, and does the person putting their name to it stand behind it". That question has a checkable answer. The first one usually does not, and it is getting less answerable every year.
Why detectors do not work
The pitch is appealing: paste in text, get a percentage. The reality is worse than the pitch in several specific ways.
They produce false positives on real human writing. Detectors typically look for text that is statistically unsurprising: common words in common orders. Plenty of humans write that way on purpose. Clear technical writing, formal academic prose, business documents that follow a template, and writing by people working in a second language all tend to be more predictable than an off-the-cuff blog post. That means the people most likely to be wrongly accused are often those least equipped to defend themselves, which is a serious fairness problem and not a hypothetical one.
They produce false negatives whenever anyone tries. Rewriting a few sentences, asking the model for a more idiosyncratic style, or simply editing the text by hand pushes output past most detectors. The people actually trying to deceive you are the ones most likely to pass.
The percentage is not a probability. "92 percent AI" does not mean there is a 92 percent chance a machine wrote it. It is an internal score on an unstated scale, and different tools give different numbers for the same paragraph. Try it: paste something you wrote yourself into three detectors and watch them disagree.
And the ground is moving. Detectors are trained against yesterday's models. Every new model release makes the training slightly stale.
Never use a detector score as evidence against a person. If you are a teacher or a manager, a score is not proof, it is a machine's guess about a machine. Ask the person to talk you through their work instead. That conversation tells you far more, and it is fair.
Signals worth noticing
If detectors are unreliable, what is left is your own reading. None of these are proof, and every one of them appears in genuine human writing too. Treat them as reasons to look closer, never as a verdict.
Fluent but hollow. The clearest tell is not a phrase, it is a feeling: everything reads smoothly and nothing lands. Paragraphs restate the heading and then move on. You reach the end and cannot say what the piece argued.
Suspicious symmetry. Three examples for every point. Every list exactly the same length. Every paragraph roughly the same size. Human writing is lumpier, because some points genuinely deserve more room.
No friction. Real expertise contains hedges, exceptions and the odd strong opinion. Machine-generated text often refuses to commit, balancing every claim so evenly that nothing is actually said. Related tell: it never mentions a downside that costs anything.
Generic specifics. Details that sound concrete but name nothing: "a leading study", "many experts", "significant improvements", "in recent years". Genuine specificity has names, dates and sources attached to it.
Confident detail that is subtly wrong. The most important signal in practice. Statistics with no source, quotations attributed to the wrong person, citations to papers that do not exist, plausible legal or medical claims that are close to right but not right. Language models generate fluent text whether or not the underlying claim is true, so wrongness arrives in the same confident tone as correctness.
Timeline and context errors. A piece that describes something recent as though it were still forthcoming, or that misses an obvious development a knowledgeable person would have mentioned.
Odd repetition of framing. The same connective structures over and over, or a habit of restating your question back at you before answering it.
Signals that are not reliable at all, despite what you may have read: em dashes, particular favoured words, or a fondness for "delve". Punctuation preferences are style, plenty of careful human writers use them, and models change their habits with every release. Accusing someone based on their punctuation is a good way to be both rude and wrong.
The skill that actually lasts: verification
Here is the reframe that matters. In a year, machine text will be harder to spot than it is today. The verification question stays exactly as answerable as it has always been, because it has nothing to do with who typed the words.
Check the checkable. Any specific claim (a number, a date, a quotation, a law, a study) can be traced to a source. Do it for the claims that would change your decision. Not every sentence, just the load-bearing ones.
Follow the citation, do not admire it. A reference is only worth something if it exists and says what the text claims it says. Invented citations are a common failure mode, and they look completely normal until you look them up.
Ask who is accountable. Is there a named author, an organisation with a reputation to lose, a correction policy, a date? Anonymous, undated content is weak evidence whoever produced it.
Look for a second, independent source. Not another article repeating the same claim, but a genuinely separate one. If every result traces back to a single origin, you have one source, not five.
Raise your standards with the stakes. A restaurant recommendation needs almost no checking. Anything touching your health, your money, your legal position or your safety needs a real source, and often a real professional.
A practical habit for chatbot answers specifically: ask the tool which parts of its answer you should verify independently, and where it might be wrong. It is not a guarantee, but a model prompted this way will often flag its own shakiest claims, which tells you where to point your attention first.
If you have to make a judgement about a person
Sometimes you genuinely have to decide: a manager reviewing a report, a teacher marking an essay, someone reading an application.
Do not lead with an accusation, and do not lead with a detector score. Ask about the work. "Talk me through how you got to this recommendation." "Which of these sources was most useful and why?" "What did you consider and reject?" Someone who did the work can answer immediately. Someone who pasted an output cannot, and neither of you has to argue about a percentage.
Then be clear about what you actually want going forward. Most people are not trying to deceive anyone, they are using an ordinary tool without knowing your expectations. "Use AI to draft and edit, but you own every fact and you must be able to explain any part of it" is a rule people can follow. "Do not use AI" is a rule you cannot enforce and probably do not mean.
The durable skill here is not detection. It is judgement: knowing which claims matter, checking those, and being clear about who is accountable for the result. That skill was valuable before any of this existed, and it will still be valuable when detectors have been quietly retired.
Our free AI Safety, Privacy and Verification course goes further into all of it: how models get things confidently wrong, what to do about your privacy, and a practical verification routine you can actually keep up.