Do AI writing detectors actually work? 5 tools tested for accuracy

6

We live in an era where a simple scroll through social media can trigger paranoia. Did a human write this? Or did an LLM spit it out?

The answer matters. Whether you are grading essays, hiring writers, or just tired of reading sterile content online, you want to know the truth. Supposedly, AI leaves fingerprints. The em dash is the usual suspect. Monotone tone. The other one.

But can tools actually spot these tics reliably?

I ran five popular AI detectors against a controlled experiment. On one side: four articles I wrote by hand (no bots). On the other: 150-word intros generated by ChatGPT, Gemini, and Claude based on my article titles.

The results? Some tools nailed it. Others missed the mark hard.

Pangram: The confident checker

Pangram claims to be the detector that “actually works.” It’s not free forever, but the free tier allows four checks a day. Paid plans start at $20.

Here’s what happened:
* My writing: 100% Human. High confidence.
* ChatGPT/Claude text: 100% AI.

Pangram correctly identified every sample. It even highlighted specific giveaways, like the repetitive “from the moment you…” structure in the AI samples.

Verdict: 4 / 4 correct.

Grammarly: The grammar giant enters the chat

You probably use Grammarly for typos. Now it flags AI too. Free accounts get three checks daily. Paid plans kick in at $12.

My samples? Zero percent AI. No patterns. Clean slate.

The AI samples? Not as black-and-white.
* Claude: 68% AI.
* Gemini: 66% AI.

Grammarly spotted the AI fingerprints. But it wasn’t fully convinced. It saw the patterns but hesitated on the verdict. Still, it got the call right.

Verdict: 4 / 4 correct.

GPTZero: Preserving the human

GPTZero’s mission is “preserving what’s human.” Free users can scan 10,000 words a month. Paid plans start around $24.

By the third tool, I was getting cocky. My writing style must be distinctly non-AI. GPTZero agreed. “Highly confident this text is entirely human.”

When fed the AI samples, it caught them too. It even pointed out which sentences looked most robotic. I struggled to see the pattern. Maybe it’s there. Maybe not. But the tool didn’t miss them.

Verdict: 4 / 4 correct.

Scribbr: The free outlier

Scribbr offers editing services and a free detector. No account needed. Just paste and check.

My writing? Cleared as 100% human. Good.

The AI samples? Also marked as 100% human-written.

This is where it broke. Scribbr claimed the ChatGPT and Claude texts were fully human. The tool seemed very confident in its wrong answer. Whatever its algorithm is, it needs an update.

Verdict: 2 / 4 correct.

Copyleaks: The half-right contender

Copyleaks checks text, images, and video. Four free scans. Paid tiers start at $17.

My text? 100% human-free.

The AI test? Mixed bag.
* Gemini: 100% AI (Correct).
* Claude: 0% AI (Incorrect).

Copyleaks caught one bot but let the other slide. It seems to struggle with calibration depending on which model you’re testing.

Verdict: 3 / 4 correct.

The final takeaway

I wasn’t sure what to expect. That’s why I tested them. But now? It’s clearer.

AI detectors are not perfect.

They worked consistently for identifying my human writing. Every single tool flagged my text as human. It seems easier to confirm what isn’t AI than what is.

The reverse wasn’t true. Scribbr failed badly. Copyleaks missed half the bots. The others were perfect in this tiny sample.

If you’re serious about catching AI writing, don’t trust one tool. Use two or three in tandem. Look for patterns. And maybe upgrade for detailed reasoning if the stakes are high.

Is there a definitive tell? Maybe. Or maybe it’s just a cat-and-mouse game that changes with every model update.

Disclosure: Popular Science’s parent company sued OpenAI in April 2024 over copyright infringement. This context may influence the landscape, though the test results above reflect the tools’ performance in isolation.

Попередня статтяHow AI Optimization Risks Worsening NAEP Score Stagnation