NewsLabs
Log in
Book a demo
← All stories
INDUSTRY

Inside Google's AI Detection Toolbox

Inside Google's AI Detection Toolbox

Written by

Marko Dobrinić
Marko DobrinićML Engineer & AI Consultant

Marko is an ML engineer with over 8 years of experience creating and applying AI across various domains. A data scientist first, turned software engineer, he loves tinkering with data and crafting real-world solutions with the help of machine learning.

LinkedIn
Inside Google's AI Detection Toolbox

Every few weeks, someone on LinkedIn declares a glorious Google victory over AI content, claiming the search giant now buries every page ChatGPT ever touched. Every few weeks, someone else announces you should publish 9000 AI articles before lunch break. More likely than not, both of these types of people have got something to sell to you.

But, when you dig a bit deeper to find the truth, Google doesn't have a single silver bullet AI detector. They’ve built a toolbox where each tool answers a different question. This is part one: the machinery, in part two we will answer whether any of it touches your rankings.

Google AI detection systems under the hood
Google AI detection systems under the hood

The invisible ink, mathematically speaking

SynthID is Google DeepMind's watermarking toolkit for text, images, video and audio which started out marking only Google's own models (Gemini, Imagen, Veo, Lyria) and by late 2026 it has now become something closer to industry plumbing.

Rigging the game

Language is a chain of discrete word choices, so you can't hide a signature by nudging pixels. Google DeepMind's 2024 Nature paper give a solution by using Tournament sampling. It’s as if during the knockout stage of the World Cup every referee has been bribed and has a secret favorite.

When Gemini picks its next word, it first draws a few possible candidates ("quick," "fast," "swift"). They face off in a bracket, and a scoring function seeds them by the last few words plus a secret key that decides each match. The winner is always a word the model might have chosen anyway, so the text reads normally. But over hundreds of words, the referee's bias adds up, and anyone holding the key can check whether the "winners" won suspiciously often. One rigged match can’t prove much but a rigged season is obvious. The checker doesn't even need to run Gemini.

The cost of implementation is quite small, where DeepMind's benchmark, showed only a 0.57% slowdown. A live test on nearly 20 million Gemini responses found users rated watermarked and unwatermarked answers the same (a test of perceived quality, not detection accuracy), so it’s also quite hard to tell.

If you ask "What's the capital of France?" then there's one possible candidate: Paris. A one-player tournament can't be rigged, so short or factual answers will carry little to no signal. Google open-sourced the method in October 2024, but we’ll come to that later.

Dots on paper

Many color laser printers will add a barely visible pattern of tiny yellow dots to every page, which enable identifying the machine that printed it. SynthID does the the same digitally, embedding a pattern into the pixels as image is being generated, and into every frame of the video. Google's SynthID-Image paper says the mark will survive cropping, compression and filters. Interestingly enough, by May 2026 Google had marked more than 100 billion images and videos, plus 60,000 years of audio, where the signature lives in the waveform.

SynthID only detects SynthID. If the mark is there, a participating model made the content. If it's missing, you’re not close to an answer, same as if a car without a VIN plate may have had it filed off, since "Not watermarked" does not mean "not AI."

The text mark is also easy to wash out. Xia Han's team showed that just by paraphrasing and using back-translation will significantly weaken detection, as if you are retyping a printed letter so the yellow dots don't come along. Omidi, Dong and Wang built an attack that breaks SynthID's simplest scoring by stacking extra tournament rounds on top, but Google's Bayesian detector holds up better. ETH Zurich's SRI Lab gave a nice overview in their blogpost - forging the mark is hard and scrubbing it is easy. Translate a paragraph to Croatian and back, and the watermark is gone. Images are much harder to launder.

Where it lives

Google built SynthID checks into the Gemini app, where people used them 50 million times, then brought them to Search on May 19, 2026. You can ask Lens, AI Mode or Circle to Search "Is this AI generated?" about an image. Same functionality has been added to Chrome. Google hasn't said whether any of this feeds rankings.

Google is not the only one using it. NVIDIA marks its Cosmos video, OpenAI began stamping SynthID on images in May 2026 and audio in July. Kakao and ElevenLabs also started using it. And because the text method is open source, Anthropic built its own version for Claude in August 2026, with its own key, citing the EU AI Act.

When you sample actual data in the wild you don’t see the trend really; a 2026 study of AI images on X found SynthID in only 18% of sampled OpenAI images, and no C2PA metadata at all, because X strips it on upload.

S-CTS and SAFE hunting for the slop farms

We know SynthID marks files, but Google's anti-slop research forgets about the files and focuses on the content factories. Old-school moderation worked like a bouncer with a list of banned faces – if you spot the same clip twice, block it. A slop farm won’t upload the same clip twice, rather, It will upload tens of thousands slightly different ones. What you have to do is track down their base of operations.

In July 2025, YouTube renamed its "repetitious content" rule to "inauthentic content," basically banning mass-produced videos from monetization while saying "AI-assisted" channels remain eligible.

Asking for a second opinion

The Scalable Cluster Termination System (S-CTS) comes from a 2026 Google Research paper built for "an online video platform" (there is no mention of YouTube). It works like a careful doctor when you walk in sick. They will want you to do an X-ray or some blood work, since cough alone won’t give you a pneumonia diagnosis and couple days off work.

The first test, the Psi-A component, is the symptom check which never sees a frame of the video. It just studies the overall plumbing - how the API is used, upload timings and generation metadata. Five hundred "different" channels uploading from the same setup within seconds of each other is like a nagging cough that won't go away. The second test, the Psi-C component, is the X-ray: it scores the videos for templated narratives and "Generative Artifacts."

Ban will happen only when both tests come back positive. A solo creator using AI has the cough but no dark spots on the X-ray, so Psi-A never flags them. If a network of 500 accounts is pumping near-identical videos from shared infrastructure, that will trigger both components.

Under the hood, the system combines each channel into a short text summary and hands it to a Gemini model that’s tuned with LoRA. LoRA is the adapter layer that enables Google to use it on top of any other video model in order to detect slop, whether it’s Sora or Kling.

Google reported 92% to 95% precision, up to 96% recall and an overturn rate below 1%, with weak spots like 68% recall on generative NSFW content. The famous 50,000 clusters and 130,000 channels come from Search Engine Journal's coverage of an earlier version, but the paper Google hosts today doesn't contain them.

The detective squad

A sibling paper, "The Synthetic Gap", describes SAFE, the Scaled Abuse Forensics Examiner. Where S-CTS returns a score, SAFE works a case like a detective unit. You get a lead investigator plus three forensic specialists covering the content, with behavior and the network map of the accounts.

The content agent hunts videos that break a rule and videos that dodge a rule while breaking its spirit, the moderation equivalent of a kid hovering his finger a millimeter from his sister's face. The behavior agent's example finding is quite specific, where "100% of channels utilize identical OS versions and upload within the same 5-second window." It’s pretty obvious behaviour that’s easy to spot.

In the fine print SAFE claims early deployment "significantly accelerates" threat detection but publishes no numbers. And like S-CTS, nothing in it mentions Google’s web Search.

Brains of the operation

SpamBrain is Google Search's machine-learning spam platform, running since 2018. This is a tool that we know least about from direct Google info, but it matters most for the websites. It works like an inbox spam filter where it doesn't care whether a person or a script sent the "you’ve won a prize" email, only that it looks and spreads like a scam.

Google's last detailed webspam report, covering 2022, said SpamBrain caught 200 times more spam sites than at launch and kept more than 99% of Search visits spam-free. What model is used, which signals and thresholds are used, Google still keeps a secret, for the same reason casinos don't hand out their card-counting manual.

Google's stance on AI text has done a full 180. In April 2022, John Mueller said he "can't claim" Google could detect AI writing, and auto-generated text was spam by definition. By February 2023, Google rewarded quality "rather than how content is produced." If you look at today's spam policies, they cover everything from "attempting to manipulate generative AI responses," to defined scaled content abuse as low-value pages at scale, "no matter how it's created." AI usage isn't a big deal anymore, but you cannot just dish out thousands of pages of crap.

The human layer

Thousands of external Search Quality Raters score sample results using a public handbook. Since January 2025, raters gave the lowest rating to mass-produced pages made with automated tools, "generative AI or otherwise." On October 1, 2026, Google added that it's "critical to manually factcheck and review all AI-generated content" before publishing. The qualities raters now look for are effort, originality, talent or skill, and accuracy. Authorship isn't on that list.

Beyond the big three

AI checks have quietly spread across Google's products. Nearly all of them read a label someone else attached or ask you to confess, and only one confirmed system looks at the content and makes its own call.

AI-related systems in other Google products
AI-related systems in other Google products

YouTube's labeler is the closest thing to an in-production AI detector Google has confirmed. Creators can contest a wrong label, and YouTube says a label alone doesn't hurt your recommendations or monetization. What they won't say is how those signals actually work in practice. Claims that they read SynthID aren't in YouTube's announcement.

If we take a look at the Ads, Google’s trust and safety chief wrote in September 2024 that C2PA signals would inform how ads policies are enforced. Nobody at Google has said anything similar about organic Search.

Then there's the 2024 leak of Google's internal Content Warehouse API documentation. Among the page-quality signals sits contentEffort: "LLM-based effort estimation for article pages." It just estimates effort, not authorship, and Google never confirmed how or whether it's used.

We couldn’t find a concrete proof of a Google’s patent or paper describing a general tool that tells human text from AI text. The AI-detection patents on Google Patents belong to other companies. Google's own published text work is mostly just watermarking, going back to a 2011 method for marking Google Translate output.

Smoke detectors, toast and AI text

Why doesn't Google just build a detector that reads a paragraph and flags the AI? Because nobody can build one that works in the wild with great accuracy.

The problem looks like a badly calibrated smoke detector. If its tuned sensitive enough to catch every fire, it will go off whenever someone makes a toast. Turn the sensitivity down, and it will miss a real fire. AI-text detectors face the same trade-off with human writing being the toast. The harder a detector tries to catch AI text, the more genuine human work it wrongly flags.

RAID, the biggest public stress test validated 12 detectors against more than 6 million texts. They used simple tricks, like swapping letters for look-alike characters, pushed error rates past 95%, and several detectors on default settings flagged over 10% of human writing as AI. OpenAI's own classifier, launched and withdrawn in 2023, caught just 26% of AI text while wrongly accusing 9% of humans, which is like a smoke detector that misses three fires in four and still goes off at breakfast.

Stanford study found seven popular detectors wrongly labeled 61% of non-native English speakers' TOEFL essays as AI on average. Clear, simple sentences look "machine-like," which punishes people writing carefully in a second language.

26%AI-generated text caught by OpenAI classifier
61%non-native English speaker text labeled as AI

Detectors could ace tests as well, where several teams topped 99% accuracy but on models they'd trained on. If we add a different model or a determined editor, which is what the web actually looks like, the numbers won’t look anything close to this. That's why the field has moved from guessing after the content has been written to marking at birth.

The law is catching up

The EU has written provenance marking into law. Since August 2, 2026, Article 50 of the EU AI Act has placed obligations on two parties. Providers, or the companies that build AI tools, must mark their output in a machine-readable, detectable format "as far as technically feasible." Deployers, which includes any newsroom publishing with those tools, carry a separate duty of disclosure.

For providers, the Code of Practice requires two layers of marking: cryptographically signed C2PA metadata and an invisible watermark. The two complement each other, metadata records detailed information about how a file was made, but a screenshot or an upload to many platforms removes it. A watermark carries less information but remains in the content itself. AI systems already on the market before August 2 have until December 2, 2026 to comply, and violations can bring fines of up to €15 million or 3% of global annual turnover.

For publishers, the relevant provision is Article 50(4), and it does not require a label on every article produced with AI tooling. As regulatory consultant Marko Đuričić explains in our EU AI Act series, the disclosure duty applies only when three conditions are met: the text is published, it is intended to inform the public, and it concerns a matter of public interest. Even then, publishers are exempt if a person has substantively reviewed the content, including checking the facts, and a named person or entity holds editorial responsibility for it. Correcting typos, having one AI system review another, or adding a nominal approval step does not qualify. If AI substantively changes the article after an editor has signed off, the exemption no longer applies. An article written by a journalist who used AI only for brainstorming, headlines or grammar may fall outside the rule entirely.

California's SB 942 took effect the same day, requiring large AI providers to offer free detection tools.

There are still unknowns

We don't know whether the slop-farm tech has reached web Search, what's inside SpamBrain, or how Europe will enforce Article 50. And we found no evidence Google can reliably spot AI text from other companies' models.

Three things to keep in editorial brain:

  1. A watermark proves a positive, never a negative.
  2. Text watermarks are the most fragile. Paraphrasing usually removes them.
  3. Detecting AI and punishing AI are different questions. Google's Search systems, as documented, chase spam and low effort and not authorship.

Basically Google can see watermarks, slop networks and spam patterns. What it does with that when it ranks your page you will be able to read in part two: “Will Google Penalize Your AI Content?” with the ranking data from Ahrefs, Semrush, SE Ranking and Lily Ray.

Written by

Marko Dobrinić
Marko DobrinićML Engineer & AI Consultant

Marko is an ML engineer with over 8 years of experience creating and applying AI across various domains. A data scientist first, turned software engineer, he loves tinkering with data and crafting real-world solutions with the help of machine learning.

LinkedIn

Get the next one in your inbox.

New posts from the NewsLabs blog, now and then. No noise.

A few emails a month. See our privacy notice.

Subscribe to our newsletter!

A few emails a month. See our privacy notice.