product

Reef 看得見那些商用 API 漏掉的假圖

在與兩套領先商用偵測服務的正面對比中,Reef 自研的偵測方法在 WildFake 上拿下最高 AUROC(0.930),並抓出一整類商用產品直接漏掉的舊世代生成器。

By Walking Whale

Reef mark — the Walking Whale reef icon

Product announcement · Reef detection

Reef sees the fakes the big APIs miss

We put Reef's own AI-image detector on the exact same benchmark as two of the best-known commercial detection APIs. On the toughest, generator-balanced test set, Reef came out on top — and it did so while catching a whole family of fakes the others walk right past.

Reef · AUROC 0.930 A*****T · 0.801 S*********e · 0.779 Same 500 images · WildFake

The short version: on a level playing field, Reef is the strongest discriminator of the three — the highest AUROC on WildFake, the best balanced operating point, and the only detector that covers the older generators the commercial products consistently miss.

93 / 100Hand Reef one real photo and one fake, and it ranks the fake as more suspicious about 93 times out of 100 (AUROC 0.930) — the highest of the three
89 / 85Catches ~89 of every 100 fakes, and correctly clears ~85 of every 100 genuine photos — the best-balanced setting of the three
9 in 10Its catch rate on the older kinds of fakes both commercial products miss almost entirely

First, the metrics in plain words

Three short definitions, and every chart below reads at a glance. No statistics background needed.

The metric What it's really asking Reef's result
AUROC Show it one real photo and one fake — how often does it pick the fake as the more suspicious of the two? (1.00 = perfect, 0.50 = a coin flip) 93 times out of 100 (0.930)
Catch rate (recall) Of all the fakes in the pile, how many does it flag? ~89 of every 100
Clear rate (specificity) Of all the genuine photos, how many does it correctly leave alone — no false alarm? ~85 of every 100
Overall grade (balanced accuracy) The fair average of catching fakes and clearing reals, so a lopsided test can't flatter the score. 0.87 — best of the three

A cleaner read than any single score

AUROC is the fair headline: it doesn't depend on where you draw the line, so it rewards raw skill at telling fake from real and doesn't flatter anyone's chosen setting. On WildFake — a balanced set of 250 fakes and 250 real photos — Reef out-ranks both commercial detectors by a clear margin.

Bar chart in Reef brand colors: AUROC on WildFake — Reef 0.930, A*****T 0.801, S*********e 0.779
AUROC on WildFake — higher is better. Reef (green) leads both commercial detectors.

In numbers: Reef 0.930, A*****T 0.801, S*********e 0.779. Reef stays strong on the second test set too — 0.906 on tiny_genimage, comfortably above S*********e and close behind A*****T. Across both collections, Reef is the only detector that never finishes last.

Where each detector wins

What we measured Who wins Reef's number
Telling fake from real (AUROC) · WildFake Reef 0.930 — highest
Best balanced overall grade · WildFake Reef 0.87 — highest
Catching the older generators both rivals miss Reef ~90–100% caught
Catching today's models (Stable Diffusion, Midjourney) All three tie ~100% — near-perfect for everyone

The coverage win — catching what the others can't

This is the heart of it. Both commercial detectors share a blind spot: the older families of image generators. Reef covers exactly that gap — catching those older fakes almost every time, where the turnkey products catch almost none. On the modern generators everyone already handles (Stable Diffusion, Midjourney) all three are near-perfect.

Grouped bar chart in Reef brand colors: catch rate by generator family on WildFake — Reef near 100 percent across the older generators the commercial products miss
WildFake — share of each generator's fakes that get flagged. Reef (green) stays high across the board; the commercial detectors collapse on the older families at the top.
Kind of generator Reef A*****T S*********e Reef's lead
ADM (older diffusion)91%9%0%+82 pts
VQVAE / token100%40%20%+60 pts
Classic GAN90%30%0%+60 pts
Advanced autoencoder90%30%40%+50 pts
Imagen100%36%55%+45 pts
Advanced GAN100%80%20%+20 pts
DDPM diffusion91%73%9%+18 pts
DDIM diffusion100%91%36%+9 pts
DALL·E100%91%100%tie
Midjourney100%100%100%tie
Stable Diffusion100%100%100%tie
VQDM diffusion70%70%40%tie

Catch rate per generator family on WildFake (share of that family's fakes flagged). Reef leads or ties in every row, and opens double-digit gaps on the eight older families up top.

The same pattern holds on the second test set. On tiny_genimage, Reef again catches the older generators (ADM, VQDM, BigGAN) that the commercial products let through, and ties everyone on the modern ones.

Grouped bar chart in Reef brand colors: catch rate by generator family on tiny_genimage — Reef matches or beats both commercial products on every family
tiny_genimage — Reef (green) matches or beats both products on every generator family, and is far ahead on ADM, VQDM and BigGAN.
Why this matters

A detector is only as good as its worst blind spot. Anyone can flag today's headline models; the images that slip through are the ones from generators a product was never tuned for. Reef closes that gap without any per-generator retraining.

The scorecard

AUROC alongside each detector's best balanced setting on WildFake — the fair "best each can do". Column names use the plain words from the table above; the technical term is in brackets. Reef leads on both raw skill and overall grade.

Grouped bar chart in Reef brand colors: catch rate, clear rate and overall grade at each detector's best-balanced setting on WildFake — Reef leads catch rate and overall grade
At each detector's best-balanced setting on WildFake. Reef (green) leads on catch rate and on the overall grade; it stays strong on clear rate too.
Detector AUROC · WildFake Catch rate (recall)* Clear rate (specificity)* Overall grade (bal. acc.)*
Reef 0.930 89% 85% 0.87
A*****T 0.801 68% 81% 0.74
S*********e 0.779 57% 98% 0.78

* Catch rate, clear rate and overall grade at each detector's own best-balanced setting on WildFake.

How Reef does it

Briefly, and without the internals: Reef reads each image through a frozen, general-purpose vision model used as-is, with no task-specific training, and compares it against a curated reference library of real and AI-made examples. There is no fine-tuning and no per-generator retraining — which is exactly why Reef generalises to generators it was never built around.

Tested the hard way

Every Reef number here comes from a strict "never seen it before" test: the detector is never shown the image collection it is being scored on. These are results from the hardest setting we could design — not a home-field one.

One eye, broader coverage

Turnkey APIs are excellent at the models they were tuned for. Reef's edge is breadth: the highest AUROC on the balanced benchmark, the best-balanced setting, and coverage of the older generators that quietly slip past everyone else — all from a single, general-purpose eye. That is the detector we're building on.

文章