Reef 看得見那些商用 API 漏掉的假圖
在與兩套領先商用偵測服務的正面對比中,Reef 自研的偵測方法在 WildFake 上拿下最高 AUROC(0.930),並抓出一整類商用產品直接漏掉的舊世代生成器。
By Walking Whale
Product announcement · Reef detection
Reef sees the fakes the big APIs miss
We put Reef's own AI-image detector on the exact same benchmark as two of the best-known commercial detection APIs. On the toughest, generator-balanced test set, Reef came out on top — and it did so while catching a whole family of fakes the others walk right past.
The short version: on a level playing field, Reef is the strongest discriminator of the three — the highest AUROC on WildFake, the best balanced operating point, and the only detector that covers the older generators the commercial products consistently miss.
First, the metrics in plain words
Three short definitions, and every chart below reads at a glance. No statistics background needed.
| The metric | What it's really asking | Reef's result |
|---|---|---|
| AUROC | Show it one real photo and one fake — how often does it pick the fake as the more suspicious of the two? (1.00 = perfect, 0.50 = a coin flip) | 93 times out of 100 (0.930) |
| Catch rate (recall) | Of all the fakes in the pile, how many does it flag? | ~89 of every 100 |
| Clear rate (specificity) | Of all the genuine photos, how many does it correctly leave alone — no false alarm? | ~85 of every 100 |
| Overall grade (balanced accuracy) | The fair average of catching fakes and clearing reals, so a lopsided test can't flatter the score. | 0.87 — best of the three |
A cleaner read than any single score
AUROC is the fair headline: it doesn't depend on where you draw the line, so it rewards raw skill at telling fake from real and doesn't flatter anyone's chosen setting. On WildFake — a balanced set of 250 fakes and 250 real photos — Reef out-ranks both commercial detectors by a clear margin.
In numbers: Reef 0.930, A*****T 0.801, S*********e 0.779. Reef stays strong on the second test set too — 0.906 on tiny_genimage, comfortably above S*********e and close behind A*****T. Across both collections, Reef is the only detector that never finishes last.
Where each detector wins
| What we measured | Who wins | Reef's number |
|---|---|---|
| Telling fake from real (AUROC) · WildFake | Reef | 0.930 — highest |
| Best balanced overall grade · WildFake | Reef | 0.87 — highest |
| Catching the older generators both rivals miss | Reef | ~90–100% caught |
| Catching today's models (Stable Diffusion, Midjourney) | All three tie | ~100% — near-perfect for everyone |
The coverage win — catching what the others can't
This is the heart of it. Both commercial detectors share a blind spot: the older families of image generators. Reef covers exactly that gap — catching those older fakes almost every time, where the turnkey products catch almost none. On the modern generators everyone already handles (Stable Diffusion, Midjourney) all three are near-perfect.
| Kind of generator | Reef | A*****T | S*********e | Reef's lead |
|---|---|---|---|---|
| ADM (older diffusion) | 91% | 9% | 0% | +82 pts |
| VQVAE / token | 100% | 40% | 20% | +60 pts |
| Classic GAN | 90% | 30% | 0% | +60 pts |
| Advanced autoencoder | 90% | 30% | 40% | +50 pts |
| Imagen | 100% | 36% | 55% | +45 pts |
| Advanced GAN | 100% | 80% | 20% | +20 pts |
| DDPM diffusion | 91% | 73% | 9% | +18 pts |
| DDIM diffusion | 100% | 91% | 36% | +9 pts |
| DALL·E | 100% | 91% | 100% | tie |
| Midjourney | 100% | 100% | 100% | tie |
| Stable Diffusion | 100% | 100% | 100% | tie |
| VQDM diffusion | 70% | 70% | 40% | tie |
Catch rate per generator family on WildFake (share of that family's fakes flagged). Reef leads or ties in every row, and opens double-digit gaps on the eight older families up top.
The same pattern holds on the second test set. On tiny_genimage, Reef again catches the older generators (ADM, VQDM, BigGAN) that the commercial products let through, and ties everyone on the modern ones.
A detector is only as good as its worst blind spot. Anyone can flag today's headline models; the images that slip through are the ones from generators a product was never tuned for. Reef closes that gap without any per-generator retraining.
The scorecard
AUROC alongside each detector's best balanced setting on WildFake — the fair "best each can do". Column names use the plain words from the table above; the technical term is in brackets. Reef leads on both raw skill and overall grade.
| Detector | AUROC · WildFake | Catch rate (recall)* | Clear rate (specificity)* | Overall grade (bal. acc.)* |
|---|---|---|---|---|
| Reef | 0.930 | 89% | 85% | 0.87 |
| A*****T | 0.801 | 68% | 81% | 0.74 |
| S*********e | 0.779 | 57% | 98% | 0.78 |
* Catch rate, clear rate and overall grade at each detector's own best-balanced setting on WildFake.
How Reef does it
Briefly, and without the internals: Reef reads each image through a frozen, general-purpose vision model used as-is, with no task-specific training, and compares it against a curated reference library of real and AI-made examples. There is no fine-tuning and no per-generator retraining — which is exactly why Reef generalises to generators it was never built around.
Every Reef number here comes from a strict "never seen it before" test: the detector is never shown the image collection it is being scored on. These are results from the hardest setting we could design — not a home-field one.
One eye, broader coverage
Turnkey APIs are excellent at the models they were tuned for. Reef's edge is breadth: the highest AUROC on the balanced benchmark, the best-balanced setting, and coverage of the older generators that quietly slip past everyone else — all from a single, general-purpose eye. That is the detector we're building on.