Finding AI-Generated Faces in the Wild

Evaluating synthetic-face detection across GAN and diffusion engines, including held-out generators and reduced image quality.

Our 2023 detector exploited the rigid facial geometry of StyleGAN images. But the generative landscape did not hold still. Stable Diffusion, DALL-E 2, and Midjourney could also generate faces, without StyleGAN’s alignment. Upload pipelines could then downscale and recompress those images.

In Finding AI-Generated Faces in the Wild, our LinkedIn team and Professor Hany Farid at UC Berkeley evaluated detection across GAN and diffusion engines, including generators withheld from training. We also tested reduced resolution and JPEG compression, using separately resolution-matched models for the small-image results. We published this coauthored paper at the Workshop on Media Forensics at CVPR 2024.

The production model described here had already been replaced by the time we published the paper.

One Classifier, Ten Engines

We trained and evaluated against 18 datasets: 120,000 real profile photos from LinkedIn members, plus another 105,900 synthetic images spanning ten generation engines. The GAN side includes generated.photos, StyleGAN 1 through 3, and EG3D; the diffusion side includes DALL-E 2, Midjourney, and Stable Diffusion 1, 2, and xl. Six engines were used for training; four were held out entirely to test generalization.

Grid of representative AI-generated face and non-face images from ten synthesis engines including generated.photos, StyleGAN 1 to 3, EG3D, DALL-E 2, Midjourney, and Stable Diffusion variants

Representative AI-generated images from the ten synthesis engines used for training and evaluation. Some engines contribute faces only; others contribute both faces and non-face images. (Figure 2 of the paper; dataset counts in Table 1.)

At a fixed 0.5% false positive rate, the classifier detected 98% of AI-generated faces from engines seen in training. We reported 84.5% detection on the held-out evaluation of 5,000 faces from four engines at the same false-positive rate (Table 2). Performance varied considerably by generator: EG3D reached 99.5% and generated.photos 95.4%, while Midjourney mostly slips through (19.4%). Adding examples from new engines to training is one possible response.

Built for the Wild

The “in the wild” part is the point of the paper. Profile photos do not arrive as pristine megapixel originals; they get downscaled and JPEG-compressed, sometimes repeatedly. Detectors that depend on fragile pixel-level traces tend to die somewhere in that pipeline.

The architecture is straightforward: images are resized to 512 pixels and fed through an EfficientNet-B1 backbone, with the backbone frozen and 6.8 million parameters of scoring layers trained on top. For compression robustness, the training data mixes uncompressed images with JPEG-compressed ones across a range of quality levels.

Two plots showing true positive rate versus image resolution and versus JPEG quality, with resolution-matched training maintaining high accuracy at small sizes

True positive rate as a function of resolution (top) and JPEG quality (bottom) at a fixed 0.5% false positive rate. In the top panel, the solid curve is the 512-trained model evaluated at lower resolutions; each point on the dashed curve uses a model trained at the matching resolution. The JPEG model was trained on uncompressed images and a range of JPEG qualities. (Figure 3 of the paper.)

A model trained only at 512 pixels lost most of its detection power on 128-pixel images. A separate model trained and evaluated at 128 pixels stayed around 90%, at the same 0.5% false positive rate. For the model trained with mixed uncompressed and JPEG images, detection degraded as JPEG quality fell from 100 to 20. Section 4 reports 94.3% TPR at quality 80 and 88.0% at quality 60, both at 0.5% FPR.

What the detector appears to use

The detector flagged none of the synthetic non-face images in this evaluation: their true positive rate was 0% (Table 2).

AI-generated faces alongside their integrated-gradient attribution maps, which concentrate on facial regions

Integrated-gradient attributions for AI-generated faces concentrate around the face and other areas of skin. The top row averages 100 StyleGAN 2 faces together; the others are individual examples. (Figure 5 of the paper.)

The attributions discussed in Section 4.1 concentrate on facial regions. These observations are consistent with the detector using facial structure. The difference between real and synthetic training images—only the real set contained some non-faces—makes that interpretation harder to isolate.

Resources