Sanyam Jain / panodiff
Preprint  —  arXiv:2507.09227  ·  under revision, Applied Soft Computing

PanoDiff-SR

Synthesizing dental panoramic radiographs using diffusion and super-resolution

Sanyam Jain1, Bruna Neves de Freitas1,2, Andreas Basse-O'Connor3, Alexandros Iosifidis4, Ruben Pauwels1

1Department of Dentistry and Oral Health, Aarhus University, Denmark 2Aarhus Institute of Advanced Studies, Aarhus University, Denmark 3Department of Mathematics, Aarhus University, Denmark 4Faculty of Information Technology and Communication Sciences, Tampere University, Finland

Every radiograph above is synthetic. None of them is a real patient. Each was shown to six dentists for twelve seconds, and the number under each one is how many of the six called it real.

1Before anything else, take the challenge

Real radiograph or synthetic? Twelve seconds per image, two buttons, no second look. Choose 50, 100 or 200 images, put your name on the live leaderboard, and see which generators fooled you.

This is not the observer study in the paper. In the paper, six dentists each judged 100 real and 100 PanoDiff-SR radiographs drawn at random, so that the comparison was fair; they reached 68.5%. Here the synthetic half is the hardest material we could find among the generators we trained — GANs and diffusion models alike — and the real half comes from the public DENTEX dataset. Every image, real or synthetic, passed the same sharpness check at 1024 × 512. The two numbers are not comparable. How the images were chosen.

Real or synthetic?

Live Keyboard: R or ← for real, S or → for synthetic, F for full screen.
Test length
Enter a name to start.
How it works
  • Each radiograph is shown for 12 seconds, as in the study. Then it disappears and you answer from memory.
  • Two choices only: real or synthetic. No “unsure” — if in doubt, go with your gut.
  • Exactly half real, half synthetic, in random order, with the same mix of generators at every test length.
  • 100- and 200-image tests have a break every 50 images. Breaks do not count towards your time.
  • Progress cannot be saved. Please finish in one sitting; leaving the page ends the test.
  • No feedback until the end, as in the study. Then you get your score, your rank and every image with its answer.

Your name, score and answers are stored (Google Firebase) and used to study how people judge synthetic radiographs. Your name and score appear publicly on the leaderboard below.

Leaderboard

connecting… One ranking for all test lengths. Updates live.
#NameScoreCorrectFakes caughtReals trustedMedian timeWhen
Loading the leaderboard…
Remove an entry

Enter the name and the delete code you used. Entries made in this browser without a code can also be removed with the button in their row.

How the images were chosen

The pool holds 300 radiographs: 150 real and 150 synthetic. Each test draws its images from it, half and half, keeping the proportions below at every length.

Quality first. Every image is a native 1024 × 512 radiograph that is as sharp as a typical good DENTEX radiograph and no grainier than one. We measure sharpness as the share of the image's contrast that is finer than a 512 × 256 image could hold, and noise with Immerkær's (1996) estimator; the thresholds are the 40th percentile of sharpness and the 95th percentile of noise among DENTEX radiographs, and they apply to the real and the synthetic half alike. Three of our generators never reach them — PanoDiff with SwinIR, FastGAN with SwinIR and FastGAN trained at full size produce visibly soft images — so they are not in the challenge.

Synthetic. We generated 7243 radiographs with each of the remaining seven models and, among the images that pass the quality check, looked for the ones that are hardest to catch. For every image we measured (1) how sure a vision transformer, trained to tell that model's images from real ones, was that it is synthetic, scored on images the transformer never trained on; and (2) the realism score of Kynkäänniemi et al. (2019), which is how deep the image sits inside the cloud of real radiographs in Inception feature space. The images that ranked highest on both went in. We removed any that were nearly identical to a training radiograph (that would be a copy, not a synthesis) and any with burned-in labels or markers, and kept only images whose closest real neighbours come from DENTEX, so that the synthetic half looks like the same kind of radiograph as the real half. Three PanoDiff-SR images from the paper's study that two of the six dentists called real are included; the study images that fooled more dentists are not sharp enough to pass the quality check.

Real. All real images come from DENTEX (Hamamci et al., 2023; CC BY-NC-SA 4.0), processed exactly as the training data were. Seven are DENTEX radiographs from the paper's study that half or more of the dentists thought were synthetic, 60 are the ones our transformers found most synthetic-looking, and the rest are an ordinary random draw from the DENTEX radiographs that pass the quality check.

SourceFamilyIn the poolWhat it is

The paper's observer study, for reference

68.5%

Mean accuracy of six dentists over 200 images, with unsure responses weighted as half a correct and half an incorrect decision.

κ = 0.18–0.43

Pairwise agreement between observers, which is fair at best. They were not converging on the same tell-tale features.

157

Judgements the observers refused to commit on. They are the hardest images, and how you count them is the one thing that moves the headline number.

2Why it is built in two stages

A panoramic radiograph is a wide, thin image of the whole jaw. Generating one directly at diagnostic resolution with a diffusion model is possible and expensive. PanoDiff-SR does not.

Machine learning in dentistry is short of data. Public panoramic radiograph (PR) datasets are few, each is tied to one device and one patient population, and none of them contains many rare findings. Synthetic radiographs are a way out of that: pooled with real ones they can balance an under-represented class, and on their own they make teaching material that carries no patient inside it.

Until now that work has been almost entirely GAN-based, and three problems keep recurring. The field of view is usually cropped to the dentoalveolar region rather than the complete arch. Observer studies still catch the synthetic images at around 78–82%, which means visible artefacts survive. And adversarial training remains fragile — mode collapse and instability are reported repeatedly in medical image synthesis.

Diffusion answers the third problem directly. It does not answer the second one for free, and it makes the first one expensive: the iterative denoising loop has to run at whatever resolution you want out, and at 1024 × 512 that cost is prohibitive on one GPU. The pipeline therefore splits the two jobs. Generation happens small; resolution is added afterwards in a single pass.

The pipeline, end to end

Paper §3, Table 2 Click a stage.
Diagram of the selected pipeline stage

    What each stage costs

    The split is only worth making if it is cheaper, so the paper measures it. Both stages were profiled back to back on the same single 47 GB GPU.

    Computational cost of the two stages

    Paper Table 2
    PanoDiff (LR synthesis)SR — HAT 4× (LR → HR)

    * wall-clock from the training logs of the runs reported in the paper. † from a dedicated profiling run measuring both stages back to back. LR: 256 × 128. HR: 1024 × 512.

    250 : 1

    Network evaluations per image. The generator runs its denoiser 250 times; the super-resolution network runs once.

    7.15 s

    End-to-end inference per radiograph, of which 6.84 s is the diffusion loop at 256 × 128 and 0.31 s is the 4× upscale.

    1 GPU

    Everything in this paper — both trainings and all 7243 generated radiographs — ran on one RTX 6000 Ada.

    3Five datasets, five devices, one corpus

    7243 radiograph files — 5653 distinct radiographs — pooled from five public sources on four continents, cropped by one fixed margin and resized to a common 1024 × 512.

    7,243files after preparation: 5,653 distinct radiographs
    paper §3.1
    5source datasets, five different devices
    paper Table 1
    5countries: China, Switzerland, DR Congo, USA, Brazil
    paper Table 1
    1024×512common resolution every image is resized to
    process_data.py

    The five sources

    Paper Table 1 Click a row for what it contributes.
    SourceImagesCountryNative resolution

    The counts are the radiograph files used from each release. For DENTEX these are its 3603 training images, spread over four label subsets that share radiographs; the 598 TSXK radiographs entered the corpus twice; and USPFORP repeats 72 of its 945 files. The 7243 files are therefore 5653 distinct radiographs. Every number on this page that involves the real set was computed on the 7243 files, as in the paper.

    Pooling five devices is not free. Each machine has its own exposure, geometry and post-processing, and a t-SNE of ResNet-50 features shows the five sources occupying distinguishable regions of feature space. The model therefore learns a mixture of device appearances, and every distribution metric in this paper is computed against that mixture.

    t-SNE of 500 random images from each of the five source datasets, forming distinguishable clusters
    Five devices, five neighbourhoods. t-SNE of 500 random images per source dataset. The clusters separate, which is exactly why the paper later stratifies its FID by device (§10) rather than reporting one pooled number and stopping.

    What preparation does to an image

    process_data.py

    One step, applied to everything: a fixed margin is cropped from every source radiograph to remove letterboxing and burned-in markers, then a Lanczos resize to 1024 × 512. The margin is the same for all five sources, and it matters beyond training: any real radiograph compared with the synthetic ones has to be framed the same way, which is why the held-out restoration test (§5) and the MedSAM probe (§9) apply exactly this step to their real images.

    crop_top_left_bottom_right = (64, 127, 90, 127)   # every source, in source pixels
    w, h = img.size
    img.crop((127, 64, w - 127, h - 90)).resize((1024, 512), LANCZOS)

    4Stage one — PanoDiff

    A 34-million-parameter U-Net, trained to strip noise out of a radiograph, then run backwards 250 times from pure noise until a jaw appears.

    Diffusion has two halves. The forward half needs no learning at all: take a real radiograph and add Gaussian noise to it on a fixed schedule until, after 1000 steps, nothing is left but noise. The reverse half is the model: a U-Net trained to look at a noisy image and say how much of it is noise. Run that network repeatedly, subtracting a little of its prediction each time, and you walk backwards from noise into an image.

    The forward process, on a real radiograph

    Illustration Drag the timestep. The schedule is the cosine one the model was trained with.
    t = 0t = 250t = 500t = 750t = 1000
    0
    Signal retained
    √̄ᾱt, the weight on the image1.000
    √̄1−ᾱt, the weight on the noise0.000

    x_t = √(ᾱ_t) · x_0 + √(1 − ᾱ_t) · ε, ε ~ N(0, I)

    The two schedules are drawn to their real definitions, so the switch shows a genuine difference: the cosine schedule, which PanoDiff was trained with, destroys the image later and more gently.

    The three phases, as the paper draws them

    Network

    U-Net with hidden dims [64, 128, 256, 512] and self-attention, 34.0 M trainable parameters.

    Downsampling

    Only two spatial halvings, by pixel-unshuffle rather than strided convolution, so the coarsest feature map is 32 × 64 and the reduction moves information into channels instead of discarding it.

    Training

    L1 objective, 1000 training timesteps, cosine β-schedule, EMA weights, 110 epochs, seed 42.

    Sampling

    DDIM with 250 inference steps — far fewer than ancestral DDPM sampling would need for the same quality.

    Diagram of the reverse phase: a trainable U-Net predicting noise, contrasted with step-by-step DDPM denoising
    The reverse phase. Top: the U-Net (flame — trainable) sees a pair (x0, xt=k) and is trained with an L1 loss to predict the noise in xt=k. Bottom, for contrast: step-by-step DDPM denoising with the same network frozen (snowflake), for which a single-step estimate is not adequate at high noise levels. Paper Figure 5.

    Watch it learn

    Because the sampling seed is fixed, the same noise vector can be decoded by each checkpoint in turn, so what changes between columns is the generator alone. The striking thing is how early the coarse structure arrives: by epoch 11 the model already draws a recognisable full arch. What the later epochs add is consistency and detail — firmer crown outlines, roots that separate from the trabecular background, and even contrast across the image.

    The same five seeds, across training

    Checkpoints ep11 → ep110 Drag the epoch, or press Play. Click any radiograph to enlarge.
    Epoch 11
    11
    All five columns are upscaled by the same super-resolution model, so what changes between epochs is the generator alone.

    What the diffusion stage is actually sensitive to

    Two things were varied on the trained checkpoint: whether sampling uses the averaged (EMA) weights or the raw ones, and how many DDIM steps the sampler takes. FID here is computed on 1000 generated against 1000 real low-resolution radiographs, so these numbers are comparable with each other but not with the full-corpus tables further down.

    Diffusion ablation

    Paper §3.4 Hover a point. The axes trade quality against sampling cost.
    Weight averaging DDIM sampler budget

    Weight averaging is doing a great deal of work. Reading the same checkpoint from raw instead of EMA weights raises FID from 67.8 to 122.5. The Inception score moves the other way (2.65 → 2.78), a clean reminder that IS rewards confident, varied ImageNet predictions without reference to the real distribution at all.

    The sampler budget is the cheap lever: 100 steps costs 5 FID points and runs 2.5× faster, 50 steps stays within 15 points at 4.9× faster, and only at 25 steps does quality fall away.

    Diffusion against a GAN, at the resolution both were trained at

    The comparison that matters for stage one is at 256 × 128, before super-resolution confuses things. FastGAN was trained on the same corpus.

    Low-resolution samples, side by side

    Paper Figure 18 Click any sample to enlarge.

    PanoDiff — FID 55.0, IS 2.98

    FastGAN — FID 85.2, IS 2.25

    FID is against the same 7243 real low-resolution radiographs; real data itself scores IS 2.90. Precision and recall say where the difference lies: FastGAN's samples reach a somewhat higher precision (0.43 against 0.28), but they cover almost none of the real distribution (recall 0.015 against 0.22) — the limited diversity adversarial training is known for, measured rather than assumed (§6). The two models were trained under their own recommended settings, so this does not equalise parameter count, training iterations or compute; the controlled version of that question is in §10.

    5Stage two — super-resolution

    A hybrid attention transformer, pre-trained on natural images and fine-tuned for 400,000 iterations on radiograph pairs, takes the 256 × 128 seed to 1024 × 512 in a single forward pass.

    The seed the diffusion model produces is a complete panoramic radiograph in composition but useless in detail: at 256 × 128 a molar is a smudge a dozen pixels across. HAT (hybrid attention transformer) combines channel attention with overlapping window self-attention, so it can use both local texture and long-range context; here it starts from released natural-image weights and is fine-tuned on PR pairs with pixel, perceptual and adversarial terms.

    The seed and the result

    Generated pairs Drag across the image. Left is the 256×128 seed, bicubic-upscaled for display.
    The same radiograph after HAT super-resolution
    The 256 by 128 diffusion seed, bicubic-upscaled
    LR seed, bicubicHAT-SR, 1024×512
    ↔

    Both halves are the same generated radiograph. The left is what stage one emitted; the right is what stage two made of it.

    What the upscale adds

    Enamel edges, the periodontal ligament space, trabecular texture in the ramus and the outlines of restorations. What it cannot add is anything the seed did not imply — the anatomy is decided at 256 × 128.

    Does fine-tuning earn its 53 hours?

    Measured honestly, on 300 real radiographs downsampled 4× and restored, the answer depends entirely on which metric you believe — and this is the clearest case on the page of a metric disagreeing with the eye.

    Super-resolution ablation

    Full sweep behind paper §3.4 Switch the metric. Bars are scaled within each metric.

    Bicubic upsampling wins on PSNR and SSIM and is visibly the worst image on the page. PSNR and SSIM reward a smooth, conservative prediction, and blur is exactly that. LPIPS, which compares deep features rather than pixels, ranks the same rows in the opposite order — and it is LPIPS that agrees with what the crops below look like. Fine-tuning takes LPIPS from 0.274 to 0.116; the last of that comes from weight averaging.

    Two magnified crops of one radiograph restored five ways: bicubic, stock HAT, fine-tuned without EMA, fine-tuned with EMA, and ground truth
    The same pixels, five ways. Left to right: bicubic upsampling, the released HAT model with no fine-tuning, our fine-tuning without EMA, our fine-tuning with EMA, and the ground truth. All crops are the same region at the same magnification. From the super-resolution ablation.

    HAT against SwinIR

    Both upscalers start from weights pre-trained on natural images and were fine-tuned on the same 7243 radiograph pairs. On generated seeds HAT-SR brings both generators closer to the real radiographs than SwinIR does (§6), and the reason is visible in the arm viewer below: the fine-tuned SwinIR returns something close to a smooth enlargement of its input. Measured with a Laplacian filter over all 7243 images, PanoDiff's samples keep about 4% of the high-frequency energy of real radiographs after SwinIR, against about 80% after HAT-SR.

    Restoring real radiographs shows where that comes from. The test uses the 185 radiographs of the DENTEX validation and test sets that are not in the training corpus (the other 115 of those 300 are pixel-identical to DENTEX training images), prepared exactly as the corpus was, downsampled fourfold and restored. With their released natural-image weights the two networks are perceptually close (LPIPS 0.310 against 0.291). Fine-tuning separates them in opposite directions: SwinIR becomes the best of the four on PSNR and SSIM, HAT-SR the best on LPIPS, and each ordering holds on all 185 images.

    Restoring 185 held-out radiographs

    Paper Table 10
    VariantPSNR ↑SSIM ↑LPIPS ↓

    Winning on the distortion measures while losing on the perceptual one is the perception–distortion trade-off, and it has been reported for dental radiographs too. Two things drive it here. SwinIR was fine-tuned with a pixel loss alone, which favours smooth averages, and HAT-SR with pixel, perceptual and adversarial terms. And SwinIR was fine-tuned on exactly this fixed fourfold downsampling, HAT-SR on varied blur, noise and compression, so the table measures restoration from this one degradation rather than super-resolution in general.

    Four upscaling arms on the same seeds

    Paper Figure 18, Tables 8–9 Pick a row, then switch the arm.
    A synthetic radiograph produced by the selected generation and upscaling arm

    Look closer: one seed, both upscalers

    Paper Figure 18 Hover over any image — the lens follows in all three.
    Low-resolution seed, bicubic-upscaled for display
    Seed 256 × 128, bicubic for display
    The same seed after HAT-SR
    HAT-SR 1024 × 512
    The same seed after SwinIR
    SwinIR 1024 × 512

    Full-resolution images; the first three samples of each generator are the columns of the paper's Figure 18. Both upscalers were fine-tuned on the same 7243 radiographs. Look at enamel edges, the periodontal ligament space and the trabecular bone of the ramus: HAT-SR reconstructs them, SwinIR returns a smooth enlargement of the seed. On a touch screen, drag across an image.

    6How close is the distribution?

    Four sets of 7243 images — real and synthetic, at low and high resolution — and every pairwise comparison the paper could make between them.

    The headline is FID 40.5 between 7243 real and 7243 synthetic high-resolution radiographs. On its own that number means very little, which is why the table below includes the comparisons that give it a scale: what two halves of the same set score against each other (the floor), and what two different real devices score against each other (the ceiling of what "different but both real" looks like).

    Every FID comparison in the paper

    Paper Table 3 Hover a bar. Lower is more similar.
    Self-comparison — the floor Real vs. synthetic Real device vs. real device High vs. low resolution

    Read the rose bar at 40.5 against the green band at 36.8–70.1. The distance from real to synthetic at high resolution sits inside the range that separates two real datasets acquired on different machines. Read it also against the grey floor at 5–6, which is what the same distribution scores against itself: 40.5 is not small in absolute terms.

    t-SNE of real and PanoDiff features at low resolution, left, and high resolution, right
    Super-resolution pulls the two clouds together. t-SNE of ResNet-50 features for real (GT) and PanoDiff (PD) radiographs. At low resolution, left, the clusters are adjacent but distinct and their centroids are 81.41 apart. At high resolution, right, the synthetic cluster sits surrounded by real points and the centroid distance falls to 34.73. Paper Figure 8.

    Inception scores

    Paper Tables 4 and 8 Switch the view.

    Distance, fidelity and coverage for every arm

    Every arm is compared with the real set at its own resolution. FID and the kernel inception distance (KID) say how far apart two sets are; KID is unbiased, so it depends less on the number of images. Precision and recall say where a distance comes from: precision is the share of generated images that lie close to real ones (fidelity), recall the share of real images that lie close to generated ones (coverage). Both generators were trained from scratch on the 7243 real radiographs, and both upscalers fine-tuned on the same 7243, so every arm has seen the same data.

    Every generated set against the real set

    Paper Table 9 Best per resolution in bold.
    Generated setFID ↓KID ×103 ↓Precision ↑Recall ↑

    Numbers are the source indices of the paper's Table 8. KID is the mean ± sd over 100 subsets of 1000 images. Precision and recall use the same Inception features with k = 3 nearest neighbours. PanoDiff with HAT-SR is the closest generated arm to the real radiographs on all four measures, each by a wide margin. The last two rows apply the same upscalers to the real low-resolution images, which isolates the upscaling step: both land far closer to the real set than any generated arm, so most of the remaining distance comes from the generator. There the order reverses — with a real input the correct output exists, a pixel-loss network returns a smooth estimate that stays close to it, and a network trained with adversarial terms resynthesises detail that is plausible but not this radiograph's. On generated seeds there is no true detail to stay close to.

    Per-image fidelity and coverage scores for every arm at low and full resolution
    Fidelity and coverage, image by image. Every image has a neighbourhood in the Inception feature space, out to its third-nearest neighbour in its own set. Top row: for each generated image, its distance to the nearest real image in units of that neighbourhood; images left of the dashed line lie inside a real neighbourhood, and their share is the precision. Bottom row: the same for each real image against the generated set; the share inside is the recall. FastGAN's curve sits slightly further inside than PanoDiff's in the top row and almost entirely outside in the bottom row: realistic-looking samples that cover almost none of the real distribution. Paper Figure 17.
    t-SNE of every generative arm in the Inception feature space, in three panels
    Every arm in one feature space. t-SNE of 1500 images per arm in the Inception feature space behind every FID and IS value, with each panel embedded separately. Crosses mark centroids and lines enclose half of each arm's peak density. Distances between clusters in a t-SNE map are not quantitative; measured in the original feature space, the distances to the real images in panel (c) order the arms exactly as their FIDs do. Paper Figure 16.

    What these metrics can and cannot say. FID compares the mean and covariance of Inception features between two sets; it is sensitive to sample size, so numbers are only comparable within a table, which is what KID is there to relieve. Precision and recall split a distance into fidelity and coverage. Inception score never looks at the real data at all — it rewards an ImageNet classifier being confident and varied, which is why a model can improve on IS while drifting away from the distribution it is supposed to be modelling. That happens twice on this page: in the diffusion ablation, and with the SwinIR arms, which score above real radiographs. None of these metrics was designed for radiographs, and none knows what a mandible is. That is what the next sections are for.

    7Six dentists, 200 radiographs, twelve seconds each

    Three early-career dentists with under ten years of practice and three with over fifteen, spanning prosthodontics, oral surgery, orthodontics and oral radiology. Each sat one session of about an hour, alone.

    Each radiograph was displayed for twelve seconds, including about two seconds of loading. For each one the observer moved a slider to one of five positions: definitely real, probably real, unsure, probably fake, definitely fake. The graded scale is deliberate — forcing a binary choice would have thrown away the information that matters most, which is how often an expert simply could not tell.

    How each observer used the scale

    Raw responses, 6 × 200 Click an observer. Bars are the raw counts, no weighting applied.

    Derived metrics, certainty-weighted

    Bars read left to right: definitely real, probably real, unsure, probably fake, definitely fake. The counts here are computed from the study's raw submission log and reproduce the paper's Table 5 exactly.

    The one convention that moves the number

    A five-level scale needs a rule for what an unsure response is worth, and it is fair to ask whether that rule is carrying the headline result. The paper reports all four candidate rules.

    Mean accuracy under four conventions

    Paper §4.1 Click a convention.

    Three of the four agree to within 3.4 points, so the reported accuracy is a property of the responses rather than of the weighting. The fourth — discarding unsure responses — lifts the mean to 78.5% and lands it exactly on the 78.2% GAN benchmark the paper's central claim rests on being below. It does that not because the radiographs became easier to detect but because the 157 hardest judgements were removed from the denominator.

    Better than guessing, and lower than for GAN radiographs. Treating each reader as one observation, the six accuracies give a 95% confidence interval of 62.6–74.3%, which excludes 50% (one-sample t-test, 8.15 standard errors above chance, p = 0.0005; exact Wilcoxon p = 0.031, the smallest attainable with six readers), and every reader's ROC area exceeds 0.5 (DeLong, all p < 10−11). So the images are not undetectable. The same six accuracies are lower than the 78.2% sensitivity reported for StyleGAN2-ADA radiographs (p = 0.008) and the 82.3% accuracy for StyleGAN3 radiographs with jaw cysts (p = 0.002). Neither study replicates this protocol, so these are comparisons with published reference points rather than a head-to-head. Paper §4.1.

    ROC and precision–recall, per observer

    Computed from the raw responses Hover a curve. Toggle a group in the legend.

    They did not agree with each other either

    Weighted Cohen's κ between every pair

    Paper Table 6 Quadratic weights, so a near-miss is penalised less than a reversal.

    Values run from 0.18 to 0.43, which is fair at best on the Landis–Koch scale, and there is no meaningful difference between the early-career and experienced groups. In a real-versus-synthetic task this is the favourable direction: high agreement would mean the observers were all spotting the same artefact. Low agreement, together with 68.5% accuracy, says a sizeable fraction of the synthetic radiographs carried no consistent tell at all. The highest pair, EC3 and EP3 at 0.43, is also the pair whose response distributions look most alike.

    What each kind of mistake looked like

    Paper Figure 12 Pick a cell of the confusion matrix, then a certainty level.

    Examples were chosen for each category by the majority of the observers' assessments, not by eye. Click either image to enlarge.

    Six pie charts, one per observer, splitting decisions into correct and incorrect at full and partial certainty
    Every decision, per observer. Correct and incorrect decisions split by whether the observer was fully or partially certain. The size of the incorrect-but-fully-certain wedge is the interesting part. Paper Figure 9.

    8What a machine sees that a dentist does not

    A vision transformer trained to separate real from synthetic reaches 97.5% test accuracy. The humans reached 68.5%. Where it looks, measured over the whole held-out set, says what kind of difference it has found.

    A ViT was trained on 7243 real and 7243 synthetic high-resolution radiographs with an 80:20 split. It separates them almost perfectly, which is worth stating plainly: these images are not indistinguishable, they are hard for a person under time pressure. Attention rollout then shows which patches drove each decision.

    Attention map viewer

    Paper Figure 13 Pick a sample, then slide between the radiograph and its map.
    A radiograph from the held-out set Its attention map
    55%

    Top three rows of the grid are real radiographs, bottom three synthetic. A dot marks the four the classifier got wrong; all four are synthetic radiographs it read as real. C is the classifier's output, 1 for real and 0 for synthetic.

    More diffuse, not misplaced. On the thirty samples above, attention on synthetic radiographs looks more diffuse than on real ones. Measured over all 2898 held-out radiographs, that holds: the entropy of the map is higher for synthetic images (6.100 against 6.044 nats), by far the largest effect in the analysis. The share of attention on anatomy, by contrast, is nearly the same for both classes. What separates them is how concentrated the attention is, not where it falls — consistent with a detector that checks whether the expected fine-scale texture is present, which in PanoDiff-SR is reconstructed by the super-resolution stage.

    Attention concentration over the full held-out set

    Paper §4.3, Figure 14 1449 real and 1449 synthetic held-out radiographs.

    Each map is scored against fixed anatomical templates — the 1000 expert teeth and jaw masks of the TUFTS dataset, averaged and thresholded at 0.5 — so the regions are identical for every image. With 1449 images per group every difference is significant; the effect sizes (rank-biserial) are the quantity to read, and only entropy (+0.607) is large.

    How the attention statistics are computed: a radiograph, its attention map, the anatomical templates, and the distributions over the held-out set
    How the attention statistics are computed. A held-out radiograph (1), the attention-rollout map the real-versus-synthetic ViT produces for it (2), and the fixed anatomical templates (3); three numbers follow for every image. Bottom (4): their distributions over the 1449 real and 1449 synthetic held-out radiographs. The entropy separates the two classes; the anatomical shares barely do. Paper Figure 14.

    9A segmentation model that has never seen the data

    The classifier above was trained to find differences. MedSAM was not. Prompted identically on real and synthetic radiographs, it segments synthetic teeth exactly as it segments real ones.

    MedSAM is a general medical segmentation model, fine-tuned from SAM on about 1.5 million image–mask pairs from other modalities and organs. Given a box around an object it returns a mask and a score for how good it believes that mask to be. If synthetic teeth look like real teeth to such a model, it should segment them with the same confidence and accuracy. It was run on all 7243 synthetic radiographs and on the 2697 real radiographs that have a manual tooth mask: 500 ADLD, 634 DENTEX, 598 TSXK and 1000 TUFTS, less 35 with an empty mask. The real radiographs were prepared exactly as the training corpus was, so both sets are framed alike — straight from the source files, their dentition would cover a quarter less of the image, and every measure would read that as anatomy. Every image received the same three kinds of box prompt: a fixed box at the same position in every image, a dentition box around the tooth region, and one box per tooth region.

    MedSAM under identical prompts

    Paper Table 7 2697 real and 7243 synthetic radiographs.
    PromptMeasureRealSyntheticrpDevice vs. others |r|

    r is the effect size of the real–synthetic difference (rank-biserial; r > 0 means higher on real; |r| < 0.1 negligible, 0.1–0.3 small, above 0.3 medium) and p the two-sided Mann–Whitney value. The last column gives, for scale, the range of |r| when each real device is compared with the other three on the same measure. Dice is the overlap with the tooth region the prompt was built from; against the manual masks MedSAM reaches 0.83 with tooth-region prompts and 0.78 with the dentition box.

    MedSAM masks on twelve real and twelve synthetic radiographs with tooth-region prompts
    Same prompts, same model, both sets. MedSAM with tooth-region prompts on real radiographs, three from each labelled device, and on PanoDiff-SR radiographs. Yellow boxes are the prompts, cyan the union of MedSAM's masks; the badge is its mean predicted mask quality. Paper Figure 15.

    MedSAM cannot tell them apart. Under all three prompts its predicted quality is almost the same on both sets — 0.464 against 0.466 with the fixed box, 0.480 against 0.472 with the dentition box, 0.505 against 0.507 per tooth region — every difference negligible (|r| ≤ 0.08), and not even in a consistent direction. The real devices differ from each other by more than that. The one effect above negligible is that synthetic dentitions break into slightly more regions (4.2 against 3.4 per image), still inside the range of the real devices. A classifier trained on the two classes separates them at 97.5% (§8); a general segmentation model does not, so what separates them is a signature a detector can learn, not a property such a model relies on. The reference regions come from the authors' own U-Net, used identically on both sets, and tooth regions are connected regions rather than individual teeth.

    10Is the two-stage split actually the right call?

    Six models trained from scratch at full resolution on the same 7243 radiographs under a comparable budget, to test the premise the whole pipeline rests on. The answer is not a clean win.

    The main purpose of PanoDiff-SR is efficient synthesis, which is why the rest of the paper compares it with FastGAN, a GAN built for low training cost that fits the same small-generator-plus-upscaler setting. Generating directly at 1024 × 512 is the control that has to be run once: StyleGAN2-ADA, StyleGAN3, FastGAN-HR, ADM, LDM and PanoDiff-HR — the same model as PanoDiff but trained at full resolution — all on the same data and with comparable iterations and training budget.

    Direct high-resolution synthesis against the two-stage pipeline

    Paper Table 11 Hover a point. Down and left is better on both axes.
    Diffusion GAN Two-stage (ours) Ring size ∝ training cost
    ModelFamilyFID ↓IS ↑Precision % ↑Recall % ↑ViT detect % ↓Cost (acc.-h)

    † The two-stage rows were timed per training step on the same accelerator as the other rows and multiplied by their number of steps; the SR stage is 61.5 of these hours. Precision (fidelity) and recall (coverage) are computed image by image as in Paper Figure 17, with k = 3.

    Four things this table settles, one of them against the paper.

    • The split works. PanoDiff-HR — the identical model trained directly at full resolution — scores FID 155.4 against the two-stage pipeline's 40.5, and costs more to train (120 against 82 accelerator-hours). Generating small and upscaling is not a compromise here, it is the thing that makes the model work at all.
    • The two StyleGANs beat it on FID. 23.5 and 19.4 against 40.5, and the real-versus-synthetic classifier finds them harder to spot (89.6% and 77.9% against 97.6%). They are also the most expensive models here, at 441 and 572 accelerator-hours against 82.
    • Their lead is fidelity, not coverage. The StyleGANs put more of their images close to real radiographs (precision 61.1% and 67.3% against 56.2%), but PanoDiff-SR covers more of the real radiographs than any other model: recall 28.8% against 21.9% and 25.5% (paired McNemar tests on the same 7243 real images, p < 10−6; the lead holds in 10 of 10 resamplings against StyleGAN2-ADA and 8 of 10 against StyleGAN3).
    • No model memorised its training set. For every model, the median distance from a generated image to its nearest training radiograph exceeds the distance between a real radiograph and its closest real neighbour (0.093; StyleGANs 0.105 and 0.102).

    Why the six dentists were not shown the GAN arms. Every published expert evaluation of GAN-generated panoramic radiographs we found reports that clinicians could tell them apart or rated them unrealistic: 78.2% sensitivity for StyleGAN2-ADA PRs (Schoenhof et al., 2024), 82.3% for StyleGAN3 PRs with jaw cysts (Fukuda et al., 2024), 81.6% mean accuracy by four dental radiologists on conditional-GAN PRs, one of them without error (Fu et al., 2026), and realism scores of 2.44–2.72 out of five (Pedersen et al., 2025). Where GANs and diffusion were compared directly, the GAN gave poor results and the diffusion images were hard to tell from real ones (Adnan et al., 2026; Kirkwood et al., 2025). The paper therefore reads the GANs' lower FID as a distributional match, not as clinical realism.

    Does it look like a jaw?

    FID and IS have no idea what a mandible is, so four global anatomical measures were computed over 1500 images per model: bilateral symmetry, the periodicity of the tooth row, the number of distinct enamel crowns, and occlusal separation.

    Anatomical plausibility

    Paper Table 12 The real column is the target, not a winner.

    * differs from real at p < 0.05, ** at p < 10−4 (Mann–Whitney U). The last row is the fraction of images falling outside the central 98% of real radiographs on any of the four measures, so the real column's 7.3% is the floor that statistic can reach. LDM's 53.9% and its 26.1 enamel crowns are the signature of the repeated tooth rows visible in the figure below. These measures reject degenerate output reliably; they are global, and they say nothing about crown-to-root proportion or root morphology.

    Random samples from each direct high-resolution baseline alongside real radiographs and the two-stage pipeline
    One row per model. Examples from each direct high-resolution model, with real radiographs and the two-stage pipeline for reference. The PanoDiff-SR panels are synthetic radiographs that a majority of the six dentists judged real; the other rows show the examples that best meet one anatomical criterion applied to every row. PanoDiff-HR, which differs from PanoDiff-SR only in generating at full resolution, is blurred and sparse, and the latent model repeats tooth rows. Paper Figure 19.

    Is the model closer to one device than another?

    The corpus pools five machines, so a fair worry is that the model has simply learned the largest source. Each device was therefore compared with the synthetic set and with the pooled real corpus at a matched n = 500, redrawn five times.

    Device-stratified FID

    Paper §4.5.3 Hover a pair of bars.
    to the synthetic set to the pooled real corpus

    Each device set was cropped and resized exactly as the training corpus was. The synthetic set is 71.5 (TUFTS) to 102.0 (TSXK) from the five devices, a spread of 1.43×, while the devices are 61.7 to 117.0 from one another, a spread of 1.90×, so the generator does not favour one device beyond how much the devices already differ. The device it matches least, TSXK, is also the one most distinct from the other four (mean FID 97.1, against 74.2–95.2 for the rest); five devices are too few to make that relation significant (Pearson r = 0.76, p = 0.13). For scale, two random halves of the real corpus give 37.8 at n = 500.

    11What comes next

    The next steps the paper sets out, each with the question it would answer.

    Search the design choices

    U-Net size, SR loss weights, number of diffusion timesteps and the dataset mix were taken from the source implementations. A systematic search over such choices has improved diffusion models for natural images (Karras et al., 2022).

    Separate architecture from objective

    HAT and SwinIR were fine-tuned with different losses and degradations. Fine-tuning SwinIR with HAT's recipe would show how much of the difference is the architecture.

    Harmonise the devices

    The five sources occupy different regions of feature space. Harmonising them, or conditioning the generator on the device, could remove this source of variation (Guan & Liu, 2022).

    A larger reader panel

    Each of the six observers read every image once. More readers, with repeated readings, would measure within-reader variation and test the effect of clinical experience (Obuchowski & Bullen, 2022).

    Patch-level memorisation

    The memorisation test compares whole images. Diffusion models can copy training images in part (Carlini et al., 2023; Somepalli et al., 2023), so a patch-level test should precede any public release.

    Synthesise lesions

    Conditioning the generator on caries or fractures would supply rare cases for detection models, as synthetic lesions have done in other imaging tasks (Frid-Adar et al., 2018).

    The claim is not that these radiographs are indistinguishable from real ones. A classifier separates them at 97.5%. The claim is that six clinicians, reading at the speed clinicians read, reached 68.5% — and that this was done on one GPU.

    12Code, models and citation

    Both stages, the six baselines of Table 11, every analysis script of the paper, the observer-study application the dentists actually used, and the checkpoints behind every figure on this page.

    Repository map

    github.com/s4nyam/PanoDiff Pick a directory.

      
                

      Diffusion checkpoints

      Six epochs of PanoDiff — 11, 33, 55, 77, 99 and 110 — the six columns of the training progression above.
      archive.org/download/panodiff/trained_models

      Super-resolution weights

      Real_HAT_GAN_SRx4_finetuned, the 400k-iteration EMA model used for every high-resolution image on this page.
      archive.org/download/panodiff/experiments

      Source datasets

      All five are public or available on request. Links and terms are in §3; none of them is redistributed here.

      Cite this work

      @misc{jain2025panodiffsr,
        title         = {PanoDiff-SR: Synthesizing Dental Panoramic Radiographs using Diffusion and Super-resolution},
        author        = {Jain, Sanyam and Neves de Freitas, Bruna and Basse-O'Connor, Andreas
                         and Iosifidis, Alexandros and Pauwels, Ruben},
        year          = {2025},
        eprint        = {2507.09227},
        archivePrefix = {arXiv},
        primaryClass  = {eess.IV},
        url           = {https://arxiv.org/abs/2507.09227}
      }
      Copied Preprint: arXiv:2507.09227