Synthesizing dental panoramic radiographs using diffusion and super-resolution
1Department of Dentistry and Oral Health, Aarhus University, Denmark 2Aarhus Institute of Advanced Studies, Aarhus University, Denmark 3Department of Mathematics, Aarhus University, Denmark 4Faculty of Information Technology and Communication Sciences, Tampere University, Finland
Every radiograph above is synthetic. None of them is a real patient. Each was shown to six dentists for twelve seconds, and the number under each one is how many of the six called it real.
Real radiograph or synthetic? Twelve seconds per image, two buttons, no second look. Choose 50, 100 or 200 images, put your name on the live leaderboard, and see which generators fooled you.
Your name, score and answers are stored (Google Firebase) and used to study how people judge synthetic radiographs. Your name and score appear publicly on the leaderboard below.
| R ← | Real |
| S → | Synthetic |
| F | Full screen on / off |
| Z | Magnifier on / off (move the pointer over the image) |
| + − | Magnification |
| I | Invert |
| 0 | Reset brightness, contrast and invert |
| ? | This help |
Take a pause: look away from the screen, stretch, blink. The clock is stopped.
| Source | Shown | You said real |
|---|
“You said real” on a synthetic row means the generator fooled you; on the DENTEX row it means you trusted a real radiograph.
How the rank is calculated. A single score ranks every run, whatever its length:
score = 100 × (correct + 5) / (images + 10)
It is your accuracy after adding ten imaginary answers at chance level, five right and five wrong. In statistical terms it is the Bayesian estimate of your true accuracy under a Beta(5, 5) prior centred on guessing. A short test is pulled more towards 50% than a long one, because a lucky streak is more likely over 50 images than over 200; with more images the score converges on your plain accuracy.
| Run | Accuracy | Score |
|---|---|---|
| 45 of 50 | 90.0% | (45 + 5) / 60 = 83.3 |
| 90 of 100 | 90.0% | (90 + 5) / 110 = 86.4 |
| 180 of 200 | 90.0% | (180 + 5) / 210 = 88.1 |
| 48 of 50 | 96.0% | (48 + 5) / 60 = 88.3 |
| 25 of 50 | 50.0% | (25 + 5) / 60 = 50.0 |
Every test is fair to compare: all lengths draw from the same pool of 300 images, always half real and half synthetic, and always with the same share of each generator and of the harder and easier real radiographs. Ties are broken by the longer test, then by the faster median decision time, then by who finished first. Each finished run is a separate entry.
| # | Name | Score | Correct | Fakes caught | Reals trusted | Median time | When | |
|---|---|---|---|---|---|---|---|---|
| Loading the leaderboard… | ||||||||
Enter the name and the delete code you used. Entries made in this browser without a code can also be removed with the button in their row.
The pool holds 300 radiographs: 150 real and 150 synthetic. Each test draws its images from it, half and half, keeping the proportions below at every length.
Quality first. Every image is a native 1024 × 512 radiograph that is as sharp as a typical good DENTEX radiograph and no grainier than one. We measure sharpness as the share of the image's contrast that is finer than a 512 × 256 image could hold, and noise with Immerkær's (1996) estimator; the thresholds are the 40th percentile of sharpness and the 95th percentile of noise among DENTEX radiographs, and they apply to the real and the synthetic half alike. Three of our generators never reach them — PanoDiff with SwinIR, FastGAN with SwinIR and FastGAN trained at full size produce visibly soft images — so they are not in the challenge.
Synthetic. We generated 7243 radiographs with each of the remaining seven models and, among the images that pass the quality check, looked for the ones that are hardest to catch. For every image we measured (1) how sure a vision transformer, trained to tell that model's images from real ones, was that it is synthetic, scored on images the transformer never trained on; and (2) the realism score of Kynkäänniemi et al. (2019), which is how deep the image sits inside the cloud of real radiographs in Inception feature space. The images that ranked highest on both went in. We removed any that were nearly identical to a training radiograph (that would be a copy, not a synthesis) and any with burned-in labels or markers, and kept only images whose closest real neighbours come from DENTEX, so that the synthetic half looks like the same kind of radiograph as the real half. Three PanoDiff-SR images from the paper's study that two of the six dentists called real are included; the study images that fooled more dentists are not sharp enough to pass the quality check.
Real. All real images come from DENTEX (Hamamci et al., 2023; CC BY-NC-SA 4.0), processed exactly as the training data were. Seven are DENTEX radiographs from the paper's study that half or more of the dentists thought were synthetic, 60 are the ones our transformers found most synthetic-looking, and the rest are an ordinary random draw from the DENTEX radiographs that pass the quality check.
| Source | Family | In the pool | What it is |
|---|
Mean accuracy of six dentists over 200 images, with unsure responses weighted as half a correct and half an incorrect decision.
Pairwise agreement between observers, which is fair at best. They were not converging on the same tell-tale features.
Judgements the observers refused to commit on. They are the hardest images, and how you count them is the one thing that moves the headline number.
A panoramic radiograph is a wide, thin image of the whole jaw. Generating one directly at diagnostic resolution with a diffusion model is possible and expensive. PanoDiff-SR does not.
Machine learning in dentistry is short of data. Public panoramic radiograph (PR) datasets are few, each is tied to one device and one patient population, and none of them contains many rare findings. Synthetic radiographs are a way out of that: pooled with real ones they can balance an under-represented class, and on their own they make teaching material that carries no patient inside it.
Until now that work has been almost entirely GAN-based, and three problems keep recurring. The field of view is usually cropped to the dentoalveolar region rather than the complete arch. Observer studies still catch the synthetic images at around 78–82%, which means visible artefacts survive. And adversarial training remains fragile — mode collapse and instability are reported repeatedly in medical image synthesis.
Diffusion answers the third problem directly. It does not answer the second one for free, and it makes the first one expensive: the iterative denoising loop has to run at whatever resolution you want out, and at 1024 × 512 that cost is prohibitive on one GPU. The pipeline therefore splits the two jobs. Generation happens small; resolution is added afterwards in a single pass.
The split is only worth making if it is cheaper, so the paper measures it. Both stages were profiled back to back on the same single 47 GB GPU.
| PanoDiff (LR synthesis) | SR — HAT 4× (LR → HR) |
|---|
* wall-clock from the training logs of the runs reported in the paper. † from a dedicated profiling run measuring both stages back to back. LR: 256 × 128. HR: 1024 × 512.
Network evaluations per image. The generator runs its denoiser 250 times; the super-resolution network runs once.
End-to-end inference per radiograph, of which 6.84 s is the diffusion loop at 256 × 128 and 0.31 s is the 4× upscale.
Everything in this paper — both trainings and all 7243 generated radiographs — ran on one RTX 6000 Ada.
7243 radiograph files — 5653 distinct radiographs — pooled from five public sources on four continents, cropped by one fixed margin and resized to a common 1024 × 512.
| Source | Images | Country | Native resolution |
|---|
The counts are the radiograph files used from each release. For DENTEX these are its 3603 training images, spread over four label subsets that share radiographs; the 598 TSXK radiographs entered the corpus twice; and USPFORP repeats 72 of its 945 files. The 7243 files are therefore 5653 distinct radiographs. Every number on this page that involves the real set was computed on the 7243 files, as in the paper.
Pooling five devices is not free. Each machine has its own exposure, geometry and post-processing, and a t-SNE of ResNet-50 features shows the five sources occupying distinguishable regions of feature space. The model therefore learns a mixture of device appearances, and every distribution metric in this paper is computed against that mixture.
One step, applied to everything: a fixed margin is cropped from every source radiograph to remove letterboxing and burned-in markers, then a Lanczos resize to 1024 × 512. The margin is the same for all five sources, and it matters beyond training: any real radiograph compared with the synthetic ones has to be framed the same way, which is why the held-out restoration test (§5) and the MedSAM probe (§9) apply exactly this step to their real images.
crop_top_left_bottom_right = (64, 127, 90, 127) # every source, in source pixels w, h = img.size img.crop((127, 64, w - 127, h - 90)).resize((1024, 512), LANCZOS)
A 34-million-parameter U-Net, trained to strip noise out of a radiograph, then run backwards 250 times from pure noise until a jaw appears.
Diffusion has two halves. The forward half needs no learning at all: take a real radiograph and add Gaussian noise to it on a fixed schedule until, after 1000 steps, nothing is left but noise. The reverse half is the model: a U-Net trained to look at a noisy image and say how much of it is noise. Run that network repeatedly, subtracting a little of its prediction each time, and you walk backwards from noise into an image.
The two schedules are drawn to their real definitions, so the switch shows a genuine difference: the cosine schedule, which PanoDiff was trained with, destroys the image later and more gently.
U-Net with hidden dims [64, 128, 256, 512] and self-attention, 34.0 M trainable parameters.
Only two spatial halvings, by pixel-unshuffle rather than strided convolution, so the coarsest feature map is 32 × 64 and the reduction moves information into channels instead of discarding it.
L1 objective, 1000 training timesteps, cosine β-schedule, EMA weights, 110 epochs, seed 42.
DDIM with 250 inference steps — far fewer than ancestral DDPM sampling would need for the same quality.
Because the sampling seed is fixed, the same noise vector can be decoded by each checkpoint in turn, so what changes between columns is the generator alone. The striking thing is how early the coarse structure arrives: by epoch 11 the model already draws a recognisable full arch. What the later epochs add is consistency and detail — firmer crown outlines, roots that separate from the trabecular background, and even contrast across the image.
Two things were varied on the trained checkpoint: whether sampling uses the averaged (EMA) weights or the raw ones, and how many DDIM steps the sampler takes. FID here is computed on 1000 generated against 1000 real low-resolution radiographs, so these numbers are comparable with each other but not with the full-corpus tables further down.
Weight averaging is doing a great deal of work. Reading the same checkpoint from raw instead of EMA weights raises FID from 67.8 to 122.5. The Inception score moves the other way (2.65 → 2.78), a clean reminder that IS rewards confident, varied ImageNet predictions without reference to the real distribution at all.
The sampler budget is the cheap lever: 100 steps costs 5 FID points and runs 2.5× faster, 50 steps stays within 15 points at 4.9× faster, and only at 25 steps does quality fall away.
The comparison that matters for stage one is at 256 × 128, before super-resolution confuses things. FastGAN was trained on the same corpus.
FID is against the same 7243 real low-resolution radiographs; real data itself scores IS 2.90. Precision and recall say where the difference lies: FastGAN's samples reach a somewhat higher precision (0.43 against 0.28), but they cover almost none of the real distribution (recall 0.015 against 0.22) — the limited diversity adversarial training is known for, measured rather than assumed (§6). The two models were trained under their own recommended settings, so this does not equalise parameter count, training iterations or compute; the controlled version of that question is in §10.
A hybrid attention transformer, pre-trained on natural images and fine-tuned for 400,000 iterations on radiograph pairs, takes the 256 × 128 seed to 1024 × 512 in a single forward pass.
The seed the diffusion model produces is a complete panoramic radiograph in composition but useless in detail: at 256 × 128 a molar is a smudge a dozen pixels across. HAT (hybrid attention transformer) combines channel attention with overlapping window self-attention, so it can use both local texture and long-range context; here it starts from released natural-image weights and is fine-tuned on PR pairs with pixel, perceptual and adversarial terms.

Both halves are the same generated radiograph. The left is what stage one emitted; the right is what stage two made of it.
Enamel edges, the periodontal ligament space, trabecular texture in the ramus and the outlines of restorations. What it cannot add is anything the seed did not imply — the anatomy is decided at 256 × 128.
Measured honestly, on 300 real radiographs downsampled 4× and restored, the answer depends entirely on which metric you believe — and this is the clearest case on the page of a metric disagreeing with the eye.
Bicubic upsampling wins on PSNR and SSIM and is visibly the worst image on the page. PSNR and SSIM reward a smooth, conservative prediction, and blur is exactly that. LPIPS, which compares deep features rather than pixels, ranks the same rows in the opposite order — and it is LPIPS that agrees with what the crops below look like. Fine-tuning takes LPIPS from 0.274 to 0.116; the last of that comes from weight averaging.
Both upscalers start from weights pre-trained on natural images and were fine-tuned on the same 7243 radiograph pairs. On generated seeds HAT-SR brings both generators closer to the real radiographs than SwinIR does (§6), and the reason is visible in the arm viewer below: the fine-tuned SwinIR returns something close to a smooth enlargement of its input. Measured with a Laplacian filter over all 7243 images, PanoDiff's samples keep about 4% of the high-frequency energy of real radiographs after SwinIR, against about 80% after HAT-SR.
Restoring real radiographs shows where that comes from. The test uses the 185 radiographs of the DENTEX validation and test sets that are not in the training corpus (the other 115 of those 300 are pixel-identical to DENTEX training images), prepared exactly as the corpus was, downsampled fourfold and restored. With their released natural-image weights the two networks are perceptually close (LPIPS 0.310 against 0.291). Fine-tuning separates them in opposite directions: SwinIR becomes the best of the four on PSNR and SSIM, HAT-SR the best on LPIPS, and each ordering holds on all 185 images.
| Variant | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|
Winning on the distortion measures while losing on the perceptual one is the perception–distortion trade-off, and it has been reported for dental radiographs too. Two things drive it here. SwinIR was fine-tuned with a pixel loss alone, which favours smooth averages, and HAT-SR with pixel, perceptual and adversarial terms. And SwinIR was fine-tuned on exactly this fixed fourfold downsampling, HAT-SR on varied blur, noise and compression, so the table measures restoration from this one degradation rather than super-resolution in general.

Full-resolution images; the first three samples of each generator are the columns of the paper's Figure 18. Both upscalers were fine-tuned on the same 7243 radiographs. Look at enamel edges, the periodontal ligament space and the trabecular bone of the ramus: HAT-SR reconstructs them, SwinIR returns a smooth enlargement of the seed. On a touch screen, drag across an image.
Four sets of 7243 images — real and synthetic, at low and high resolution — and every pairwise comparison the paper could make between them.
The headline is FID 40.5 between 7243 real and 7243 synthetic high-resolution radiographs. On its own that number means very little, which is why the table below includes the comparisons that give it a scale: what two halves of the same set score against each other (the floor), and what two different real devices score against each other (the ceiling of what "different but both real" looks like).
Read the rose bar at 40.5 against the green band at 36.8–70.1. The distance from real to synthetic at high resolution sits inside the range that separates two real datasets acquired on different machines. Read it also against the grey floor at 5–6, which is what the same distribution scores against itself: 40.5 is not small in absolute terms.
Every arm is compared with the real set at its own resolution. FID and the kernel inception distance (KID) say how far apart two sets are; KID is unbiased, so it depends less on the number of images. Precision and recall say where a distance comes from: precision is the share of generated images that lie close to real ones (fidelity), recall the share of real images that lie close to generated ones (coverage). Both generators were trained from scratch on the 7243 real radiographs, and both upscalers fine-tuned on the same 7243, so every arm has seen the same data.
| Generated set | FID ↓ | KID ×103 ↓ | Precision ↑ | Recall ↑ |
|---|
Numbers are the source indices of the paper's Table 8. KID is the mean ± sd over 100 subsets of 1000 images. Precision and recall use the same Inception features with k = 3 nearest neighbours. PanoDiff with HAT-SR is the closest generated arm to the real radiographs on all four measures, each by a wide margin. The last two rows apply the same upscalers to the real low-resolution images, which isolates the upscaling step: both land far closer to the real set than any generated arm, so most of the remaining distance comes from the generator. There the order reverses — with a real input the correct output exists, a pixel-loss network returns a smooth estimate that stays close to it, and a network trained with adversarial terms resynthesises detail that is plausible but not this radiograph's. On generated seeds there is no true detail to stay close to.
What these metrics can and cannot say. FID compares the mean and covariance of Inception features between two sets; it is sensitive to sample size, so numbers are only comparable within a table, which is what KID is there to relieve. Precision and recall split a distance into fidelity and coverage. Inception score never looks at the real data at all — it rewards an ImageNet classifier being confident and varied, which is why a model can improve on IS while drifting away from the distribution it is supposed to be modelling. That happens twice on this page: in the diffusion ablation, and with the SwinIR arms, which score above real radiographs. None of these metrics was designed for radiographs, and none knows what a mandible is. That is what the next sections are for.
Three early-career dentists with under ten years of practice and three with over fifteen, spanning prosthodontics, oral surgery, orthodontics and oral radiology. Each sat one session of about an hour, alone.
Each radiograph was displayed for twelve seconds, including about two seconds of loading. For each one the observer moved a slider to one of five positions: definitely real, probably real, unsure, probably fake, definitely fake. The graded scale is deliberate — forcing a binary choice would have thrown away the information that matters most, which is how often an expert simply could not tell.
Bars read left to right: definitely real, probably real, unsure, probably fake, definitely fake. The counts here are computed from the study's raw submission log and reproduce the paper's Table 5 exactly.
A five-level scale needs a rule for what an unsure response is worth, and it is fair to ask whether that rule is carrying the headline result. The paper reports all four candidate rules.
Three of the four agree to within 3.4 points, so the reported accuracy is a property of the responses rather than of the weighting. The fourth — discarding unsure responses — lifts the mean to 78.5% and lands it exactly on the 78.2% GAN benchmark the paper's central claim rests on being below. It does that not because the radiographs became easier to detect but because the 157 hardest judgements were removed from the denominator.
Better than guessing, and lower than for GAN radiographs. Treating each reader as one observation, the six accuracies give a 95% confidence interval of 62.6–74.3%, which excludes 50% (one-sample t-test, 8.15 standard errors above chance, p = 0.0005; exact Wilcoxon p = 0.031, the smallest attainable with six readers), and every reader's ROC area exceeds 0.5 (DeLong, all p < 10−11). So the images are not undetectable. The same six accuracies are lower than the 78.2% sensitivity reported for StyleGAN2-ADA radiographs (p = 0.008) and the 82.3% accuracy for StyleGAN3 radiographs with jaw cysts (p = 0.002). Neither study replicates this protocol, so these are comparisons with published reference points rather than a head-to-head. Paper §4.1.
Values run from 0.18 to 0.43, which is fair at best on the Landis–Koch scale, and there is no meaningful difference between the early-career and experienced groups. In a real-versus-synthetic task this is the favourable direction: high agreement would mean the observers were all spotting the same artefact. Low agreement, together with 68.5% accuracy, says a sizeable fraction of the synthetic radiographs carried no consistent tell at all. The highest pair, EC3 and EP3 at 0.43, is also the pair whose response distributions look most alike.
Examples were chosen for each category by the majority of the observers' assessments, not by eye. Click either image to enlarge.
A vision transformer trained to separate real from synthetic reaches 97.5% test accuracy. The humans reached 68.5%. Where it looks, measured over the whole held-out set, says what kind of difference it has found.
A ViT was trained on 7243 real and 7243 synthetic high-resolution radiographs with an 80:20 split. It separates them almost perfectly, which is worth stating plainly: these images are not indistinguishable, they are hard for a person under time pressure. Attention rollout then shows which patches drove each decision.
Top three rows of the grid are real radiographs, bottom three synthetic. A dot marks the four the classifier got wrong; all four are synthetic radiographs it read as real. C is the classifier's output, 1 for real and 0 for synthetic.
More diffuse, not misplaced. On the thirty samples above, attention on synthetic radiographs looks more diffuse than on real ones. Measured over all 2898 held-out radiographs, that holds: the entropy of the map is higher for synthetic images (6.100 against 6.044 nats), by far the largest effect in the analysis. The share of attention on anatomy, by contrast, is nearly the same for both classes. What separates them is how concentrated the attention is, not where it falls — consistent with a detector that checks whether the expected fine-scale texture is present, which in PanoDiff-SR is reconstructed by the super-resolution stage.
Each map is scored against fixed anatomical templates — the 1000 expert teeth and jaw masks of the TUFTS dataset, averaged and thresholded at 0.5 — so the regions are identical for every image. With 1449 images per group every difference is significant; the effect sizes (rank-biserial) are the quantity to read, and only entropy (+0.607) is large.
The classifier above was trained to find differences. MedSAM was not. Prompted identically on real and synthetic radiographs, it segments synthetic teeth exactly as it segments real ones.
MedSAM is a general medical segmentation model, fine-tuned from SAM on about 1.5 million image–mask pairs from other modalities and organs. Given a box around an object it returns a mask and a score for how good it believes that mask to be. If synthetic teeth look like real teeth to such a model, it should segment them with the same confidence and accuracy. It was run on all 7243 synthetic radiographs and on the 2697 real radiographs that have a manual tooth mask: 500 ADLD, 634 DENTEX, 598 TSXK and 1000 TUFTS, less 35 with an empty mask. The real radiographs were prepared exactly as the training corpus was, so both sets are framed alike — straight from the source files, their dentition would cover a quarter less of the image, and every measure would read that as anatomy. Every image received the same three kinds of box prompt: a fixed box at the same position in every image, a dentition box around the tooth region, and one box per tooth region.
| Prompt | Measure | Real | Synthetic | r | p | Device vs. others |r| |
|---|
r is the effect size of the real–synthetic difference (rank-biserial; r > 0 means higher on real; |r| < 0.1 negligible, 0.1–0.3 small, above 0.3 medium) and p the two-sided Mann–Whitney value. The last column gives, for scale, the range of |r| when each real device is compared with the other three on the same measure. Dice is the overlap with the tooth region the prompt was built from; against the manual masks MedSAM reaches 0.83 with tooth-region prompts and 0.78 with the dentition box.
MedSAM cannot tell them apart. Under all three prompts its predicted quality is almost the same on both sets — 0.464 against 0.466 with the fixed box, 0.480 against 0.472 with the dentition box, 0.505 against 0.507 per tooth region — every difference negligible (|r| ≤ 0.08), and not even in a consistent direction. The real devices differ from each other by more than that. The one effect above negligible is that synthetic dentitions break into slightly more regions (4.2 against 3.4 per image), still inside the range of the real devices. A classifier trained on the two classes separates them at 97.5% (§8); a general segmentation model does not, so what separates them is a signature a detector can learn, not a property such a model relies on. The reference regions come from the authors' own U-Net, used identically on both sets, and tooth regions are connected regions rather than individual teeth.
Six models trained from scratch at full resolution on the same 7243 radiographs under a comparable budget, to test the premise the whole pipeline rests on. The answer is not a clean win.
The main purpose of PanoDiff-SR is efficient synthesis, which is why the rest of the paper compares it with FastGAN, a GAN built for low training cost that fits the same small-generator-plus-upscaler setting. Generating directly at 1024 × 512 is the control that has to be run once: StyleGAN2-ADA, StyleGAN3, FastGAN-HR, ADM, LDM and PanoDiff-HR — the same model as PanoDiff but trained at full resolution — all on the same data and with comparable iterations and training budget.
| Model | Family | FID ↓ | IS ↑ | Precision % ↑ | Recall % ↑ | ViT detect % ↓ | Cost (acc.-h) |
|---|
† The two-stage rows were timed per training step on the same accelerator as the other rows and multiplied by their number of steps; the SR stage is 61.5 of these hours. Precision (fidelity) and recall (coverage) are computed image by image as in Paper Figure 17, with k = 3.
Four things this table settles, one of them against the paper.
Why the six dentists were not shown the GAN arms. Every published expert evaluation of GAN-generated panoramic radiographs we found reports that clinicians could tell them apart or rated them unrealistic: 78.2% sensitivity for StyleGAN2-ADA PRs (Schoenhof et al., 2024), 82.3% for StyleGAN3 PRs with jaw cysts (Fukuda et al., 2024), 81.6% mean accuracy by four dental radiologists on conditional-GAN PRs, one of them without error (Fu et al., 2026), and realism scores of 2.44–2.72 out of five (Pedersen et al., 2025). Where GANs and diffusion were compared directly, the GAN gave poor results and the diffusion images were hard to tell from real ones (Adnan et al., 2026; Kirkwood et al., 2025). The paper therefore reads the GANs' lower FID as a distributional match, not as clinical realism.
FID and IS have no idea what a mandible is, so four global anatomical measures were computed over 1500 images per model: bilateral symmetry, the periodicity of the tooth row, the number of distinct enamel crowns, and occlusal separation.
* differs from real at p < 0.05, ** at p < 10−4 (Mann–Whitney U). The last row is the fraction of images falling outside the central 98% of real radiographs on any of the four measures, so the real column's 7.3% is the floor that statistic can reach. LDM's 53.9% and its 26.1 enamel crowns are the signature of the repeated tooth rows visible in the figure below. These measures reject degenerate output reliably; they are global, and they say nothing about crown-to-root proportion or root morphology.
The corpus pools five machines, so a fair worry is that the model has simply learned the largest source. Each device was therefore compared with the synthetic set and with the pooled real corpus at a matched n = 500, redrawn five times.
Each device set was cropped and resized exactly as the training corpus was. The synthetic set is 71.5 (TUFTS) to 102.0 (TSXK) from the five devices, a spread of 1.43×, while the devices are 61.7 to 117.0 from one another, a spread of 1.90×, so the generator does not favour one device beyond how much the devices already differ. The device it matches least, TSXK, is also the one most distinct from the other four (mean FID 97.1, against 74.2–95.2 for the rest); five devices are too few to make that relation significant (Pearson r = 0.76, p = 0.13). For scale, two random halves of the real corpus give 37.8 at n = 500.
The next steps the paper sets out, each with the question it would answer.
U-Net size, SR loss weights, number of diffusion timesteps and the dataset mix were taken from the source implementations. A systematic search over such choices has improved diffusion models for natural images (Karras et al., 2022).
HAT and SwinIR were fine-tuned with different losses and degradations. Fine-tuning SwinIR with HAT's recipe would show how much of the difference is the architecture.
The five sources occupy different regions of feature space. Harmonising them, or conditioning the generator on the device, could remove this source of variation (Guan & Liu, 2022).
Each of the six observers read every image once. More readers, with repeated readings, would measure within-reader variation and test the effect of clinical experience (Obuchowski & Bullen, 2022).
The memorisation test compares whole images. Diffusion models can copy training images in part (Carlini et al., 2023; Somepalli et al., 2023), so a patch-level test should precede any public release.
Conditioning the generator on caries or fractures would supply rare cases for detection models, as synthetic lesions have done in other imaging tasks (Frid-Adar et al., 2018).
Both stages, the six baselines of Table 11, every analysis script of the paper, the observer-study application the dentists actually used, and the checkpoints behind every figure on this page.
Six epochs of PanoDiff — 11, 33, 55, 77, 99 and 110 — the six columns of the training
progression above.
archive.org/download/panodiff/trained_models
Real_HAT_GAN_SRx4_finetuned, the 400k-iteration EMA model used for every high-resolution
image on this page.
archive.org/download/panodiff/experiments
All five are public or available on request. Links and terms are in §3; none of them is redistributed here.
@misc{jain2025panodiffsr,
title = {PanoDiff-SR: Synthesizing Dental Panoramic Radiographs using Diffusion and Super-resolution},
author = {Jain, Sanyam and Neves de Freitas, Bruna and Basse-O'Connor, Andreas
and Iosifidis, Alexandros and Pauwels, Ruben},
year = {2025},
eprint = {2507.09227},
archivePrefix = {arXiv},
primaryClass = {eess.IV},
url = {https://arxiv.org/abs/2507.09227}
}