Problem setup
Given \(N\) cone-beam projections \(\{(p_i,\theta_i)\}_{i=1}^{N}\) acquired at gantry angles \(\theta_i\), possibly non-uniformly over a full or partial arc, we seek a single feed-forward mapping \(f_\phi:\{(p_i,\theta_i)\}_{i=1}^{N}\mapsto\mathbf{V}\) to the CBCT volume \(\mathbf{V}\), with parameters \(\phi\) shared across view counts, angular distributions, and arc spans — one model for every acquisition regime seen in training, without view-count-specific retraining. The simulated geometry follows a clinical Varian Halcyon on-board imager rather than an idealised orbit.
One model, 2–100 views
Angular-binned aggregation gives a fixed-size representation regardless of projection count, so a single architecture spans ultra-sparse to well-sampled regimes.
Geometry-aware
Multiscale 2D features are back-projected into 3D grids using the exact cone-beam geometry — no idealised orbit assumptions.
FDK-anchored
A full 3D FDK reconstruction of the same projections enters the decoder through voxel-wise gates at every scale, keeping outputs tied to the measurements.
Clinically validated
Best PSNR/SSIM at every view count on 15 real Varian Halcyon HyperSight scans, against vendor reconstructions, in 0.3–1.3 s per volume.
Interactive volume viewer
Full 256×256×128 reconstructed volumes (2 mm isotropic), decoded in your browser. Choose a patient and view count, then scrub through slices with the slider or the mouse wheel over an image; adjust the display window with the presets. Volumes load on demand (1–6 MB each) and are cached once loaded. Click an image to enlarge.
Qualitative comparisons
Pick a patient (synthetic or real projections) and a view count. To keep the page light, only the middle slice of each of the three planes is shown here; the volume viewer above has full volumes. Feed-forward baselines (DIF-Net, DIF-Gaussian, ILV) are architecture-limited to ≤20 views, so at 50 and 100 views only the per-scan optimisation panel is shown. Rows: axial / coronal / sagittal.
Feed-forward methods
Per-scan optimisation methods
Reconstruction vs. number of views
One model, synthetic projections, swept from 10 to 100 views on the Halcyon 211° arc. Drag the slider to change the view count; the polar plot shows the sampled gantry angles. Display window −400 to 600 HU. Click any panel to zoom.
Additional acquisition trajectories
The same architecture, training procedure and loss, trained once per trajectory with only the view-selection rule changed: two orthogonal views, four fixed gantry angles, and 90°/120° short arcs (arc start free on the 360° circle). Synthetic projections, reserved test cases. Display window −400 to 600 HU. Click any panel to zoom.
Abstract
Cone-beam computed tomography (CBCT) is widely used to guide radiotherapy delivery. Acquiring fewer projections reduces imaging dose and scan time but makes reconstruction increasingly ill-posed. We present EUReCA (End-to-end Unified REconstruction across Clinical Acquisitions), a geometry-aware feed-forward model for clinical linac CBCT. A single network reconstructs acquisitions ranging from 8 to 100 views — substantially broader than previously demonstrated by feed-forward methods. It is a unified framework which easily generalizes to diverse clinical acquisitions via only changing the training input. EUReCA encodes projections using a shared multiscale backbone and back-projects the resulting features into multiresolution 3D grids. Learned aggregators combine an arbitrary number of views at fixed computational cost at different resolution, a global 3D transformer refines the coarse representation, and a multilevel decoder recovers fine detail using both projection features and a reference Feldkamp–Davis–Kress (FDK) reconstruction. On a routine clinical patient-setup protocol, EUReCA achieves 29.9–33.2 dB PSNR with 10–100 views, outperforming all applicable feed-forward baselines. At 100 views, it also surpasses the strongest classical per-scan method in that method's well-sampled regime while reconstructing in 0.3–1.3 s. On 15 retrospectively downsampled clinical scans evaluated against vendor reconstructions, EUReCA performs best at every tested view count. These results establish geometry-aware, measurement-anchored feed-forward reconstruction as a practical approach to variable-view linac CBCT and a promising step toward a general-purpose reconstruction model.
Method
Data
21,000+ CT volumes
Seven public datasets plus one internal cohort, spanning thorax, abdomen, and pelvis — among the largest corpora assembled for feed-forward CBCT reconstruction.
Quality & realism filtering
Voxel spacing < 5 mm on every axis and craniocaudal extent > 300 mm, so DRRs are high-fidelity and rays never cross an artificial body boundary. 12,255 volumes retained.
Clean case-level split
80 / 10 / 10 train / val / test by patient; only the reserved test split is reported, shared by every baseline.
Real clinical scans
15 Varian Halcyon HyperSight acquisitions, evaluated against the vendor reconstructions.
| Dataset | Region | Original | Filtered | Train | Val | Test |
|---|---|---|---|---|---|---|
| MELA | Thorax (mediastinum) | 1,100 | 916 | 731 | 94 | 91 |
| LUNA16 | Thorax (lung) | 880 | 611 | 489 | 60 | 62 |
| RibFrac | Thorax (ribs) | 660 | 618 | 503 | 60 | 55 |
| AMOS22 | Abdomen | 2,383 | 796 | 654 | 75 | 67 |
| RSNA-2023 | Abdomen (trauma) | 4,138 | 2,462 | 1,922 | 259 | 281 |
| AbdomenAtlas | Abdomen | 9,262 | 5,998 | 4,828 | 586 | 584 |
| TotalSegmentator | Whole body | 1,228 | 756 | 603 | 75 | 78 |
| Prostate (internal) | Pelvis | 1,580 | 98 | 74 | 16 | 8 |
| Total | — | 21,231 | 12,255 | 9,804 | 1,225 | 1,226 |
Results
Simulated projections — routine patient-setup protocol
PSNR (dB) / SSIM, mean ± std over test cases. Feed-forward methods: full reserved test set (n ≈ 1328), matched train/test view counts. Per-scan optimisation methods: fixed 20-scan subset (SAX-NeRF @100v n = 14, wall-clock timeout); the "ours" row in that panel is re-scored on the same 20 scans. "—" = architecture cannot reach that view count (ILV additionally OOMs at 50 views).
Feed-forward methods
| Method | 10 views | 20 views | 50 views | 100 views | Time 10/20/50/100v |
|---|---|---|---|---|---|
| DIF-Net | 23.91 | 25.19 | 26.71 | — | 1.73 / 5.00 / 14.9 s / — |
| DIF-Gaussian | 24.39 | 23.30 | 22.76 | — | 2.06 / 5.64 / 16.9 s / — |
| ILV | 29.42 | 30.17 | — | — | 0.38 / 0.50 s / OOM / OOM |
| EUReCA (ours) | 29.93 | 31.52 | 32.77 | 33.22 | 0.32 / 0.41 / 0.71 / 1.33 s |
Per-scan optimisation methods
| Method | 10 views | 20 views | 50 views | 100 views | Time 10/20/50/100v |
|---|---|---|---|---|---|
| FDK | 15.34 | 17.76 | 21.13 | 22.45 | 0.11 / 0.10 / 0.08 / 0.17 s |
| R²-Gaussian | 20.55 | 23.06 | 24.77 | 24.86 | 4.6 / 4.5 / 4.6 / 15.0 min |
| IntraTomo | 20.53 | 23.71 | 25.25 | 25.78 | 2.9 / 5.1 / 12.1 / 17.4 min |
| NAF | 21.76 | 23.89 | 25.41 | 25.84 | 6.5 / 11.9 / 27.3 / 38.3 min |
| SAX-NeRF | 22.41 | 24.41 | 25.93 | 27.92 | 0.45 / 0.49 / 1.19 / 3.56 h |
| EUReCA (ours) | 29.61 | 31.08 | 32.24 | 31.56 | 0.32 / 0.41 / 0.71 / 1.33 s |
Real clinical projections — 15 Halcyon HyperSight scans
PSNR (dB) / SSIM against vendor reconstructions. At 10–100 views each method sees only 2.5–30% of the 335–407 acquired projections.
| Method | Type | 10 views | 20 views | 50 views | 100 views |
|---|---|---|---|---|---|
| EUReCA (ours) | FF | 25.75 / 0.780 | 26.44 / 0.806 | 26.94 / 0.829 | 27.24 / 0.841 |
| ILV | FF | 25.19 / 0.725 | 25.62 / 0.751 | OOM | OOM |
| DIF-Net | FF | 23.57 / 0.666 | 23.82 / 0.679 | 24.72 / 0.696 | n/a |
| DIF-Gaussian | FF | 24.28 / 0.665 | 23.19 / 0.557 | 22.53 / 0.506 | n/a |
| FDK | OPT | 15.70 / 0.190 | 17.78 / 0.259 | 20.70 / 0.431 | 22.32 / 0.582 |
| R²-Gaussian | OPT | 22.58 / 0.427 | 24.36 / 0.504 | 25.43 / 0.611 | 25.64 / 0.651 |
| IntraTomo | OPT | 24.20 / 0.705 | 25.28 / 0.756 | 26.03 / 0.790 | 26.59 / 0.815 |
| NAF | OPT | 24.32 / 0.706 | 25.45 / 0.758 | 26.45 / 0.805 | 26.91 / 0.827 |
| SAX-NeRF | OPT | 23.62 / 0.678 | 24.68 / 0.720 | 25.45 / 0.771 | 25.77 / 0.799 |
Across acquisition geometries (same model, trained per geometry)
| Acquisition | Views | PSNR (dB) | SSIM |
|---|---|---|---|
| Two orthogonal views | 2 | 25.43 | 0.833 |
| Four fixed views (0/45/90/135°) | 4 | 28.23 | 0.879 |
| 90° short arc | 15 / 30 / 45 | 27.84 / 28.36 / 28.71 | 0.882 / 0.893 / 0.900 |
| 120° short arc | 15 / 30 / 45 | 29.12 / 29.91 / 30.05 | 0.896 / 0.911 / 0.913 |
Data scaling
Citation
@article{zhu2026eureca,
title = {{EUReCA}: End-to-End Unified {CBCT} Reconstruction Across Clinical Acquisitions},
author = {Zhu, Jiening and Zhang, Chengzhu and Fan, Jason and Kuo, LiCheng and Cai, Weixing
and He, Xiuxiu and Cervino, Laura and Moran, Jean and Li, Xiang and Li, Tianfang and Fu, Yabo},
journal = {arXiv preprint},
year = {2026}
}