EUReCA: End-to-End Unified CBCT Reconstruction Across Clinical Acquisitions

Jiening Zhu, Chengzhu Zhang, Jason Fan, LiCheng Kuo, Weixing Cai, Xiuxiu He, Laura Cervino, Jean Moran, Xiang Li, Tianfang Li, Yabo Fu

Memorial Sloan Kettering Cancer Center

Problem setup

Given \(N\) cone-beam projections \(\{(p_i,\theta_i)\}_{i=1}^{N}\) acquired at gantry angles \(\theta_i\), possibly non-uniformly over a full or partial arc, we seek a single feed-forward mapping \(f_\phi:\{(p_i,\theta_i)\}_{i=1}^{N}\mapsto\mathbf{V}\) to the CBCT volume \(\mathbf{V}\), with parameters \(\phi\) shared across view counts, angular distributions, and arc spans — one model for every acquisition regime seen in training, without view-count-specific retraining. The simulated geometry follows a clinical Varian Halcyon on-board imager rather than an idealised orbit.

Halcyon acquisition geometry
Halcyon acquisition geometry. (a) 211° half-scan arc with the arc start free on the 360° circle; coloured markers show independent sparse view draws. (b) Cone-beam projection geometry, reconstruction volume, and detector pixel.

One model, 2–100 views

Angular-binned aggregation gives a fixed-size representation regardless of projection count, so a single architecture spans ultra-sparse to well-sampled regimes.

Geometry-aware

Multiscale 2D features are back-projected into 3D grids using the exact cone-beam geometry — no idealised orbit assumptions.

FDK-anchored

A full 3D FDK reconstruction of the same projections enters the decoder through voxel-wise gates at every scale, keeping outputs tied to the measurements.

Clinically validated

Best PSNR/SSIM at every view count on 15 real Varian Halcyon HyperSight scans, against vendor reconstructions, in 0.3–1.3 s per volume.

Interactive volume viewer

Full 256×256×128 reconstructed volumes (2 mm isotropic), decoded in your browser. Choose a patient and view count, then scrub through slices with the slider or the mouse wheel over an image; adjust the display window with the presets. Volumes load on demand (1–6 MB each) and are cached once loaded. Click an image to enlarge.

Patient — synthetic projections
Patient — real projections (Halcyon)
Object — synthetic projections (out of distribution)
Views
Plane
Show beside ours
Window (HU)
Slice: 64 / 127
Angular coverage · 10 views
EUReCA · 10 views
FDK · 10 views
Ground truth
Input projection
Projection 1/10 · θ=0°

Reconstruction vs. number of views

One model, synthetic projections, swept from 10 to 100 views on the Halcyon 211° arc. Drag the slider to change the view count; the polar plot shows the sampled gantry angles. Display window −400 to 600 HU. Click any panel to zoom.

Patient
Plane
Show beside ours
Views: 10
102030405060708090100
Sampled gantry angles
Angular coverage
EUReCA reconstruction
EUReCA · 10 views
FDK reconstruction
FDK · 10 views
Ground truth
Ground truth

Additional acquisition trajectories

The same architecture, training procedure and loss, trained once per trajectory with only the view-selection rule changed: two orthogonal views, four fixed gantry angles, and 90°/120° short arcs (arc start free on the 360° circle). Synthetic projections, reserved test cases. Display window −400 to 600 HU. Click any panel to zoom.

Patient
Trajectory
Draw
Plane
Show beside ours
Sampled gantry angles
Angular coverage
EUReCA reconstruction
EUReCA ·
FDK reconstruction
FDK ·
Ground truth
Ground truth
 

Abstract

Cone-beam computed tomography (CBCT) is widely used to guide radiotherapy delivery. Acquiring fewer projections reduces imaging dose and scan time but makes reconstruction increasingly ill-posed. We present EUReCA (End-to-end Unified REconstruction across Clinical Acquisitions), a geometry-aware feed-forward model for clinical linac CBCT. A single network reconstructs acquisitions ranging from 8 to 100 views — substantially broader than previously demonstrated by feed-forward methods. It is a unified framework which easily generalizes to diverse clinical acquisitions via only changing the training input. EUReCA encodes projections using a shared multiscale backbone and back-projects the resulting features into multiresolution 3D grids. Learned aggregators combine an arbitrary number of views at fixed computational cost at different resolution, a global 3D transformer refines the coarse representation, and a multilevel decoder recovers fine detail using both projection features and a reference Feldkamp–Davis–Kress (FDK) reconstruction. On a routine clinical patient-setup protocol, EUReCA achieves 29.9–33.2 dB PSNR with 10–100 views, outperforming all applicable feed-forward baselines. At 100 views, it also surpasses the strongest classical per-scan method in that method's well-sampled regime while reconstructing in 0.3–1.3 s. On 15 retrospectively downsampled clinical scans evaluated against vendor reconstructions, EUReCA performs best at every tested view count. These results establish geometry-aware, measurement-anchored feed-forward reconstruction as a practical approach to variable-view linac CBCT and a promising step toward a general-purpose reconstruction model.

Method

EUReCA method overview
Method overview. (A) Geometry-aware lifting: each of the N projections is encoded by a shared pretrained backbone into a three-scale feature pyramid; features at every scale are back-projected into 3D grids using the exact cone-beam geometry and pooled across views by angular-binned attention. (B) Coarse-to-fine 3D decoder: the coarse volume is refined by a global transformer, then decoded through MLP and ×2 upsampling stages with multi-scale fusion. (C) Analytic FDK reference: an in-frame FDK reconstruction of the same projections is encoded by a 3D CNN and merged into the decoder at every scale through gated fusion.
Variable-view aggregation over fixed angular bins
Variable-view aggregation over fixed angular bins. (A) The N view features are organised into K = 36 fixed angular bins by gantry angle; unobserved bins are represented explicitly, yielding a sequence whose length is independent of N. (B) Per 3D location, bins are aggregated scale-specifically: a transformer with a class token at the coarse scale, learned softmax-weighted mixing at mid and fine scales, fused with mean/max statistics.

Data

21,000+ CT volumes

Seven public datasets plus one internal cohort, spanning thorax, abdomen, and pelvis — among the largest corpora assembled for feed-forward CBCT reconstruction.

Quality & realism filtering

Voxel spacing < 5 mm on every axis and craniocaudal extent > 300 mm, so DRRs are high-fidelity and rays never cross an artificial body boundary. 12,255 volumes retained.

Clean case-level split

80 / 10 / 10 train / val / test by patient; only the reserved test split is reported, shared by every baseline.

Real clinical scans

15 Varian Halcyon HyperSight acquisitions, evaluated against the vendor reconstructions.

DatasetRegionOriginalFilteredTrainValTest
MELAThorax (mediastinum)1,1009167319491
LUNA16Thorax (lung)8806114896062
RibFracThorax (ribs)6606185036055
AMOS22Abdomen2,3837966547567
RSNA-2023Abdomen (trauma)4,1382,4621,922259281
AbdomenAtlasAbdomen9,2625,9984,828586584
TotalSegmentatorWhole body1,2287566037578
Prostate (internal)Pelvis1,5809874168
Total21,23112,2559,8041,2251,226

Results

Simulated projections — routine patient-setup protocol

PSNR (dB) / SSIM, mean ± std over test cases. Feed-forward methods: full reserved test set (n ≈ 1328), matched train/test view counts. Per-scan optimisation methods: fixed 20-scan subset (SAX-NeRF @100v n = 14, wall-clock timeout); the "ours" row in that panel is re-scored on the same 20 scans. "—" = architecture cannot reach that view count (ILV additionally OOMs at 50 views).

Feed-forward methods

Method10 views20 views50 views100 viewsTime 10/20/50/100v
DIF-Net23.91 ± 3.29 / 0.724 ± 0.10625.19 ± 3.89 / 0.746 ± 0.11626.71 ± 4.11 / 0.786 ± 0.1111.73 / 5.00 / 14.9 s / —
DIF-Gaussian24.39 ± 4.73 / 0.741 ± 0.18823.30 ± 3.56 / 0.611 ± 0.14822.76 ± 2.49 / 0.560 ± 0.0782.06 / 5.64 / 16.9 s / —
ILV29.42 ± 2.02 / 0.875 ± 0.04230.17 ± 2.06 / 0.891 ± 0.0370.38 / 0.50 s / OOM / OOM
EUReCA (ours)29.93 ± 2.18 / 0.908 ± 0.03131.52 ± 2.28 / 0.928 ± 0.02632.77 ± 2.52 / 0.944 ± 0.02333.22 ± 2.61 / 0.952 ± 0.0220.32 / 0.41 / 0.71 / 1.33 s

Per-scan optimisation methods

Method10 views20 views50 views100 viewsTime 10/20/50/100v
FDK15.34 ± 4.16 / 0.252 ± 0.03417.76 ± 5.46 / 0.338 ± 0.05621.13 ± 7.26 / 0.579 ± 0.09122.45 ± 7.97 / 0.740 ± 0.1040.11 / 0.10 / 0.08 / 0.17 s
R²-Gaussian20.55 ± 6.95 / 0.550 ± 0.09623.06 ± 8.31 / 0.661 ± 0.10324.77 ± 9.30 / 0.768 ± 0.12024.86 ± 9.35 / 0.786 ± 0.1184.6 / 4.5 / 4.6 / 15.0 min
IntraTomo20.53 ± 7.71 / 0.572 ± 0.26923.71 ± 8.52 / 0.751 ± 0.13825.25 ± 9.40 / 0.822 ± 0.12425.78 ± 9.73 / 0.845 ± 0.1252.9 / 5.1 / 12.1 / 17.4 min
NAF21.76 ± 6.93 / 0.595 ± 0.26023.89 ± 8.64 / 0.771 ± 0.11025.41 ± 9.50 / 0.828 ± 0.12425.84 ± 9.76 / 0.845 ± 0.1256.5 / 11.9 / 27.3 / 38.3 min
SAX-NeRF22.41 ± 7.99 / 0.752 ± 0.09624.41 ± 8.97 / 0.805 ± 0.11025.93 ± 9.88 / 0.845 ± 0.12327.92 ± 8.76 / 0.887 ± 0.1060.45 / 0.49 / 1.19 / 3.56 h
EUReCA (ours)29.61 ± 1.91 / 0.911 ± 0.02831.08 ± 2.55 / 0.928 ± 0.02832.24 ± 2.88 / 0.942 ± 0.02731.56 ± 2.87 / 0.943 ± 0.0270.32 / 0.41 / 0.71 / 1.33 s

Real clinical projections — 15 Halcyon HyperSight scans

PSNR (dB) / SSIM against vendor reconstructions. At 10–100 views each method sees only 2.5–30% of the 335–407 acquired projections.

MethodType10 views20 views50 views100 views
EUReCA (ours)FF25.75 / 0.78026.44 / 0.80626.94 / 0.82927.24 / 0.841
ILVFF25.19 / 0.72525.62 / 0.751OOMOOM
DIF-NetFF23.57 / 0.66623.82 / 0.67924.72 / 0.696n/a
DIF-GaussianFF24.28 / 0.66523.19 / 0.55722.53 / 0.506n/a
FDKOPT15.70 / 0.19017.78 / 0.25920.70 / 0.43122.32 / 0.582
R²-GaussianOPT22.58 / 0.42724.36 / 0.50425.43 / 0.61125.64 / 0.651
IntraTomoOPT24.20 / 0.70525.28 / 0.75626.03 / 0.79026.59 / 0.815
NAFOPT24.32 / 0.70625.45 / 0.75826.45 / 0.80526.91 / 0.827
SAX-NeRFOPT23.62 / 0.67824.68 / 0.72025.45 / 0.77125.77 / 0.799

Across acquisition geometries (same model, trained per geometry)

AcquisitionViewsPSNR (dB)SSIM
Two orthogonal views225.430.833
Four fixed views (0/45/90/135°)428.230.879
90° short arc15 / 30 / 4527.84 / 28.36 / 28.710.882 / 0.893 / 0.900
120° short arc15 / 30 / 4529.12 / 29.91 / 30.050.896 / 0.911 / 0.913

Data scaling

Convergence-matched data-size scaling
Convergence-matched data-size scaling. PSNR (a) and SSIM (b) versus number of unique training cases (log scale) at N = 10/20/50 views, best-validation checkpoints.

Citation

@article{zhu2026eureca,
  title   = {{EUReCA}: End-to-End Unified {CBCT} Reconstruction Across Clinical Acquisitions},
  author  = {Zhu, Jiening and Zhang, Chengzhu and Fan, Jason and Kuo, LiCheng and Cai, Weixing
             and He, Xiuxiu and Cervino, Laura and Moran, Jean and Li, Xiang and Li, Tianfang and Fu, Yabo},
  journal = {arXiv preprint},
  year    = {2026}
}