Pruning SD3's text stream and adding twin ControlNets halves satellite elevation error
Satellite stereo-photogrammetry is cheap and scalable; LiDAR is accurate and expensive. The gap between them is the problem this paper attacks directly. The approach takes photogrammetric Digital Surface Models produced by the CARS pipeline from Pléiades-HR imagery over French cities, then uses a modified Stable Diffusion 3 architecture to push those noisy, void-riddled elevation maps toward LiDAR quality.
The core architectural move is aggressive: the text stream of SD3's multimodal transformer is pruned entirely, cutting the transformer from roughly 2 billion to 1 billion parameters while keeping the pretrained image-stream weights intact. Two separate ControlNets handle the conditioning inputs — one for the stereo DSM, one for the optical imagery — with their residuals summed element-wise before injection into the backbone. The result is an image-only generator that never needed a text encoder to begin with, runs faster, and still inherits the visual priors baked into SD3 from natural-image training.
The normalization problem is less glamorous but arguably more important. Elevation patches have wildly varying means and tiny local variance relative to the full dataset range, which destabilizes standard flow-matching training. The paper's patch-wise normalization extracts scale and offset statistics from the conditioning DSM at training time and from the input DSM at inference, normalizing the target into a range compatible with the loss scale used during SD3's original training. No clipping, no zero-variance fallback — the smallest standard deviation encountered in the LiDAR data was under five thousandths of a meter, and the scheme handled it without special casing.
The numbers are concrete. Dense Urban RMSE falls from 6.00 m on the calibrated stereo input to 3.45 m with DSM plus RGB conditioning across eight in-context French cities. In Bordeaux, held out entirely from training, the same metric drops from 4.16 m to 2.77 m. Roads show similar clear gains. The story is less clean elsewhere: Bordeaux Vegetation MAE is slightly worse with the full model than with the raw input, and Croplands gains in-context are large but come with wide city-bootstrap intervals, suggesting a few cities with unusually bad stereo errors are driving the headline number.
The backbone comparison is worth noting. A pruned SD3 initialized from pretrained weights achieves an FD-DINOv2 score of 169.4 against LiDAR patches; the best scratch-trained run reaches only 692.8. That gap motivates the transfer approach, though the paper is careful not to conflate feature-distribution distance with downstream elevation accuracy.
Limitations are stated plainly. The method requires vertical co-registration as a precondition — it corrects local errors, it does not estimate an unknown vertical datum. Surface compatibility between DSM and LiDAR acquisitions is assumed but only loosely enforced via a 90th-percentile RMSE filter. No code or model weights are released. Bordeaux is one city, and all experiments stay within the Pléiades-HR/CARS acquisition setting. Transfer to other sensors, countries, or photogrammetric pipelines is explicitly deferred.
A disciplined adaptation of Stable Diffusion 3 to elevation refinement that halves Dense Urban DSM error in held-out tests — with honest accounting of where it fails.
Sources & links
Related on SkillFed
Voyager pairs GPT-4 with a growing skill library and a self-verification loop in Minecraft; its own ablations show task-ordering and outcome-checking, not raw model calls, drive…
Mirage-1 abstracts GUI trajectories into a three-tier skill hierarchy and lets Monte Carlo Tree Search write successful runs back into it, gaining 32.3% on AndroidWorld and 79.6%…