Command Palette
Search for a command to run...
تحسين نماذج السطح الرقمية الفوتوغرامترية باستخدام نماذج الانتشار المدربة مسبقًا والتكييف متعدد الوسائط
تحسين نماذج السطح الرقمية الفوتوغرامترية باستخدام نماذج الانتشار المدربة مسبقًا والتكييف متعدد الوسائط
Antoine Lorentz Stéphane May Valentine Bellet Dawa Derksen Bastien Nespoulous
الملخص
يمكن إنتاج نماذج السطح الرقمية واسعة النطاق (DSMs) بتكلفة فعّالة من الصور الفضائية عبر المسح التصويري المجسّم. غير أن الخرائط ثلاثية الأبعاد الناتجة كثيرًا ما تكون ملوثة بالضوضاء والقيم الشاذة والفجوات. وفي المقابل، يوفر الليدار الجوي قياسات ارتفاع عالية الدقة لكن بتكلفة أعلى بكثير. في هذا العمل، ندرس نماذج الانتشار المكيَّفة بكل من نماذج السطح الرقمية الفوتوغرامترية وصور Pléiades لتحسين نماذج السطح الرقمية المسجلة رأسيًا. ونقدم بنية معدلة من Stable Diffusion 3 مع تدفق نصي مشذّب واستراتيجية تطبيع على مستوى الرقع، مما يتيح تدريبًا مستقرًا على بيانات الليدار ونقلًا من الصور الطبيعية إلى خرائط الارتفاع. وتُظهر التجارب في مدن فرنسية أن التكييف متعدد الوسائط يحسّن دقة الارتفاع، حيث يخفض الجذر التربيعي لمتوسط مربع الخطأ (RMSE) في المناطق الحضرية الكثيفة من 6.00 إلى 3.45 م في المدن المدرجة في سياق التدريب، ومن 4.16 إلى 2.77 م في مدينة بوردو المستبعدة من التدريب.
One-sentence Summary
Researchers from Thales and CNES introduce a modified Stable Diffusion 3 architecture with a pruned text stream and patch-wise normalization that refines photogrammetric DSMs using Pléiades imagery as multimodal conditioning, improving elevation accuracy with Dense Urban RMSE reductions from 6.00 to 3.45 m in the in-context French cities and from 4.16 to 2.77 m in held-out Bordeaux.
Key Contributions
- Introduces a patch-wise normalization strategy that stabilizes diffusion model training on locally low-variance elevation data from LiDAR.
- Develops an image-only architecture by pruning the text stream from Stable Diffusion 3's multimodal transformer, reducing transformer parameter count from roughly 2B to 1B while retaining pretrained visual representations.
- Uses separate ControlNets to condition on photogrammetric DSMs and Pléiades imagery, with experiments in French cities showing DSM plus RGB conditioning reduces Dense Urban RMSE from 6.00 to 3.45 m in-context and from 4.16 to 2.77 m in held-out Bordeaux.
Introduction
State-of-the-art satellite stereo-photogrammetry pipelines such as CARS, S2P, MicMac, and NASA ASP can produce large-scale digital surface models at low cost, but their outputs generally suffer from high noise, missing coverage from occlusions or low-texture regions, and matching errors. Aerial LiDAR is more accurate but is expensive and often impractical for frequent, large-area mapping, especially in remote areas. Prior generative terrain work has mainly targeted void filling or text-driven generation, while related depth completion, GAN restoration, and super-resolution methods operate under different task assumptions rather than systematically refining dense noisy photogrammetric DSMs. The authors introduce a conditional diffusion approach that adapts a pretrained Stable Diffusion 3 backbone to refine photogrammetric DSMs toward LiDAR quality under vertical co-registration and surface-compatibility assumptions. Their main contributions are a patch-wise elevation normalization strategy for stable training, an image-only SD3 variant that removes the text stream to reduce transformer parameters from about 2B to 1B, and separate ControlNets for conditioning on the DSM and Pléiades-HR optical imagery.
Dataset
Dataset sources and scale
- The authors use a multi-city dataset assembled from IGN LiDAR-HD point clouds, orthorectified Pléiades-HR imagery, and CARS stereo DSMs.
- Raw LiDAR coverage spans 20 French cities: Amiens, Angers, Arcachon, Biarritz, Bordeaux, Clermont-Ferrand, Grenoble, Lyon, Marseille, Montpellier, Nancy, Nantes, Nice, Paris, Poitiers, Reims, Rennes, St-Etienne, Strasbourg, and Toulouse.
- Pléiades-HR imagery and DSMs are available for 9 French cities: Amiens, Arcachon, Biarritz, Bordeaux, Montpellier, Nice, Paris, Strasbourg, and Toulouse.
- All products are processed as 512 × 512 pixel patches at 0.5 m resolution in Lambert93.
- Table 2 summarizes modality availability by city and distinguishes geographic patch locations from Pléiades acquisition-patch pairs.
LiDAR and DSM preprocessing
- LiDAR-HD point clouds are rasterized to 0.5 m elevation rasters by median aggregation with the CARS-rasterize plugin.
- The filling workflow processes 2048 × 2048 pixel blocks with 50-pixel overlap. Connected LiDAR NoData components below an 800-pixel sieve threshold are filled. Eligible pixels are filled from a 7 × 7 neighborhood using the mean of values below the local 40th percentile, for at most five iterations.
- The percentile and neighborhood rules were empirical choices informed by tests around buildings, not a sensitivity-tested optimization.
- Any 512 × 512 target crop containing remaining LiDAR NaN is discarded.
- Stereo DSMs are supplied by the CARS workflow at 0.5 m resolution. Orthorectification through Reference3D and reprojection to Lambert93 provide the geometric starting point, but do not by themselves ensure vertical co-registration.
- For each DSM acquisition, the authors estimate an acquisition-level vertical offset from overlap with LiDAR. They build a global histogram of valid DSM-LiDAR differences over ±100 m with 4096 equal-width bins and subtract the modal bin center.
- The model assumes a vertically calibrated DSM input. It also assumes the DSM and LiDAR represent a compatible surface, so patches with calibrated RMSE above the city-specific 90th percentile are excluded from conditioning.
- High-RMSE excluded patches form a separate curation cohort of 4064 geographic patches. Of these, 18 have no Pléiades acquisition listed, 17 in Montpellier and one in Bordeaux. The remaining 4046 RGB-supported patches are used for the stress evaluation in Appendix D.
RGB preprocessing
- Pléiades RGB imagery is orthorectified but has no common radiometric calibration across acquisitions.
- When Airbus does not supply pansharpened imagery, the authors perform pansharpening with GDAL.
- For each acquisition and RGB band, intensities are clipped using the 2nd and 98th percentiles.
- The NIR band is discarded because the model accepts only three input channels.
- Multiple temporal acquisitions over the same location are included, but they remain grouped as one geographic sample. One acquisition is drawn randomly during collation, so additional dates do not increase that location’s sampling weight.
Patch extraction and data splits
- Patch extraction uses non-overlapping 512 × 512 windows with stride 512. Border remainders are dropped.
- Each city, x, y location defines one geographic patch. Adjacent patches may touch but share no pixels, and no geographic buffer is applied.
- All Pléiades acquisitions attached to a location remain grouped in the same split.
- The LiDAR-only training set, used for backbone adaptation, uses the 19 non-held-out LiDAR cities.
- LiDAR+Pléiades locations without DSM conditioning are split into approximately 90% training and 10% validation.
- Fully paired LiDAR+DSM+Pléiades patches within each non-held-out city are split into 70% training, 10% validation, and 20% test.
- All fully paired Bordeaux patches are reserved as the held-out test city. Bordeaux locations without all three modalities are not used for training or validation.
- No geographic patch appears in more than one split. A training or validation patch may be reused in the corresponding partition across successive stages when additional modalities are available.
- Exact split counts by city and stage are reported in Appendix C, Table A4.
Land cover stratification
- Conditional elevation-error metrics are stratified by land cover using the Theia Land Cover 2021 map at 10 m resolution, reprojected with nearest-neighbor sampling to 0.5 m.
- One source land-cover label therefore covers about 20 × 20 evaluation pixels, providing coarse stratification rather than pixel-pure object boundaries.
- The original Theia classes are grouped into five broader categories: Dense Urban, Sparse Built-Up, Roads, Croplands, and Vegetation.
- Table 3 reports the number of pixels per source land-cover class and dataset split for locations containing LiDAR-HD, Pléiades-HR, and stereo DSM data.
- The same finite LiDAR/DSM support is used for all methods when computing reported MAE and RMSE values.
Method
The authors leverage flow matching to learn a continuous transformation between a noise distribution p0 and a data distribution p1 through an Ordinary Differential Equation (ODE):
dtdxt=vt(xt)where the time-dependent velocity field vt transports probability mass along a density path pt. In multisample flow matching, the model is trained by sampling pairs (x0,x1) from a joint distribution q(x0,x1) whose marginals match the source and target distributions. A trajectory xt is constructed using a deterministic interpolation function, typically a straight line for rectified flow: xt=(1−t)x0+tx1. The corresponding conditional velocity field is vt(xt∣x1)=x1−x0, and the neural network is trained to match this field via the Joint Conditional Flow Matching (JCFM) objective.
Directly modeling transport from standard Gaussian noise to raw elevation patches causes instability due to poorly scaled training targets. To mitigate this, the authors introduce a patch-wise noise scaling scheme. They define a scale s and an offset u, shifting and scaling the noise distribution to match local data characteristics: x0∼N(u1,s2I). By reparameterizing with standard Gaussian noise ϵ∼N(0,I), the source noise is expressed as x0=sϵ+u, and the normalized target data as x^1=(x1−u)/s. The transport trajectory in normalized space becomes x^t=(1−t)ϵ+tx^1, with a target velocity field vt(xt)=s(x^1−ϵ). The neural network is reparameterized to predict the velocity in normalized space, v^θ(x^t,t,s,u)≈x^1−ϵ. The final training objective is normalized by the global scalar E[s2]:
L(θ)=Et,ϵ,x1[E[s2]s2∥v^θ(x^t,t,s,u)−(x^1−ϵ)∥2]During sampling, the model integrates standard Gaussian noise through the normalized probability flow and applies an explicit denormalization step using s and u to generate the target data x1.
The authors adapt Stable Diffusion 3 Medium (SD3), a dual-stream diffusion transformer, for elevation map generation. Elevation patches are adapted to an RGB-like format by duplicating the normalized single-channel patch across three channels for the Variational AutoEncoder (VAE). At inference, the three decoded channels are averaged before conversion to meters. Two sinusoidal encoders followed by Multi-Layer Perceptrons (MLPs) embed the conditioning variables s and u, adding the resulting embeddings to the timestep embedding.
To obtain an efficient, image-only backbone, the text stream of the Multimodal Diffusion Transformer (MM-DiT) is pruned. Joint cross-attention blocks attending over text tokens are replaced with self-attention over image tokens. Projection matrices, feed-forward networks, LayerNorm modules, and modulation layers specific to the text stream are removed, along with external text encoders. This reduction lowers the transformer size from 2B to 1B parameters.
To incorporate auxiliary modalities like Digital Surface Models (DSM) and Pléiades RGB imagery, the authors adopt a multi-branch ControlNet strategy. Two distinct ControlNets are instantiated, one for each modality. At each transformer block, modality-specific residuals are added element-wise before injection into the pruned SD3 backbone.
The dataset comprises LiDAR-HD, orthorectified Pléiades-HR imagery, and CARS stereo DSMs processed as 512×512 pixel patches at 0.5 m resolution. All rasters are reprojected into the Lambert93 coordinate system. Patch extraction precedes splitting, with non-overlapping windows defining geographic patches. The partitioning strategy ensures that patches from the same geographic location remain in the corresponding partition as modalities are added, while all fully paired Bordeaux patches form a held-out test set.
Training proceeds in three stages: backbone adaptation, single-modality ControlNet training, and sequential multimodal ControlNet training. For backbone adaptation, the pruned SD3 architecture is trained on the LiDAR training set. When initializing from pretrained weights, a learning rate of 1×10−5 is used. For the DSM-only ControlNet, the backbone is frozen, and the ControlNet is trained on the paired LiDAR+DSM+Pléiades set with a batch size schedule increasing from 120 to 360. For multimodal conditioning, the Pléiades ControlNet is trained first on the LiDAR+Pléiades set, then frozen, followed by the addition and training of the DSM ControlNet on the fully paired set. Both ControlNet stages use a learning rate of 1×10−5.
Experiment
The study evaluates a pruned SD3-based elevation refinement model by comparing pretrained and scratch backbones, measuring a frozen VAE round-trip diagnostic on LiDAR patches, and incrementally adding stereo DSM and Pléiades RGB conditioning across in-context cities and a held-out Bordeaux test set. Pretrained initialization performs substantially better than the best scratch configuration, while the VAE diagnostic shows that the image representation retains elevation structure with measurable but modest distortion. The conditioning results indicate that the stereo DSM provides a useful geometric anchor and that adding RGB generally improves refinement, especially in dense urban areas, although benefits vary by land cover and are not guaranteed everywhere, particularly in homogeneous cropland and some vegetation. Overall, the experiments support optical imagery as a complement to stereo geometry while highlighting limitations related to acquisition mismatch, spatial variability, and the constrained scope of the evaluation.
Across the compared capabilities, prior methods are largely specialized: several address void filling, Marigold-DC addresses sparse-depth completion, DSM super-resolution addresses resolution change, and DSM-to-LoD2 addresses dense DSM correction with a LoD2-like target. The proposed method shares dense DSM correction with DSM-to-LoD2 but adds RGB conditioning and treats void filling as a secondary optional capability. It is not associated with sparse-depth completion, resolution change, or LoD2-like targets. Diff-DEM, Dfilled, and GAN-based void filling methods all share a void-filling association but are not marked for dense DSM correction. The proposed method and DSM-to-LoD2 both address dense DSM correction, but only the proposed method combines it with RGB conditioning and a secondary optional void-filling capability.
The dataset shows heterogeneous multimodal coverage across the reported French cities. LiDAR patches are present for every listed city, while Pléiades and stereo DSM patches are available only in a subset. In cities with Pléiades coverage, the all-acquisition patch count exceeds the location-level count, indicating repeated acquisitions over time. LiDAR patches are available for every reported city, but Pléiades and stereo DSM coverage is absent for Angers, Clermont-Ferrand, and Grenoble. Bordeaux has the largest all-acquisition Pléiades patch count among the reported cities, while Amiens has the largest LiDAR patch count. Stereo DSM patch counts are consistently smaller than LiDAR and Pléiades counts in the cities where stereo DSM data are present. All-acquisition Pléiades counts are higher than location-level Pléiades counts across covered cities, reflecting multiple image dates per location.
The land cover splits are dominated by dispersed urban and deciduous forest pixels, while dense urban, beaches and dunes, and permanent snow are far less represented. Bordeaux, the held-out city, has lower pixel counts across all classes, with particularly limited dense urban and coastal coverage. The validation split is consistently larger than the Bordeaux train split for all classes with nonzero coverage. Dispersed urban is the most common class in both the in-context data and the Bordeaux test split, exceeding dense urban by a wide margin. Deciduous forest is the second most common in-context class but has much lower held-out coverage in Bordeaux. Dense urban is scarce in the Bordeaux test split relative to other built and forest classes. Beaches and dunes remain a small class in every split, and glaciers or permanent snow have zero pixels.
The table compares training configurations and computational cost for backbone adaptation, single-modality ControlNet training, and sequential multimodal ControlNet training. Backbone fine-tuning and from-scratch training use the same step and batch settings but differ in learning rate, while ControlNet stages share a two-phase batch-size schedule. Multimodal training adds ControlNets sequentially and incurs the highest total cost because the DSM stage runs after the Pléiades ControlNet is frozen. Backbone adaptation from pretrained weights and from scratch use the same step and batch settings, but from-scratch training uses a higher learning rate. Single-modality ControlNet training keeps the backbone frozen and follows a two-phase schedule with a batch-size increase, making the later phase more computationally expensive. Multimodal training adds ControlNets sequentially, freezing the Pléiades ControlNet before DSM ControlNet training, and has the highest total computational cost among the compared configurations.
A text-stream-pruned SD3 model initialized with pretrained image-stream weights achieves a much lower FD_DINOv2 than the same architecture trained from scratch. The pretrained initialization is roughly 4x better under this feature-distribution distance, and the metric is lower-is-better. This suggests SD3 RGB pretraining provides a clear transfer advantage for elevation-map generation in the pruned setup. Pretrained SD3 image-stream weights give the best FD_DINOv2, substantially below the best scratch-trained configuration. Random initialization leads to a much higher feature-distribution distance, indicating that RGB pretraining helps the pruned architecture.
The experiments evaluate a dense DSM correction method with RGB conditioning and optional void filling, distinguishing it from prior specialized approaches. Dataset and land cover analyses show multimodal coverage is uneven across French cities, with LiDAR available everywhere while Pléiades and stereo DSM appear only in subsets, and splits are dominated by dispersed urban and deciduous forest while several classes are scarce or absent. Training comparisons indicate sequential multimodal ControlNet training incurs the highest computational cost, and pretrained SD3 image-stream initialization provides a clear transfer advantage over from-scratch training for pruned elevation-map generation.