Files
lidar_rendu/docs/GROUND_CLASSIFICATION.md
Antoine fb892ea9f2 Translate the whole project to English and fix outdated comments and help
Comments, docstrings, logs, CLI help, map UI, legends, PDF sheet, scripts,
compose files and AGENTS.md are now English. Data keys stay unchanged
(relief_oriente, densite_sol, visualisations/, API JSON keys, link params).
Wrong comments and help defaults found along the way are corrected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 23:16:45 +02:00

161 lines
8.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Ground classification — options, benchmark and references
> **Status (up to date)**: the pipeline default is `--ground-classification ign`
> with `--ign-classes sol` (IGN vendor ground class 2), extracted directly with
> laspy (`_extract_ign_ground` in `dtm.py`), with PDAL as fallback. `auto`,
> `smrf` and `csf` remain selectable via `--ground-classification`. Map-triggered
> generation always uses `ign` (no classification/reconciliation setting is
> exposed in the map UI).
Synthesis note for choosing/improving the ground-detection algorithm.
Context: IGN LiDAR HD tiles (Lambert 93), areas of strong relief / rock outcrops /
dense forest where the ground is under-classified and the DTM shows large holes.
Reference tile: `LHD_FXX_0999_6778_PTS_LAMB93_IGN69`
- 65,047,919 points, ~1 km², resolutions 0.5 m and 0.2 m.
- Class breakdown (vendor pre-classification):
class 2 (ground) **25.79%**, class 5 (high vegetation) **63.2%**,
class 3 (low vegetation) 7.85%, class 4 (medium vegetation) 2.66%,
class 1 0.47%, class 6 0.01%. No points in class 0.
- Holes in the existing DTM (before correction): **43.5%** at 0.5 m,
**44.1%** at 0.2 m (tile extent, header bounds).
## Measured benchmark (PDAL, 1 km²)
| Method | Time | Ground points | Ground surface* | Holes* |
|---|---|---|---|---|
| **IGN** (vendor pre-classification) | **9.4 s** | 25.8% | 84.6% | 15.4% |
| **SMRF** | 326.3 s | 36.1% | 90.4% | 9.6% |
| **CSF** | 355.2 s | 17.6% | 46.9% | 53.1% |
\* "Ground surface" computed over the point extent (point cloud bounding box),
at 0.5 m. Final DTM hole percentages (header bounds, larger) are higher: see
the reference tile above.
## A. Geometric filters (current PDAL stack)
- **IGN** (vendor pre-classification, class 2): the fastest (~9 s). Reliable
where the vendor has confidence; holes under dense forest / steep relief.
No parameter to tune.
- **SMRF** — Pingel, Clarke & McBride 2013, *ISPRS J. Photogramm. Remote
Sens.* 77:21-30. **Raster-based** filter (operates on a DSM, not on points),
hence faster than point-based filters; **minimizes type I errors**
(ground omission) → well suited when ground is scarce (forest). Best
coverage of the three here (90.4%) but ~5.4 min/tile.
- **CSF** — Zhang et al. 2016, *Remote Sensing* 8(6):501. Inverted cloth
draped over the point cloud; simple, accurate, but **the cloth no longer
touches the ground on steep/hilly terrain** → poor classification. Slowest
here and worst on this tile. Best reserved for urban areas.
- **PTD/PTIN** (Progressive TIN Densification) — Axelsson 2000, ISPRS
Congress. **Literature's winner**: most robust on complex terrain +
forest (Moudrý et al. 2020, *Measurement* 150:107047; Cai et al. 2019,
*Remote Sensing* 11(9):1037) and the fastest (lidR benchmark: PTD ~20 s
vs CSF ~156 s vs PMF ~1800 s). **NOT available in the PDAL version
bundled in this image** (`filters.ground` / TIN missing) — would need to
be added to use it (or via lidR / a custom implementation).
## B. Fast hybrid (original plan, only partly kept)
> Only step 1 below is in use today. Step 2 (lowest-return floor) was
> dropped and step 3 (`_interpolate_holes`) is no longer called by the DTM
> builder: see "Synthesis / decision" for the current gap handling.
**PTD / Wack & Wimmer** principle (Wack & Wimmer 2002, *ISPRS Archives*
XXXIV/3A:293-296: DTM from lowest return, excluding the lowest 1% per cell
to discard outliers):
1. **Base = IGN pre-classification** (class 2, ~9 s, reliable and official).
2. **Measured gap filling**: for each cell with no ground point, take the
**robust lowest return** (min of the 99% of points in the cell) → adds
*measured* ground where the vendor failed (rock outcrops, clearings,
forest floor).
3. **Topographic inpainting** of the remaining gaps (terrain-aware
interpolation, `dtm.py:_interpolate_holes`, still present as a helper
but not called by `create_dtm_fast`).
Expected at the time: **continuous** DTM (0% holes), robust in
forest/relief, **~10-15 s/tile** instead of 326-355 s. No GPU dependency,
no training.
## C. AI / ML models (supervised — require labels)
Caveat (Qin et al. 2023, *ISPRS J. Photogramm. Remote Sens.*
202:246-261): **everything is supervised**; the main risk is
**generalization** — a model trained on one region degrades elsewhere.
No fully unsupervised DL filter published to date.
**Point-based (3D):**
| Model | Year | Architecture | Accuracy | Speed (~/km², GPU) |
|---|---|---|---|---|
| KPConv / RandLA-Net (Qin, OpenGF) | 2021 | KPConv / RandLA-Net | 97.8% OA, DTM RMSE 0.20 m, ground IoU 95% | 0.5-2.5 min |
| PFCN (Jin, *IEEE JSTARS* 13:3958) | 2020 | point-FCN | Te 1.73%, Kappa 93.9% | ~1/3 the cost of PointNet++ |
| Terrain-Net (Li, *Remote Sensing* 14(22):5798) | 2022 | KPConv + self-attention | OA 98%, mIoU 0.933 | parameter-free at transfer |
| MSVC (Štroner, *Remote Sensing* 17(4):615) | 2025 | 9x9x9 voxel DNN | beats CSF on F-score | — |
**Rasterized (directly output the DTM — closest to our need):**
| Model | Year | Architecture | Result |
|---|---|---|---|
| Precursor (Rizaldy, *ISPRS Annals* IV-2:231) | 2018 | 2D FCN | Te 5.22%, 78x faster |
| DeepTerRa / ALS2DTM (Lê, *IEEE JSTARS* 15:2778) | 2022 | GAN pix2pix (U-Net) | DTM RMSE < 1 m, filter + interpolation in one pass |
| DSM2DTM (Bittner, *ISPRS Annals* X-1/W1-2023:925) | 2023 | U-Net (EfficientNet) | non-ground mask + per-pixel ground height |
**Training datasets**: OpenGF (Qin et al., CVPRW 2021,
arXiv:2101.09641 — 47.7 km², 542 M pts); ALS2DTM (Lê et al. 2022,
arXiv:2206.03778 — 52 km², 1.66 billion pts, urban/forest/mountain).
**Costs / obstacles for our case**: (1) labels → to be generated as
pseudo-labels (high-quality SMRF/PTD output on a representative sample of
our tiles) or pre-training on OpenGF/ALS2DTM; (2) generalization across the
diverse terrain of LiDAR HD (plain/forest/mountain/urban); (3)
infrastructure: checkpoint + GPU inference path in the Docker image.
**Best fit if going the AI route**: a **rasterized U-Net (DSM2DTM-style)** —
rasterize the point cloud into multi-channel grids (elevation, slope,
curvature, density, return statistics), output a ground mask + ground
height. 2D = very fast and trivial to deploy on GPU, merges filtering +
interpolation. KPConv/RandLA-Net is more accurate in pure 3D but heavier to
deploy.
## Synthesis / decision
- "Fast" hard constraint + low maintenance → **fast hybrid (B)**
(~10-40 s/tile, zero training, zero GPU). ← **chosen approach, base
IMPLEMENTED** (step 1 only, see below)
- Base = IGN pre-classification (fast: ~5 s with the direct laspy
extraction, vs ~13.5 s through PDAL). `auto` prefers it as soon as
≥ 20% of points are classified as ground (threshold lowered from 30%
to 20%).
- **This gap-filling note is superseded**: gap filling in the DTM is no
longer a distance-based `fillnodata` pass over small holes. It is now
a morphological closing bounded to the point envelope
(`_fill_small_gaps` in `dtm.py`): the closing radius follows the local
point spacing (1.5 × the spacing measured over 5 m, staged at
1/1.5/2/3 m), nothing is extended beyond measured pixels, and islands
under 1 m² are removed. Large holes (dense forest, steep relief where
ground is under-classified) still remain as nodata (dark grey in the
oriented relief, hatched in the PDF export).
Deliberately no floor at the lowest return: under dense canopy that
return is vegetation, which would print trees into the DTM.
- Maximum quality in hard cases (steep + dense), ~1-2 min/tile + GPU +
training accepted → **rasterized U-Net (C)**. (not implemented)
- Best geometric filter available in PDAL → **SMRF (A)** (best coverage
90.4% but 5.4 min/tile). Selectable via `--ground-classification smrf`.
- Literature's absolute "winner" (fast + robust) → **PTD/PTIN
(A)**: to be integrated (not in the current PDAL stack).
## References
- Axelsson (2000), PTIN/PTD, ISPRS Congress.
- Pingel, Clarke & McBride (2013), SMRF, ISPRS J. P&RS 77:21-30.
- Zhang et al. (2016), CSF, Remote Sensing 8(6):501.
- Wack & Wimmer (2002), lowest-return DTM, ISPRS Archives XXXIV/3A.
- Moudrý et al. (2020), CSF/PTIN/PMF/SMRF comparison, Measurement 150:107047.
- Cai et al. (2019), CS+PTD, Remote Sensing 11(9):1037.
- Qin et al. (2021), OpenGF, CVPR Workshops (arXiv:2101.09641).
- Qin et al. (2023), dataset + evaluation + survey, ISPRS J. P&RS 202:246-261.
- Lê et al. (2022), DeepTerRa/ALS2DTM, IEEE JSTARS 15:2778 (arXiv:2206.03778).
- Bittner et al. (2023), DSM2DTM, ISPRS Annals X-1/W1-2023:925.
- lidR book (PTD/CSF/PMF comparison): https://r-lidar.github.io/lidRbook/gnd.html