Files
lidar_rendu/AGENTS.md
Jacquin Antoine d0dc8d90e9 Make the worker batch timeout configurable via LIDAR_BATCH_TIMEOUT
The pool of tile workers used to cancel every remaining tile after a
hardcoded 2-hour wall clock, silently truncating large batches (a
670-tile completion run lost its last 348 tiles that way). The timeout
now defaults to unlimited and can be capped per deployment with the
LIDAR_BATCH_TIMEOUT environment variable (seconds); the local worker
compose sets it to 6 hours.

💘 Generated with Crush

Assisted-by: Crush:glm-5.2
2026-09-28 00:03:52 +02:00

30 KiB
Raw Permalink Blame History

Workflow

  • install: docker build -t lidar-lidar . (deps baked into the image)
  • one-command local start (GPU optional): ./start.sh (map on 8973), ./start.sh process [options], ./start.sh logs|stop — creates input//output/ as the current uid, adds docker-compose.cpu.yml (gpus: !reset) when Docker cannot reach any GPU.
  • build: docker build -t lidar-lidar .
  • build the lightweight map (Raspberry Pi, two-machine deployment — see docs/DEPLOY_WEBAPP.md): docker compose -f docker-compose.maps.yml up -d --build (image Dockerfile.maps, no PDAL/GPU)
  • build the tile generator (processing machine): docker compose -f docker-compose.worker.yml up -d --build (service worker = mapserve on the full image: API + tiles + inventory for the remote maps)
  • test all: ./run.sh --test (rebuilds the image automatically before the tests; with a direct docker run, rebuild by hand first)
  • test file: docker run --rm lidar-lidar python3 -m pytest -v --pyargs lidar_pipeline.tests.<module>
  • test case: docker run --rm lidar-lidar python3 -m pytest -v --pyargs lidar_pipeline.tests.<module>::<TestClass>::<test_method>
  • lint: not configured
  • format: not configured
  • after every edit: ./run.sh --test
  • RULE 1 — always run through docker compose (never a direct docker run): local map/API → docker compose up -d --build serve (port 8973, mapserve on the full image); one-off processing → docker compose run --rm --build process [options]; logs → docker compose logs -f serve; stop → docker compose down.
  • RULE 2 — ALWAYS --build: the code is baked into the image (never mounted). Without --build, up/run reuse the existing image and the OLD code runs. --build is almost instant thanks to the cache (.dockerignore keeps input/ and output/ out of the context). After an edit, docker compose up -d --build serve recreates the container on fresh code.
  • quick test without rebuild (code mounted over the image): docker run --rm -e PYTHONPATH=/app -v $(pwd)/lidar_pipeline:/app/lidar_pipeline lidar-lidar python3 -m pytest --pyargs lidar_pipeline.tests -q (~3 min; prefix with timeout 600, and NO | tail pipe, which hides the progress)
  • debug: ./run.sh --debug (file:line logging); container shell: docker run --rm -it -v $(pwd)/input:/data/input -v $(pwd)/output:/data/output --entrypoint bash lidar-lidar
  • ./run.sh passes its own defaults to the CLI: -r 0.5 and -w 1 unless given (the CLI alone defaults to -r 0.2, -w auto); image quality, format and ground classification fall through to the CLI defaults (60, avif, ign).
  • updating a remote lightweight map (git checkout + local maps override, see docker-compose.maps.override.yml.example): ssh <host> "cd <checkout> && git pull && docker compose -f docker-compose.maps.yml -f docker-compose.maps.override.yml up -d --build" (procedure in docs/DEPLOY_WEBAPP.md).
  • backfilling the quality sidecars (tiles processed before the sidecar existed): the fixed command of process (compose) is replaced entirely as soon as any argument follows the service name, so --quality-backfill alone does not work — use docker compose run --rm --build process python3 -m lidar_pipeline /data/input -o /data/output --quality-backfill.

Conventions

  • A single web server: mapserve.py (image lidar-maps, port 8975 lightweight / 8973 worker). The old webapp (webapp.py, export.py, index.html/_APP_JS) has been removed: tile generation (the scope of the historical webapp — /api/preview, /api/generate, /api/status, /api/stop, /api/queue/clear, /api/cell) lives in mapserve.py, the interface in web/map.{html,css,js} (read by mapui.py, _MAP_* constants kept, written as app.css/app.js by write_map_assets() and baked into the images). On the lightweight image without LIDAR_GENERATION_URL, /api/status answers available: false and the interface hides the Generation tab.
  • Single-panel interface (web/map.{html,css,js}): one tabbed panel View / Tile / PDF export / Generation (hidden when the generator is unavailable or not allowed) / Share, replacing the old stack of blocks. Below 720 px wide, the panel becomes a bottom sheet with three heights (closed/half/full; handle #sheetGrip dragged or simply tapped, or the active tab tapped again). A click on a tile selects it (outline) and fills the Tile tab (extent, IGN, strip alignment) without switching the displayed tab; clicking the same tile again deselects it, clicking another moves the selection. Escape deselects the tile, then collapses the panel (icon strip on desktop, closed sheet on phone — the strip stays usable, a click on a tab expands the panel again). Keyboard shortcuts: 1–5 (tabs, no effect when the tab is hidden), P (next display mode), Escape. Relief intensity (&I=, contrast 0.5–2×, default 1×): a personal comfort setting stored in lidarMapView_v2.intensity and shared in the link, never frozen as a server default (★ Set as default ignores it). Two themes, light/dark (lidar-theme, auto by default, follows the system), and every shared setting set on :root as CSS variables (no hard-coded colour in the components). localStorage keys (through lsGet/lsSet, silent in private browsing): lidarMapView_v2 (view/layers), lidar-print (export settings), lidar-panel (tab, sheet height, collapse to the icon strip), lidar-theme.
  • index.py = catalogue + shared registries (no interface any more): VIZ_LABELS/VIZ_LEGENDS, display defaults (DEFAULT_VIZ/PRECISION_VIZ/VIEW_MODES), PANEL_VIZ/KEYWORD_TO_STEP, scan_tiles/cells_with_all_viz, thumbnails + sub-tiles + the index_tiles.json inventory (build_index). The inventory is served by mapserve's /api/tiles to the lightweight machines (LIDAR_SOURCE_URL).
  • Generation is 0.2 m only (policy): /api/generate (GENERATE_RESOLUTIONS in mapserve.py), the compose process command and the CLI -r default all produce 0.2 m exclusively; 0.5 m stays available via an explicit -r 0.5 (note: ./run.sh still passes -r 0.5 by default). Completeness detection (complete_cells) requires the viz at 0.2 m only.
  • North-to-south generation: tiles are processed by decreasing row (row = northing in km), then increasing column — find_laz_files (pipeline.py) for batch passes and _resolve_request (mapserve.py) for runs started from the map. Workers take the files in submission order: the map fills from top to bottom during a run (an explicit --file on the CLI keeps the user's order). Generation parallelism: LIDAR_WORKERS (10 in docker-compose.yml, auto in docker-compose.worker.yml and when unset).
  • Full sub-tiling: _CARTO_SUBTILED_VIZ (empty in index.py) splits ALL layers into 500 m quadrants at 0.2 m; all in AVIF 4:2:0 q75 (_SUBTILE_AVIF_QUALITY), encoded ONCE from the original raster: tif_to_crop calls index.write_subtiles (through rendering._write_subtiles_from) right after the tile (so they are newer than it, and build_index does not re-encode them; the old chain tile q60 → sub-tile q55 stacked two losses: 18.1 dB/SSIM 0.84 against 19.5 dB/0.93). 4:2:0 tops out around 19.5 dB on the relief (per-pixel hue averaged over 2 × 2): only 4:4:4 would go further (q75: 28 dB, ~1.6× heavier). Exception: flat level maps (_SUBTILE_LOSSLESS_VIZ, the precision layer) in lossless WebP (.webp, accepted by tiles.py). Pillow ignores lossless=True for AVIF (q75, lossy); in RGB no AVIF setting is exact. A layer that fails to split falls back to the whole tile (_fallback_full_dalle) without penalising the others.
  • XYZ tiles reusable outside the project: /tiles/{layer}/{z}/{x}/{y}.png follows the OpenStreetMap scheme (256 px, EPSG:3857, RGBA PNG transparent outside coverage, CORS *) — JOSM, iD, QGIS, uMap and MapLibre consume it as is, with discovery through TileJSON / WMTS / josm.imagery.xml. The 512 px variant (@2x.avif) is reserved for the internal interface (half as many requests over HTTP/1.1). Any change to the URL template breaks client configurations: changing it requires an explicit decision.
  • Naming: all code identifiers, comments, logs, help and user-facing strings are English; data keys inherited from the French era (relief_oriente, densite_sol, visualisations/, --ign-classes sol, stats key telechargees…) are stable contracts and must not be renamed.
  • Adding a visualization requires 4 edits: (1) generate_X() in visualizations.py, (2) entry in VIZ_STEPS in pipeline.py, (3) entry in COLORMAPS in rendering.py (or RGB_LEGENDS for an RGB output), (4) entry in VIZ_LEGENDS in index.py (title/legend/description + sampled cmap gradient — single text source merged into COLORMAPS at import, also served in /api/map/meta and the TileJSON). Missing any one breaks the pipeline.
  • generate_* signature is strict: (dem_file, basename, vis_dir, resolution, shared=None) returning Path on success, None on failure. IGN overlays (ortho, topo) omit shared.
  • Return None on failure, never raise: dtm.py, visualizations.py, and ign.py all return None to let the pipeline continue. Raising aborts the entire file.
  • Logger is always logging.getLogger("lidar"), never __name__. All modules route through this single logger so worker processes can configure it.
  • Joint scan-line adjustment (3rd pass of strip alignment, STRIP_ALIGN_VERSION 3): each scan line (~150 lines/s, split by _scan_line_ids: flyback of the scan_angle sawtooth) can be offset AND tilted (roll) by 1 to 15 cm — isolated sunken lines, a whole pass tilted and visible as a band (measured on LHD_FXX_0999_6882). _joint_line_corrections (dtm.py) re-fits each line (a + b·u) against the consensus of the OTHER strips in its 1 m cell (local plane through the slope), least squares trimmed at 5 MAD via bincount, damped step 0.5, stop below 1 mm, re-centring (no global drift), estimated on 1 point in 3; CuPy when the GPU is active, numpy fallback. On top of that, a calibration profile per strip and per degree of scan angle (STRIP_ANGLE_BIN, shared by all lines of the strip, mean removed, slope kept because the per-line tilt is undetermined when the swath only partly crosses the tile; sparse bins = neighbouring value) removes the non-linear across-swath errors (±1.5 cm measured at the edge), a source of steps parallel to the flight line at the swath edge. Lines without overlap: _scan_line_corrections_beam (against the surface of their own strip, line-to-line component only, σ 3 lines). Validated on never-seen blocks (10 m checkerboard). Jitter by time windows (2nd pass) is no longer computed when scan_angle exists (redundant). CPU cost ~26 s per tile for the whole alignment (73 s before). Sidecar: lines, line_window, line_cell, line_model.
  • Filling gaps between points, bounded by the envelope (GAP_FILL_VERSION 3, _fill_small_gaps in dtm.py): at 0.2 m ~80 % of the pixels contain no point. The old fillnodata at 1 m filled every pixel within 1 m of a point — each isolated point became a flat 2 m disc (hue atan2(0,0) = saturated pink in the oriented relief) and each hole received an extrapolated 1 m band (coloured fringe). Now: morphological closing of the measured pixels (nothing is extended outwards) whose radius follows the local point spacing (GAP_RADIUS_K 1.5 × spacing measured over 5 m, steps GAP_RADII_M 1/1.5/2/3 m, 1 m minimum to plug car-sized holes), then islands < GAP_MIN_ISLAND_M2 (1 m²) removed. Morphology with numpy slices (_morph_step, ~10× scipy); ~5–9 s per 5000² tile. Version stored in the GeoTIFF tag LIDAR_GAP_FILL (2 = envelope, 3 = + density sidecar file): absent/different ⇒ DTM regenerated (_gap_fill_matches in pipeline.py). rasterio.fill.fillnodata writes into its input: always pass it a copy.
  • Fixed-scale openness: generate_openness normalises with frozen references OPENNESS_POS_REF / OPENNESS_NEG_REF (degrees, medians measured on 15 tiles spread over the territory) instead of a per-tile z-score — same openness = same colour, seamless mosaic.
  • Filename special-cases in _expected_output_path(): pos_open → positive_openness, neg_open → negative_openness, hillshade → hillshade_multi.
  • Imposed generation settings: the map no longer offers a classification or stitching setting — _build_command (mapserve.py) always passes --ground-classification ign --ign-classes sol --edge-buffer 100; any ground_class/ign_classes/reclassify/edge_buffer fields sent are ignored. CLI defaults aligned (ign, 100).
  • IGN classification by direct extraction: _extract_ign_ground (dtm.py) filters the classes with laspy (returns ≥ 1) and writes the ground LAS without PDAL (4.9 s instead of 13.5 s per tile); the read done by the automatic detection is reused (_LAST_READ). PDAL remains the fallback.
  • Pyramid updated during rendering: the pipeline rebuilds the inventory after each tile by default (incremental_index unless --no-index, whatever the launcher), with a deferred pass when the 3 s debounce skips a tile. On the map side, _bg_poller watches the inventory every BG_WATCH_S = 10 s (local file or upstream payload) and starts a scan as soon as it changes; tiles of modified source tiles go to the front of the queue (_bg_enqueue(front=True)), ahead of the backlog.
  • Full pyramid generated ahead of time: by default tiles.zoom_cached stores ALL levels up to native (TILE_CACHE_MAX_Z = TILE_MAX_NATIVE_Z 19, TILE_EVEN_LEVELS off) and background maintenance — enabled with LIDAR_TILE_BACKGROUND=1, set in the Pi override — generates them ahead of time up to 18@2x (default of LIDAR_TILE_BACKGROUND_MAX_Z); on the Pi it DOWNLOADS them from the worker. Map capped at zoom 19 (maxZoom, 1 screen px = 1 LiDAR px, never an upscaled tile). Reason: on-the-fly rendering on the Pi (~170 ms/tile, 3 at a time) gave 1.5–6 s per screen at zooms 17–19. Bounded maintenance queue (queue_max): each source tile keeps in _bg["incomplete"] the index of its first tile rejected for lack of room and resumes from there on the next scans (before: the first scan filled the queue and the rest of the pyramid was never generated). ~110,000 @2x tiles for 3,240 source tiles (~5 GB). Reduced storage still possible (LIDAR_TILE_EVEN_LEVELS=1 → EvenLevelTileLayer, even levels only, native 19 requested as is).
  • Bounded map memory (Pi, mem_limit: 1g): _open_source (tiles.py) caches a detached copy of each source (an open AVIF image keeps its decoder, ~18 MB more per 2500² quadrant) and counts 4 bytes/pixel (PIL stores RGB on 32 bits); Dockerfile.maps sets MALLOC_MMAP_THRESHOLD_=1048576 so that glibc returns large buffers to the system. Without these two points, browsing at high zoom (levels rendered on the fly) climbed to ~700 MB RSS and the container was killed in a loop by the OOM killer (502 on the Traefik side).
  • Standalone map (Pi without worker): upstream (LIDAR_SOURCE_URL/LIDAR_MAPS_URL) down ⇒ the map is still served from disk. _remote_payload (tiles.py) also remembers the failure (TTL 60 s, timeout 10 s: otherwise every request paid the network timeout again and /api/map/meta stopped answering), then the local index (_build_index) takes over; circuit breakers on sources (_SOURCE_OFFLINE, 60 s) and tiles (_UPSTREAM, re-armed on expiry); a fetched source is dated to its upstream version (os.utime) so the local index keeps the same dates; no upstream request for a tile without a source tile when the inventory comes from upstream. In cache-only mode, a stale tile keeps being served (no-store, X-Tile-Pending) until the new one arrives.
  • Pyramid downloaded in the background: with LIDAR_MAPS_URL, maintenance (_bg_process_one) DOWNLOADS the tiles of the stored levels from upstream (stat telechargees), without waiting for them to be displayed, and with no pause between two downloads (the worker copes); local rendering as a fallback when upstream does not answer, which alone is followed by LIDAR_TILE_BACKGROUND_PAUSE (Pi resources).
  • First appearance of source tiles (_apply_seen, index_xyz/.sources_seen.json): effective date of a source = max(version, first entry in the inventory). A source tile written before but inventoried after an XYZ tile (upstream index TTL, inventory debounce) therefore invalidates that tile — otherwise a permanent hole at that level, on disk, in memory and in the browser (unchanged ?v= stamp). Persistent registry; absent with an existing cache (upgrade) ⇒ everything is invalidated once.
  • Browser cost of the map (measured in Chrome, Intel Iris Xe): the darkening filter of the OSM base map (.base-dark) is set on the layer container, never on each tile — identical rendering (per-pixel operations), GPU process 92 % → 56 % while panning and 99 % → 53 % on wheel zoom. LiDAR layers use keepBuffer: 1 (512 px tiles = 1 MB decoded each). Measured with no notable effect: the cards' backdrop-filter, isolation/mix-blend-mode set to "normal", will-change. JS heap 2–5 MB, idle CPU ~1 %. AVIF decoding ~12 ms/tile against ~5 ms for WebP (accepted trade-off: AVIF storage).
  • Fast AVIF encoding: AVIF_SPEED = 9 (rendering.py, tiles.py, _SUBTILE_AVIF_SPEED in index.py) — a 5000 × 5000 px tile encoded in 0.6 s instead of 4 s (+3 % size, −0.3 dB); encoding was the longest step of rendering a layer.
  • Workers bounded by VRAM: each pool process takes a fixed slot when it is created (_init_worker_slot in pipeline.py, multiprocessing queue) computed by gpu_worker_slots (gpu.py): at most (free VRAM − LIDAR_GPU_RESERVE_MIB 512) / LIDAR_GPU_WORKER_MIB (2048, estimated peak: CUDA context + joint CuPy alignment in float64 + EDT of _fill_nans) workers per GPU, at least one, interleaved; the excess runs on the CPU (force_cpu). The old round-robin by file number put 6 workers on each 8 GB RTX 5060 with LIDAR_WORKERS=auto (12): OOM. Unreadable VRAM (nvidia-smi failing): unbounded round-robin.
  • Preparation on the GPU: _fill_nans (cupyx distance transform) and the SharedDEM gradients run on the GPU when it is active (CPU fallback). NUMBA_CACHE_DIR=/tmp/numba-cache (Dockerfile) avoids recompiling the numba kernels in every worker.
  • Default output is AVIF, not WebP. Use --format webp for WebP. Quality default is 60 (visually lossless on smooth colour ramps, ~÷3 vs q98).
  • Vertical alignment of flight strips: each tile mixes several passes (1–2 PointSourceId per pass), sometimes vertically biased by a few cm (±2.5 cm measured on 1000_6882). create_dtm_fast measures the robust offset of each strip (ground points, 1 m cells, median surface iterated 3×) and subtracts offsets ≥ 0.5 cm (STRIP_ALIGN_THRESHOLD in dtm.py) before rasterization. Offsets computed per tile (they drift along a flight line: never a global table), memoised per ground LAS, recorded in DTM/*_dtm*_stripalign.json (cache sidecar: absent, or different version/threshold/parameters ⇒ DTM regenerated). Can be disabled: --no-strip-align.
  • Intra-strip jitter (2nd pass of the alignment): successive scan lines of the SAME pass can be randomly offset vertically (sensor vibration / high-frequency trajectory noise) — a constant offset per strip is not enough. _strip_jitter_offsets splits each strip into GPS time windows (STRIP_JITTER_BIN = 0.1 s, time origin specific to each strip), measures the robust offset of each window against the median surface of the OTHER strips (shared 1 m cells, ≥ STRIP_JITTER_MIN_CELLS = 40 cells), smooths the series (rolling median STRIP_JITTER_SMOOTH = 5 windows), clamps it to ± STRIP_JITTER_MAX (10 cm), then interpolates it at each point's GPS time (_apply_strip_jitter); where two strips overlap, each gets a series (each absorbs its share). Requires the gps_time dimension (silently skipped otherwise). Sidecar version 2 (series in jitter), covered by --no-strip-align.
  • Layers produced and served = PANEL_VIZ (index.py, currently ('relief_oriente', 'densite_sol')): pipeline without --only/--skip (panel_steps()), generation started from the map (_panel_viz_steps in mapserve.py) and served layers (tiles.available_layers filters: panel, /tiles/…, TileJSON, WMTS, JOSM). The other visualizations can still be computed with --only. Map tests lift the restriction via _setup(..., panel=None).
  • Precision (densite_sol) and map display: create_dtm_fast writes the density of the kept ground points (stitching band included) to DTM/*_dtm*_density.tif (1 m cells, 3 × 3 mean, pts/m²); generate_densite_sol quantises it into 16 levels (density_levels: fixed log scale, level k from 0.25 × 2^(k/2) pts/m²) and keeps it at 1 m (1000² tile, ~200 KB in lossless WebP, rendering.LOSSLESS_GRAY_KEYWORDS); nearest-neighbour tiles (tiles.NEAREST_LAYERS) to keep 16 crisp greys at zooms 18–19. Interface (web/map.js, View tab): no more layer stack — one main layer (DEFAULT_VIZ) and three modes relief / precision / compare (VIEW_MODES, sliding bar, link &P=compare&C=left:right:%); both (old product) read as relief. "How to read" text: the reading field of VIZ_LEGENDS (map + PDF).
  • Oriented relief (relief_oriente, the only relief layer of the map): a single RGB image (3-band uint8 GeoTIFF, rendered as is like ortho/topo via RGB_KEYWORDS in rendering.py) — CIELAB lightness = local positive openness (DTM − Gaussian RELIEF_DETREND_M = 10 m, radii RELIEF_RADII_M = 5/10/20 m, 16 directions) 65 % + shading 35 %; hue = aspect, fixed chroma (RELIEF_CHROMA). Fixed log scale RELIEF_OPEN_RANGE (no per-tile statistic) and a total support of ~60 m (20 m maximum ray radius + 4σ = 40 m of the 10 m Gaussian detrend), still < the 100 m stitching band: seamless tiles. Fast: detrend + radii on a decimated grid ~0.8 m (RELIEF_GRID_M), a dedicated kernel that only accumulates the mean of the angles (_mean_horizon_*: CuPy RawKernel on GPU, parallel numba on CPU, numpy fallback), fused colourisation (numba) or vectorised without trigonometry (CuPy) through an L* × hue table (_relief_lut). ~5 s of computation excluding preparation on a 12-core CPU. Any constant change changes the rendering: regenerate the tiles (--only relief_oriente --force).
  • Downsampled openness: generate_openness computes the ray tracing (the most expensive step: 532 s/tile at 0.2 m on CPU) on a block-decimated grid (OPENNESS_DOWNSAMPLE = 2: block max for positive, min for negative — preserves the reliefs that bound the horizon) then resamples bilinearly. Cost ÷ factor³: 532 s → 40 s (×13). Archaeological signal preserved (corr. 0.93 after smoothing); the sub-metre noise texture disappears. --openness-downsample 1 = full resolution. SVF is NOT affected (anisotropic openness no longer exists).
  • Edge stitching between tiles: large-kernel renderings (openness/SVF: 100 m radii; LRM: 15 m) truncate their window at the tile edge — bands of artefacts at every tile change. --edge-buffer N (default 100; always applied by generation from the map, EDGE_BUFFER_METERS = 100 m in mapserve.py, no checkbox any more; 0 = off on the command line) rasterizes the DTM over the nominal 1 km tile aligned on the grid plus an N m band filled with the ground points of the 8 neighbouring LAZ files (_neighbor_ground_points in dtm.py: streaming PDAL, crop + IGN class filter; missing neighbour = automatic download from the IGN catalogue before the run, kept apart in input/edge_neighbors/ so as not to swell the corpus of global passes (deduplicated over the whole batch, _fetch_edge_neighbors in pipeline.py); not found or failure = empty band). Visualizations compute on the extended extent, then rendering.py (_core_tile_window, via tif_to_crop/tif_to_png) crops the outputs to the exact 1 km tile read from the LHD name — the AVIF images stay aligned 1 km squares in the mosaic. Buffer recorded in the GeoTIFF tag LIDAR_EDGE_BUFFER of the DTM: changing --edge-buffer invalidates the DTM cache automatically (tag absent = 0). Neighbouring bands are not strip-aligned (context only, cropped out of the final image). Cost: ~7 s of reading per neighbour + ~44 % more pixels at 100 m/0.2 m. Name outside the LHD pattern: option ignored (header bounds, no cropping).
  • Tests use lazy imports inside each test function, never at module top, to avoid importing CuPy/GDAL at import time.
  • _-prefixed names are critical private: _create_ground_pipeline, _fallback_to_smrf, _fill_nans, _init_gpu, _process_file_standalone — do not call from outside their module.
  • build_index() writes the inventory + the source levels: output/index_tiles.json (source tiles, layers, versioned URLs — served by /api/tiles), thumbnails index_thumbs/ (≈3.9 m/px + _mid 1.56 m/px) and quadrants index_subtiles/ — levels of the XYZ pyramid (tiles.py). Each tile of the current run carries its WGS84 corners for the progress frames. No HTML any more: the interface lives in mapui.py.
  • PDF export (export_pdf.py): Pillow + pyproj + reportlab, no numpy — runs in the lightweight image alone. /api/export/frame (frame) and /api/export/pdf (sheet) in mapserve.py, never delegated to LIDAR_GENERATION_URL, single export lock (_export_lock, 429 when busy). Composed directly in L93 (tiles.sources_in_bbox/load_source, crop + Lanczos resampling), without going through the XYZ tiles. Sheet: L93 grid (step depends on the scale), WGS84 corners, north arrow accounting for meridian convergence, graphic and numeric scale, 8-point orientation rose, legend text taken from VIZ_LEGENDS, quality inset (white hatching = "not recorded", a "Missing data (no relief)" line for tiles outside coverage), 300 dpi in A4 / 250 dpi in A3 (bounds the Pi's memory). NODATA_RGB copies visualizations.RELIEF_NODATA_RGB (numpy is absent from the lightweight image); outside the tiles' extent = hatched white.
  • Quality sidecar (quality.py): output/quality/{basename}.json (QUALITY_VERSION), ground density per 50 m cell and acquisition dates, computed after the DTM (ensure_quality, including on an already cached DTM — it then measures the input LAZ directly); batch backfill --quality-backfill (see Workflow); copied into the quality table of index_tiles.json (build_index) and onto the lightweight machines (tiles._persist_remote_quality, from the upstream inventory).

Architecture Notes (from code audit 2025-09)

Module structure & data flow

  • cli.py → pipeline.py (LidarArchaeoPipeline) → per-file: dtm.py (classify + rasterize) → visualizations.py (17 products) → rendering.py (GeoTIFF→AVIF) → index.py (catalogue: thumbnails + inventory)
  • gpu.py provides CuPy/NumPy proxy (xp), lazy init, OOM fallback. safe_gpu_call wraps all non-IGN viz calls.
  • mapserve.py (FastAPI, image lidar-maps) serves the XYZ pyramid (tiles.py), the interface (mapui.py), the /api/tiles inventory + static source tiles for the lightweight machines, and the generation API: pipeline as a subprocess on the full image, delegation via LIDAR_GENERATION_URL on the lightweight image (Raspberry Pi). Single job + persistent queue (_job_lock is an RLock: /api/status calls _queue_summary() under the lock). LIDAR_API_TOKEN (worker) / LIDAR_REMOTE_TOKEN (lightweight) protect the calls; LIDAR_REGEN_CIDR restricts generation to the local network.
  • progress.py writes JSONL events (O_APPEND, atomic) read by mapserve for live progress (frames on the map).

Key design decisions (intentional, do not "fix")

  • _res_suffix hardcodes 0.5 as "no suffix": coupled to index.py parsing (_strip_res_suffix defaults to 0.5 when no suffix). Changing requires sidecar metadata.
  • GPU scoring (major*1000 + minor*100 + mem_mi): compute capability priority is intentional — a newer GPU with less VRAM is preferred.
  • mapserve.py reads env at import time: deployment-focused single-purpose server, env is set once in docker-compose.
  • Repeated try/except in visualizations (same pattern in every generate_*): intentional convention for uniform return None behavior.
  • _d8_accumulate_numba defines @njit inside the function: cache=True makes subsequent calls fast; the Python function object creation is negligible.
  • pkill -9 -f "pdal pipeline <output>/temp/" in the cli.py signal handler: belt-and-suspenders alongside os.killpg. Scoped to this run's temp directory to avoid killing unrelated PDAL processes (e.g. a concurrent generation job).

Performance characteristics

  • _priority_flood uses numba JIT binary heap (single int64 array, flat view for elevation). Python heapq fallback if numba unavailable.
  • _d8_accumulate_numba uses numba with argsort top-down sweep. Python fallback exists.
  • Ray-tracing (SVF, openness): processes one direction at a time to limit VRAM. Auto-falls back to CPU on OOM via _ray_trace_horizons.
  • Multi-resolution: 0.5 m has no filename suffix whatever its position; every other resolution (including the 0.2 m default) uses a _r0p2 style suffix. Ground classification done once, shared across resolutions.
  • ProcessPoolExecutor batch runs are unlimited by default; set LIDAR_BATCH_TIMEOUT (seconds) to cap a batch's wall clock and cancel the remaining workers past it.

Numba usage pattern

  • Defined at function scope with @njit(cache=True) — first call compiles (~2-3s), subsequent calls hit disk cache.
  • Must use flat 1D array views (arr.ravel()) for integer indexing — 2D arrays with a single int index return a row slice in nopython mode.
  • Pattern: try numba → return None on ImportError → caller falls back to pure Python.

Commit & Pull Request Guidelines

Commits use imperative tense, short single-line subjects (~60–80 chars), no prefixes or scopes. Compound commits are common — multiple related changes joined by commas or "and". Examples: Fix multi-GPU with lazy CuPy init + rendering improvements, Add multi-resolution support and remove PDF generation, Fix corrupted COPC detection, add CSF→SMRF fallback, improve MSRM colormap, add SVF and anisotropic openness.

No PR template, no CI pipeline, no issue tracker. This is a standalone Docker project with no formal PR process.