Auto-detect best GPU with sm_120 skip: prefer RTX 5060 Ti, fall back to 4060 Ti when nvcc/CuPy does not support sm_120 yet

This commit is contained in:
Antoine Jacquin
2026-05-31 23:14:23 +02:00
parent 8490b5ddf0
commit 67d024a64e
2 changed files with 55 additions and 18 deletions

29
AGENTS.md Normal file
View File

@ -0,0 +1,29 @@
## Workflow
- install: `docker build -t lidar-lidar .` (deps baked into image)
- build: `docker build -t lidar-lidar .`
- test all: `./run.sh --test`
- test file: `docker run --rm lidar-lidar python3 -m pytest -v --pyargs lidar_pipeline.tests.<module>`
- test case: `docker run --rm lidar-lidar python3 -m pytest -v --pyargs lidar_pipeline.tests.<module>::<TestClass>::<test_method>`
- lint: not configured
- format: not configured
- after every edit: `./run.sh --test`
- debug: `./run.sh --debug` (file:line logging); container shell: `docker run --rm -it -v $(pwd)/input:/data/input -v $(pwd)/output:/data/output --entrypoint bash lidar-lidar`
## Conventions
- **Bilingual naming**: all code identifiers are English; every user-facing string, log message, argparse help, and comment is French.
- **Adding a visualization requires 3 edits**: (1) `generate_X()` in `visualizations.py`, (2) entry in `VIZ_STEPS` in `pipeline.py`, (3) entry in `COLORMAPS` in `rendering.py`. Missing any one breaks the pipeline.
- **`generate_*` signature is strict**: `(dem_file, basename, vis_dir, resolution, shared=None)` returning `Path` on success, `None` on failure. IGN overlays (`ortho`, `topo`) omit `shared`.
- **Return `None` on failure, never raise**: `dtm.py`, `visualizations.py`, and `ign.py` all return `None` to let the pipeline continue. Raising aborts the entire file.
- **Logger is always `logging.getLogger("lidar")`**, never `__name__`. All modules route through this single logger so worker processes can configure it.
- **Filename special-cases** in `_expected_output_path()`: `pos_open` → `positive_openness`, `neg_open` → `negative_openness`, `hillshade` → `hillshade_multi`.
- **Default output is AVIF**, not WebP. Use `--format webp` for WebP. Quality default is 98.
- **Tests use lazy imports inside each test function**, never at module top, to avoid importing CuPy/GDAL at import time.
- **`_`-prefixed names are critical private**: `_create_ground_pipeline`, `_fallback_to_smrf`, `_fill_nans`, `_init_gpu`, `_process_file_standalone` — do not call from outside their module.
## Commit & Pull Request Guidelines
Commits use imperative tense, short single-line subjects (~60–80 chars), no prefixes or scopes. Compound commits are common — multiple related changes joined by commas or "and". Examples: ``Fix multi-GPU with lazy CuPy init + rendering improvements``, ``Add multi-resolution support and remove PDF generation``, ``Fix corrupted COPC detection, add CSF→SMRF fallback, improve MSRM colormap, add SVF and anisotropic openness``.
No PR template, no CI pipeline, no issue tracker. This is a standalone Docker project with no formal PR process.

View File

@ -50,7 +50,7 @@ def _pick_gpu() -> int | None:
mem_mi = int(parts[3]) mem_mi = int(parts[3])
major, minor = (int(x) for x in cap_str.split('.')) major, minor = (int(x) for x in cap_str.split('.'))
score = major * 1000 + minor * 100 + mem_mi score = major * 1000 + minor * 100 + mem_mi
gpus.append((idx, name, cap_str, mem_mi, score)) gpus.append((idx, name, cap_str, mem_mi, score, major))
if not gpus: if not gpus:
return None return None
@ -58,13 +58,21 @@ def _pick_gpu() -> int | None:
_NUM_GPUS = len(gpus) _NUM_GPUS = len(gpus)
gpus.sort(key=lambda g: g[4], reverse=True) gpus.sort(key=lambda g: g[4], reverse=True)
# Use the best GPU (highest compute capability) # Try GPUs in order of capability.
best = gpus[0] # CuPy 13.4 + CUDA 11.8 JIT works for sm_89 (RTX 40xx).
_best_gpu_id = best[0] # sm_120 (RTX 50xx) is NOT supported by any nvcc yet (CUDA ≤ 12.9).
_gpu_name = best[1] for gpu in gpus:
_gpu_mem_gb = best[3] // 1024 idx, name, cap_str, mem_mi, score, major = gpu
HAS_GPU = True if major >= 12:
return _best_gpu_id logger.debug(f" GPU {idx}: {name} (sm_{cap_str}) — non supporté par nvcc/CuPy")
continue
_best_gpu_id = idx
_gpu_name = name
_gpu_mem_gb = mem_mi // 1024
HAS_GPU = True
return _best_gpu_id
_gpu_reason = "aucun GPU compatible (sm_120+ non supporté par nvcc/CuPy)"
except (FileNotFoundError, subprocess.TimeoutExpired, Exception): except (FileNotFoundError, subprocess.TimeoutExpired, Exception):
return None return None
@ -102,19 +110,19 @@ def _init_gpu():
return return
try: try:
# MUST set CUDA_VISIBLE_DEVICES before importing CuPy.
# If we don't, CuPy creates its CUDA context on device 0 (sm_89)
# and pre-compiled kernels won't work on device 1 (sm_120).
os.environ['CUDA_VISIBLE_DEVICES'] = str(_best_gpu_id)
import cupy as _real_cupy import cupy as _real_cupy
import cupyx.scipy.ndimage as _real_cupy_ndimage import cupyx.scipy.ndimage as _real_cupy_ndimage
# Select the target GPU using Device API # Warm-up: verify kernel execution works on this GPU
with _real_cupy.cuda.Device(_best_gpu_id): _test = _real_cupy.array([1.0, 2.0, 3.0], dtype=_real_cupy.float32)
# Warm-up: verify kernel execution works on this GPU _result = _real_cupy.sum(_test * _test)
_test = _real_cupy.array([1.0, 2.0, 3.0], dtype=_real_cupy.float32) _ = _result.get()
_result = _real_cupy.sum(_test * _test) del _test, _result
_ = _result.get()
del _test, _result
# Make this device the default for all subsequent operations
_real_cupy.cuda.Device(_best_gpu_id).use()
_xp = _real_cupy _xp = _real_cupy
_cp = _real_cupy _cp = _real_cupy