GPU multi-processing fix:
- gpu.py: revert to CUDA_VISIBLE_DEVICES approach with lazy CuPy init
(Device.use() caused CUDA_ERROR_NO_BINARY_FOR_GPU on GPU 1)
- CuPy is imported lazily on first to_gpu() call, allowing
CUDA_VISIBLE_DEVICES to be set before CUDA context creation
- nvidia-smi used for GPU count detection (no CUDA import needed)
- pipeline.py: add tip message suggesting -w N when multiple GPUs detected
Rendering improvements:
- Title: split into bold title (14pt) + italic description (10pt)
- North arrow: moved inside data area (top-right) with transparent
background — no longer overlaps title
- Colorbar: full height (compass gap removed), ScalarFormatter with
useOffset=False to prevent scientific notation on small values
Performance:
- rendering.py: save matplotlib figure to BytesIO instead of temp PNG
file — eliminates disk I/O between matplotlib and PIL
- visualizations.py: cap max_dist at 300 for ray-tracing (SVF,
openness, aniso_open) — avoids 500+ iterations at 0.2m resolution
- pipeline.py: deduplicate n_gpus calculation in parallel path