--- library_name: pytorch tags: - point-cloud - semantic-segmentation - point-transformer-v3 - lidar - 3d datasets: - RPTU-FGMB/DeKH license: cc-by-nc-sa-4.0 --- # BIMStruct3D-segmentation Point-cloud semantic segmentation model (PT-v3m1) from the [BIMStruct3D](#citation) scan-to-BIM pipeline. [![License: CC BY-NC-SA 4.0](https://img.shields.io/badge/License-CC%20BY--NC--SA%204.0-lightgrey.svg)](https://creativecommons.org/licenses/by-nc-sa/4.0/) [![arXiv](https://img.shields.io/badge/arXiv-2604.24311-b31b1b.svg)](https://arxiv.org/abs/2604.24311) [![EC3 2026](https://img.shields.io/badge/EC3-2026-blue.svg)](https://ec-3.org/publication/ec32026_197/) [![Dataset](https://img.shields.io/badge/%F0%9F%A4%97%20Dataset-DeKH-ffd21e.svg)](https://huggingface.co/datasets/RPTU-FGMB/DeKH) This folder is a self-contained package: model weights, config, a minimal copy of the [Pointcept](https://github.com/Pointcept/Pointcept) codebase (which this model was trained with), and one script (`segment_scan.py`) that takes a raw `.las`/`.laz`/`.ply` scan and produces a segmented file of the same format back out. ## What's in here ``` configs/model_config.py model architecture + class list weights/model_best.pth trained checkpoint pointcept/ trimmed copy of the Pointcept codebase (see below) segment_scan.py chunk -> infer -> merge inference script pyproject.toml / uv.lock environment definition ``` ## Model - Backbone: **PT-v3m1** (Point Transformer V3), ~46M parameters. - Pretrained jointly on Structured3D + ScanNet + S3DIS (Point Prompt Training / PPT-v1m1) with a `RandomColorDrop` augmentation (points randomly lose their color during training, so the model doesn't over-rely on it). - Fine-tuned on **CV4AEC** training data (indoor/construction point clouds). - 10 output classes: `clutter, floor, ceiling, wall, column, door, window, stairs, railing, lights` (see `configs/model_config.py` -> `data.names`). The `pointcept/` directory here is deliberately **not** a full checkout of the Pointcept codebase or of the internal repo this model was trained in — it's trimmed to only what's needed to build this specific model and run inference (backbone + generic dataset/transform machinery), so it doesn't pull in unrelated dataset loaders or the CV4AEC-specific dataset class (`CV4AECDataset`) at all -- `segment_scan.py` builds tiles as plain `pointcept.datasets.DefaultDataset` items instead, so you don't need the CV4AEC dataset or its raw data to run this. If you want the full upstream codebase instead (e.g. to train further), see the notes at the bottom. ## Setup Requires a CUDA GPU (tested on an RTX 4090) and [`uv`](https://docs.astral.sh/uv/): ```bash uv sync ``` This creates `.venv/` with Python 3.12, PyTorch 2.7 (CUDA 12.6), spconv, torch-scatter/torch-cluster, and the handful of other packages `pointcept/` needs -- pinned in `uv.lock` for a reproducible install. `flash-attn` is deliberately **not** installed (it requires a slow from-source build); the model falls back to standard attention, which this config already reflects (`enable_flash=False`). ## Running it ```bash uv run segment_scan.py --input your_scan.las --output your_scan_segmented.las ``` `--input`/`--output` accept `.las`, `.laz`, or `.ply` (mix and match -- input and output don't need to be the same format). What it does: 1. Reads the point cloud and estimates per-point normals (the model expects color + normals as input features). 2. Splits it into overlapping ~20m tiles (large scans don't fit the model in one shot) -- `--tile-size` / `--overlap` to change this. 3. Runs each tile through the model with 10-way test-time augmentation (5 scales x flip/no-flip), softmax-averaged. 4. Merges tiles back into one point cloud: for points covered by more than one tile (the overlap region), takes a majority vote across tiles. 5. Writes a file with the original coordinates/colors plus a per-point `classification` field (0-9, matching the class list above -- for `.ply` this is a custom scalar field, readable in CloudCompare/MeshLab) and a companion `_labels.npy`. On a ~6.6M point scan this takes about 8 minutes end-to-end on an RTX 4090 (most of it is the 10x test-time-augmentation passes per tile). Use `--tile-size` to trade off tile count vs. points-per-tile if you need to fit tighter GPU memory, or edit `TEST_CFG["aug_transform"]` at the top of `segment_scan.py` to drop some of the TTA passes for a faster, slightly less accurate run. ## If you want the full upstream Pointcept instead This delivery's `pointcept/` is intentionally minimal. If you need the rest (other model architectures, training scripts, the full dataset zoo), clone and drop `configs/model_config.py` + `weights/model_best.pth` into it -- the checkpoint loads cleanly against current upstream (`strict=True`), with one config-level rename to be aware of: newer Pointcept renamed the backbone's `cls_mode` argument to `enc_mode` (already reflected in `configs/model_config.py` here). ## License - **Code** (`segment_scan.py`, `pointcept/`, configs): MIT. - **Weights** (`weights/model_best.pth`): CC-BY-NC-SA-4.0. ## Citation If you use this model, please cite the paper it was built for: ```bibtex @inproceedings{chamseddine2026bimstruct3d, author = {Chamseddine, Mahdi and Kaufmann, Fabian and Schellen, Marius and Glock, Christian and Stricker, Didier and Rambach, Jason}, booktitle = {Proceedings of the European Conference on Computing in Construction (EC3)}, organization = {European Council on Computing in Construction}, title = {BIMStruct3D: A Fully Automated Hybrid Learning Scan-to-BIM Pipeline with Integrated Topology Refinement}, year = {2026}, doi = {10.35490/EC3.2026} } ``` Please also credit the [CV4AEC Scan-to-BIM Competition 2024](https://cv4aec.github.io/cvpr2024) ([challenge repo](https://github.com/GradientSpaces/cv4aec-challenge), Gradient Spaces @ Stanford) if you use data or results tied to that benchmark -- this is the specific challenge whose data the model was fine-tuned on. ## Related dataset This model was used to produce the semantic segmentation underlying the BIMStruct3D pipeline's results on the **DeKH (German Hospital) dataset**: real hospital point clouds (four scenes across three buildings) with semantic annotations and ground-truth IFC BIM models, released alongside the paper above. - Dataset: - License: CC BY-NC-SA 4.0 (non-commercial) ## Acknowledgement This research was funded by the European Union as part of the projects: HumanTech (Grant Agreement 101058236) and ShieldBOT (Grant Agreement 101235093).