superpoint-p150
Magic Leap SuperPoint port on one Tenstorrent Blackhole p150a. Weights: magic-leap-community/superpoint · Paper: arXiv:1712.07629 · Upstream code: magicleap/SuperPointPretrainedNetwork · Port: changh95/tt-superpoint
Runs on p150 (mesh P150).
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a tt-metal / ttnn Python environment at tt-metal 8b98410e730 with patches/tt-metal-eth-dispatch.patch applied (git -C $TT_METAL_HOME apply patches/tt-metal-eth-dispatch.patch, then rebuild tt-metal).
hf download changh95/superpoint-p150 --include "code/*" "patches/*" --local-dir superpoint-p150 && cd superpoint-p150
pip install -e code/ # the Python API only
pip install -e "code/[server,test]" # also the HTTP server (fastapi, uvicorn, pydantic) and pytest
# pip install -e code/ (in an environment with ttnn / tt-metal)
from tt_superpoint import SuperPoint
with SuperPoint.from_pretrained(device_id=0) as model: # loads weights, opens the chip, captures traces
out = model("code/sample_data/house_in_field_1080p.jpg", max_keypoints=1024)
print(len(out)) # number of keypoints N
print(out.keypoints[:3]) # (N, 2) [x, y] in original image pixels
print(out.scores[:3]) # (N,) scores, descending
print(out.descriptors.shape) # (N, 256) L2-normalised descriptors
outs = model(["a.jpg", "b.jpg"]) # list in -> list out (PIL / numpy / torch inputs also accepted)
- The first call downloads the weights (
magic-leap-community/superpoint, 5 MB) into your HF cache. from_pretrainedcompiles the kernels and captures the traces. Without cached kernels, start-up takes 10-60 s.from_pretrainedwarms up the common variants (sizes 1920x1080, 1600x900, 1280x720 and 640x480; nms_radius 4 and 0; JPEG and PNG files). Your first call is then as fast as later calls. Start-up takes about 5 s with cached kernels.- To prepare other variants, call
model.warmup(size=(w, h))ormodel.warmup(nms_radius=r). Or setwarmup_variants="minimal"for a faster start-up. - The relative path
code/sample_data/...in the snippet assumes that the current directory is the repo root. From another directory, use an absolute path. - The package detects the ETH-dispatch patch and then uses a 12×10 compute grid. Without the patch, the package uses Tensix dispatch and shows a
RuntimeWarning. The outputs are the same.
| Inputs | File path, bytes, PIL.Image, numpy (H, W) or (H, W, C), torch (H, W), (C, H, W) or (H, W, C); uint8 or float in [0, 1]; any size. A list, or a 4-D batch, gives a list of outputs. |
| Keyword options | max_keypoints (1024; -1 = all above threshold), keypoint_threshold (0.005), nms_radius (4; 1–8 on device, 0 or > 8 on host), return_descriptors (True), bgr (False; use True for cv2.imread arrays), num_workers (4; lists only). |
Output (SuperPointOutput) |
keypoints float32 (N, 2) [x, y] in original image pixels; scores float32 (N,), descending; descriptors float32 (N, 256) or None; image_size, scale, device_nms, timing_ms. out["keypoints"], out.numpy() and out.to_dict() also work. |
- Throughput:
model(list_of_images, num_workers=4)decodes the next images on host threads while the device runs the current image.model.iter(frames)does the same for a generator, for example video frames.precompile_sizes=[(W, H), ...]captures the device resize for your camera sizes at startup. - Lifetime: the
withblock callsmodel.close()at the end.close()releases the traces. It also closes the chip thatfrom_pretrainedopened. A lock serialises the device work, so threads can share one model. - The output is equal to the
/predictresponse of the HTTP server for the same image and parameters. - Full reference:
code/PYTHON.md. Runnable example:code/examples/quickstart.py. The example writeskeypoints.png,keypoints.jsonanddescriptors.npy.
Serving (HTTP)
tt-model pull changh95/superpoint-p150 --with-weights
tt-model serve changh95/superpoint-p150
- Weights
magic-leap-community/superpointgo to your HF cache; the image does not contain them. - Serves on port 20000 (or the next free port); ready when the log says
Application startup complete.
With tt-cli:
tt serve changh95/superpoint-p150
printf '{"image":"%s"}' "$(base64 -w0 code/sample_data/house_in_field_1080p.jpg)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/superpoint-p150
POST /predict:image(base64 PNG/JPEG); optionalmax_keypoints(1024,-1= all above threshold),keypoint_threshold(0.005),nms_radius(4),return_descriptors(true).GET /health,GET /info.- Response:
num_keypoints,keypoints([x, y]in original image pixels, sorted by descendingscores),scores,original_size,image_size(480×640),scale,descriptors,serving_path,timing_ms. descriptors.datais a base64 NPZ:np.load(io.BytesIO(base64.b64decode(data)))["descriptors"]gives(N, 256)float16 rows, L2-normalised, in keypoint order.- Binary routes in
code/models/server/app.py(the current container image does not have them yet). Both routes use the same query parameters and give the same response as/predict:POST /predict_raw: the body is the PNG/JPEG file bytes.POST /predict_plane?height=H&width=W: the body is the R plane of the image asH*Wraw uint8 bytes.
Demo
Top-500 keypoints on the 480×640 network frame of code/sample_data/house_in_field_1080p.jpg (media/sample.png, natural image) |
|---|
![]() |
Demo & Performances
Warm, batch 1, 480×640 network frame, 1600×900 JPEG for served requests. All numbers are measured in the p150 configuration: dispatch on ETH cores, 1 command queue, 12×10 compute grid. We measured on a Blackhole chip (2026-10-03, code/ at the optimized build). The Python model() row is from 2026-10-04 on a BH Galaxy chip in the same configuration. On 2026-10-05 we measured all rows again in this configuration, on a host with more load (load 15–28). The ranges below include these values. Device code did not change between the three dates.
| Metric | Performance |
|---|---|
End-to-end predict() (base64 + JPEG decode, device, post-processing; 100 requests) |
17.5–18.6 ms median |
device_forward in predict() (full-size plane upload, device resize, one trace, D2H) |
1.55–1.67 ms median · 1.49 ms min |
HTTP POST /predict to the server in code/ (2026-10-05; 1600×900 JPEG; 50 requests; client wall time) |
23.0–24.8 ms median · 21.6 ms min · server total 17.8–18.6 ms |
Python model() call (2026-10-04; 1600×900 R plane, device resize · JPEG file path incl. decode; 200 / 100 calls) |
1.58 ms median · 1.49 ms min · JPEG path 16.8 ms median · 16.3 ms min |
| Request from the 8-bit 480×640 plane (H2D + trace + D2H + host decode) | 0.91 ms median (0.87–0.97) · 0.876 ms min |
| Trace replay, input resident (network + NMS + keypoint list + descriptor sampling) | 0.50 ms median · 0.48 ms min |
| Device time per trace replay (device profiler, 23 ops) | 462.8 µs (462.2–463.6 µs on 2026-10-05) |
RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in eager PyTorch 2.11, batch 1, with the same host pre/post-processing. The GPU host is a different machine. Full table: GPU_COMPARISON.md.
| RTX 5090 precision | GPU served-like total | vs ours (18.0 ms) | GPU upload + forward + readback vs our request (0.91 ms) | GPU forward, input resident, vs our trace (0.50 ms) |
|---|---|---|---|---|
| fp32 strict | 22.75 ms | Blackhole 1.26× faster | 2.94 ms: Blackhole 3.23× faster | 2.26 ms: Blackhole 4.52× faster |
| bf16 autocast | 21.45 ms | Blackhole 1.19× faster | 1.59 ms: Blackhole 1.74× faster | 0.91 ms: Blackhole 1.83× faster |
| fp16 autocast | 21.90 ms | Blackhole 1.22× faster | 1.53 ms: Blackhole 1.68× faster | 0.86 ms: Blackhole 1.71× faster |
bf16 autocast + torch.compile (CUDA graphs) |
not measured | – | 1.29 ms: Blackhole 1.42× faster | 0.59 ms: Blackhole 1.18× faster |
fp16 autocast + torch.compile (CUDA graphs) |
not measured | – | 1.21 ms: Blackhole 1.33× faster | 0.51 ms: parity (1.01×) |
The RTX 5090 table above uses the 2026-10-03 medians. With the 2026-10-05 medians (served total 18.2–18.6 ms, request 0.96–0.97 ms, trace 0.50 ms), Blackhole is 1.15–1.25× faster on the served total, 1.25–3.05× faster on upload + forward + readback (1.25–1.26× against the fp16 compiled GPU), and from parity (fp16 compiled) to 4.5× faster on the trace. See GPU_COMPARISON.md.
The Blackhole trace also does the keypoint list and the descriptor sampling. The GPU does these steps on the host in a further 2.2 ms. Our request uploads only the 8-bit image plane and reads back only the keypoints and their descriptors. The GPU uploads an fp32 tensor and reads back the full score and descriptor maps. The host JPEG decode is the largest part of the served total on both sides. The previous release was 3.2× slower than the bf16 GPU on the device forward.
Caveats
- Does not scale to multiple p150a in a mesh configuration. Current build enforces 12x10 tensix cores for compute, which is enabled by moving dispatch functions that were originally in 10 tensix cores to ETH cores. This implementation therefore assumes that chip-to-chip ethernet communication is not required by the user. The tt-metal change is in
patches/tt-metal-eth-dispatch.patch. - The HTTP server in
code/opens the chip like the Python API: ETH dispatch, 1 command queue, 12×10 grid (SP_DISPATCH=auto, the default).SP_DISPATCH=workeris an opt-in for BH Galaxy only. It uses Tensix dispatch, which gives an 11×10 compute grid on a p150. The current container image is older and opens the chip with Tensix dispatch. - Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
- Accuracy against the fp32 torch reference (natural image): score PCC 0.998583, descriptor PCC 0.999280, F1 0.9890. Some custom kernels sum in a different order than the ttnn conv. Thus the outputs are not bit-identical to the previous release. See
OPT_REPORT.md. nms_radius4 runs in the main trace. Radii 1–8 use a per-radius device NMS trace, which compiles on first use. Radius 0 or above 8 uses the host NMS.- Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
- Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404.
Licensing
- Weights: magic-leap-community/superpoint,
other(Magic Leap SuperPoint licence, noncommercial research use only); not redistributed here. - Port and serving code (
code/): Apache-2.0 (SPDX headers on the modules), from changh95/tt-superpoint, distributed under the same upstream terms since a port cannot grant more than its upstream does. patches/tt-metal-eth-dispatch.patchchanges tt-metal, which is Apache-2.0.
Provenance
These are the exact sources the container image was built from. code/ has since been updated (2026-10-03 optimized build; see OPT_REPORT.md) and is newer than the image. The Python package (code/tt_superpoint, 2026-10-04) and the binary routes are only in code/. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:
| component | built from |
|---|---|
| tt-metal | 8b98410e730bb504fea43a88609756e34821d91d |
code/ digest (image) |
ae768681d4aa677d (sha256, first 16 hex digits; the current code/ differs) |
| built | 2026-09-13T15:21:44+00:00 by tt-model 0.1.0 |
Model tree for changh95/superpoint-p150
Base model
magic-leap-community/superpoint