superpoint-p150

Magic Leap SuperPoint port on one Tenstorrent Blackhole p150a. Weights: magic-leap-community/superpoint · Paper: arXiv:1712.07629 · Upstream code: magicleap/SuperPointPretrainedNetwork · Port: changh95/tt-superpoint

Runs on p150 (mesh P150).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart (Python)

Prerequisite: a tt-metal / ttnn Python environment at tt-metal 8b98410e730 with patches/tt-metal-eth-dispatch.patch applied (git -C $TT_METAL_HOME apply patches/tt-metal-eth-dispatch.patch, then rebuild tt-metal).

hf download changh95/superpoint-p150 --include "code/*" "patches/*" --local-dir superpoint-p150 && cd superpoint-p150
pip install -e code/                    # the Python API only
pip install -e "code/[server,test]"     # also the HTTP server (fastapi, uvicorn, pydantic) and pytest
# pip install -e code/   (in an environment with ttnn / tt-metal)
from tt_superpoint import SuperPoint

with SuperPoint.from_pretrained(device_id=0) as model:   # loads weights, opens the chip, captures traces
    out = model("code/sample_data/house_in_field_1080p.jpg", max_keypoints=1024)
    print(len(out))                # number of keypoints N
    print(out.keypoints[:3])       # (N, 2) [x, y] in original image pixels
    print(out.scores[:3])          # (N,) scores, descending
    print(out.descriptors.shape)   # (N, 256) L2-normalised descriptors

    outs = model(["a.jpg", "b.jpg"])  # list in -> list out (PIL / numpy / torch inputs also accepted)
  • The first call downloads the weights (magic-leap-community/superpoint, 5 MB) into your HF cache.
  • from_pretrained compiles the kernels and captures the traces. Without cached kernels, start-up takes 10-60 s.
  • from_pretrained warms up the common variants (sizes 1920x1080, 1600x900, 1280x720 and 640x480; nms_radius 4 and 0; JPEG and PNG files). Your first call is then as fast as later calls. Start-up takes about 5 s with cached kernels.
  • To prepare other variants, call model.warmup(size=(w, h)) or model.warmup(nms_radius=r). Or set warmup_variants="minimal" for a faster start-up.
  • The relative path code/sample_data/... in the snippet assumes that the current directory is the repo root. From another directory, use an absolute path.
  • The package detects the ETH-dispatch patch and then uses a 12×10 compute grid. Without the patch, the package uses Tensix dispatch and shows a RuntimeWarning. The outputs are the same.
Inputs File path, bytes, PIL.Image, numpy (H, W) or (H, W, C), torch (H, W), (C, H, W) or (H, W, C); uint8 or float in [0, 1]; any size. A list, or a 4-D batch, gives a list of outputs.
Keyword options max_keypoints (1024; -1 = all above threshold), keypoint_threshold (0.005), nms_radius (4; 1–8 on device, 0 or > 8 on host), return_descriptors (True), bgr (False; use True for cv2.imread arrays), num_workers (4; lists only).
Output (SuperPointOutput) keypoints float32 (N, 2) [x, y] in original image pixels; scores float32 (N,), descending; descriptors float32 (N, 256) or None; image_size, scale, device_nms, timing_ms. out["keypoints"], out.numpy() and out.to_dict() also work.
  • Throughput: model(list_of_images, num_workers=4) decodes the next images on host threads while the device runs the current image. model.iter(frames) does the same for a generator, for example video frames. precompile_sizes=[(W, H), ...] captures the device resize for your camera sizes at startup.
  • Lifetime: the with block calls model.close() at the end. close() releases the traces. It also closes the chip that from_pretrained opened. A lock serialises the device work, so threads can share one model.
  • The output is equal to the /predict response of the HTTP server for the same image and parameters.
  • Full reference: code/PYTHON.md. Runnable example: code/examples/quickstart.py. The example writes keypoints.png, keypoints.json and descriptors.npy.

Serving (HTTP)

tt-model pull  changh95/superpoint-p150 --with-weights
tt-model serve changh95/superpoint-p150
  • Weights magic-leap-community/superpoint go to your HF cache; the image does not contain them.
  • Serves on port 20000 (or the next free port); ready when the log says Application startup complete.

With tt-cli:

tt serve changh95/superpoint-p150
printf '{"image":"%s"}' "$(base64 -w0 code/sample_data/house_in_field_1080p.jpg)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/superpoint-p150
  • POST /predict: image (base64 PNG/JPEG); optional max_keypoints (1024, -1 = all above threshold), keypoint_threshold (0.005), nms_radius (4), return_descriptors (true).
  • GET /health, GET /info.
  • Response: num_keypoints, keypoints ([x, y] in original image pixels, sorted by descending scores), scores, original_size, image_size (480×640), scale, descriptors, serving_path, timing_ms.
  • descriptors.data is a base64 NPZ: np.load(io.BytesIO(base64.b64decode(data)))["descriptors"] gives (N, 256) float16 rows, L2-normalised, in keypoint order.
  • Binary routes in code/models/server/app.py (the current container image does not have them yet). Both routes use the same query parameters and give the same response as /predict:
    • POST /predict_raw: the body is the PNG/JPEG file bytes.
    • POST /predict_plane?height=H&width=W: the body is the R plane of the image as H*W raw uint8 bytes.

Demo

Top-500 keypoints on the 480×640 network frame of code/sample_data/house_in_field_1080p.jpg (media/sample.png, natural image)

Demo & Performances

Warm, batch 1, 480×640 network frame, 1600×900 JPEG for served requests. All numbers are measured in the p150 configuration: dispatch on ETH cores, 1 command queue, 12×10 compute grid. We measured on a Blackhole chip (2026-10-03, code/ at the optimized build). The Python model() row is from 2026-10-04 on a BH Galaxy chip in the same configuration. On 2026-10-05 we measured all rows again in this configuration, on a host with more load (load 15–28). The ranges below include these values. Device code did not change between the three dates.

Metric Performance
End-to-end predict() (base64 + JPEG decode, device, post-processing; 100 requests) 17.5–18.6 ms median
device_forward in predict() (full-size plane upload, device resize, one trace, D2H) 1.55–1.67 ms median · 1.49 ms min
HTTP POST /predict to the server in code/ (2026-10-05; 1600×900 JPEG; 50 requests; client wall time) 23.0–24.8 ms median · 21.6 ms min · server total 17.8–18.6 ms
Python model() call (2026-10-04; 1600×900 R plane, device resize · JPEG file path incl. decode; 200 / 100 calls) 1.58 ms median · 1.49 ms min · JPEG path 16.8 ms median · 16.3 ms min
Request from the 8-bit 480×640 plane (H2D + trace + D2H + host decode) 0.91 ms median (0.87–0.97) · 0.876 ms min
Trace replay, input resident (network + NMS + keypoint list + descriptor sampling) 0.50 ms median · 0.48 ms min
Device time per trace replay (device profiler, 23 ops) 462.8 µs (462.2–463.6 µs on 2026-10-05)

RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in eager PyTorch 2.11, batch 1, with the same host pre/post-processing. The GPU host is a different machine. Full table: GPU_COMPARISON.md.

RTX 5090 precision GPU served-like total vs ours (18.0 ms) GPU upload + forward + readback vs our request (0.91 ms) GPU forward, input resident, vs our trace (0.50 ms)
fp32 strict 22.75 ms Blackhole 1.26× faster 2.94 ms: Blackhole 3.23× faster 2.26 ms: Blackhole 4.52× faster
bf16 autocast 21.45 ms Blackhole 1.19× faster 1.59 ms: Blackhole 1.74× faster 0.91 ms: Blackhole 1.83× faster
fp16 autocast 21.90 ms Blackhole 1.22× faster 1.53 ms: Blackhole 1.68× faster 0.86 ms: Blackhole 1.71× faster
bf16 autocast + torch.compile (CUDA graphs) not measured – 1.29 ms: Blackhole 1.42× faster 0.59 ms: Blackhole 1.18× faster
fp16 autocast + torch.compile (CUDA graphs) not measured – 1.21 ms: Blackhole 1.33× faster 0.51 ms: parity (1.01×)

The RTX 5090 table above uses the 2026-10-03 medians. With the 2026-10-05 medians (served total 18.2–18.6 ms, request 0.96–0.97 ms, trace 0.50 ms), Blackhole is 1.15–1.25× faster on the served total, 1.25–3.05× faster on upload + forward + readback (1.25–1.26× against the fp16 compiled GPU), and from parity (fp16 compiled) to 4.5× faster on the trace. See GPU_COMPARISON.md.

The Blackhole trace also does the keypoint list and the descriptor sampling. The GPU does these steps on the host in a further 2.2 ms. Our request uploads only the 8-bit image plane and reads back only the keypoints and their descriptors. The GPU uploads an fp32 tensor and reads back the full score and descriptor maps. The host JPEG decode is the largest part of the served total on both sides. The previous release was 3.2× slower than the bf16 GPU on the device forward.

Caveats

  • Does not scale to multiple p150a in a mesh configuration. Current build enforces 12x10 tensix cores for compute, which is enabled by moving dispatch functions that were originally in 10 tensix cores to ETH cores. This implementation therefore assumes that chip-to-chip ethernet communication is not required by the user. The tt-metal change is in patches/tt-metal-eth-dispatch.patch.
  • The HTTP server in code/ opens the chip like the Python API: ETH dispatch, 1 command queue, 12×10 grid (SP_DISPATCH=auto, the default). SP_DISPATCH=worker is an opt-in for BH Galaxy only. It uses Tensix dispatch, which gives an 11×10 compute grid on a p150. The current container image is older and opens the chip with Tensix dispatch.
  • Every image is resized to 480×640 (bilinear, /255, channel 0); one image per request, batch 1, requests serialised on the chip.
  • Accuracy against the fp32 torch reference (natural image): score PCC 0.998583, descriptor PCC 0.999280, F1 0.9890. Some custom kernels sum in a different order than the ttnn conv. Thus the outputs are not bit-identical to the previous release. See OPT_REPORT.md.
  • nms_radius 4 runs in the main trace. Radii 1–8 use a per-radius device NMS trace, which compiles on first use. Radius 0 or above 8 uses the host NMS.
  • Weights are research-only: the Magic Leap SuperPoint licence allows academic or non-profit organisation NONCOMMERCIAL research use.
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.

Licensing

Provenance

These are the exact sources the container image was built from. code/ has since been updated (2026-10-03 optimized build; see OPT_REPORT.md) and is newer than the image. The Python package (code/tt_superpoint, 2026-10-04) and the binary routes are only in code/. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:

component built from
tt-metal 8b98410e730bb504fea43a88609756e34821d91d
code/ digest (image) ae768681d4aa677d (sha256, first 16 hex digits; the current code/ differs)
built 2026-09-13T15:21:44+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for changh95/superpoint-p150

Finetuned
(1)
this model

Collection including changh95/superpoint-p150

Paper for changh95/superpoint-p150