mast3r-p150

NAVER DUSt3R ViT-L two-view 3D reconstruction (the MASt3R backbone) port on one Tenstorrent Blackhole p150a. Weights: naver/DUSt3R_ViTLarge_BaseDecoder_512_dpt · Paper: arXiv:2312.14132 · Upstream code: naver/mast3r · Port: changh95/tt-mast3r

Runs on p150 (mesh P150).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart (Python)

Prerequisite: a Python environment with tt-metal / ttnn at tt-metal 8b98410e730, built with patches/tt-metal-eth-dispatch.patch. In that environment, download this repo and install its code/ directory:

hf download changh95/mast3r-p150 --exclude "image/*" --local-dir mast3r-p150 && cd mast3r-p150
pip install -e "code/[pose]"      # the [pose] extra adds OpenCV for return_pose=True
pip install -e "code/[pose,server,test]"   # also the HTTP server and the tests (see PYTHON.md)
from mast3r_p150 import Mast3rP150

with Mast3rP150.from_pretrained(device_id=0) as model:
    out = model("media/source_1.png", "media/source_2.png")
    print(out.pts3d1.shape, out.conf1.shape)    # (512, 512, 3) (512, 512)
    xyz, rgb = out.point_cloud(conf_min=3.0)    # fused coloured cloud, camera-1 frame
    out.save_ply("pointcloud.ply")

The paths media/source_1.png and media/source_2.png are relative to the repo root. Run the snippet from the repo root, or give absolute paths.

The first call of from_pretrained downloads the NAVER weights to your HF cache. It also compiles the kernels and captures the metal trace (about 65 s). The package does not install ttnn. You do not need the HTTP server.

Mast3rP150.from_pretrained() prepares the plain call, the pose call (return_pose=True) and predict_pairs before it returns. Thus the first call is as fast as the later calls (within 10 %). Start-up takes about 8 s with a warm kernel cache. Use warmup_variants=["pair"] for a faster start-up, or model.warmup("crop") to prepare more variants later.

name type / default meaning
input view1, view2 path, PNG/JPEG bytes, PIL.Image, numpy or torch array (uint8 0..255 or float 0..1), DUSt3R view dict Any image size. View 1 is the reference camera.
option return_pose False Also calculate the pose of view 2 (PnP-RANSAC, needs OpenCV).
option intrinsics None The 3×3 K of view 2 in its original pixels. PnP then uses this K.
option conf_pct 50.0 PnP uses the top conf_pct % of view-2 pixels by confidence.
option preprocess "pad" pad adds gray padding to make a square. crop cuts a centre square.
output pts3d1, pts3d2 (512, 512, 3) float32 One 3-D point for each pixel. Both views are in the camera-1 frame, in DUSt3R scale (not metres).
output conf1, conf2 (512, 512) float32 Confidence 1 + exp(c) >= 1.
output depth1, depth2, rgb1, rgb2, masks (512, 512) arrays Depth (pts3d[..., 2]), network input pixels, and the photo mask (False on the gray padding).
output pose, preprocess, timing_ms Pose or None, 2 dicts, dict R, t (X_cam2 = R @ X_cam1 + t), focal1, focal2; the pixel mapping to the 512 grid; time per stage.
  • Helpers on the result: point_cloud(), save_ply(), save_npz() (the server npz format).
  • Throughput: model.predict_pairs([(a, b), (c, d), ...]) overlaps host work with the device. It gives the same arrays as one call for each pair.
  • DUSt3R global alignment: model.inference(pairs) returns the DUSt3R inference() output dict for global_aligner.
  • Lifetime: use with ... as model: or call model.close(). One model can be open in a process at a time.
  • Full reference: PYTHON.md. Runnable example: python code/examples/quickstart.py. It writes pointmaps.npz, pointcloud.ply, depth.png and summary.json.

Serving (HTTP)

tt-model pull  changh95/mast3r-p150 --with-weights
tt-model serve changh95/mast3r-p150       # or: tt serve changh95/mast3r-p150
printf '{"image1":"%s","image2":"%s"}' "$(base64 -w0 media/source_1.png)" "$(base64 -w0 media/source_2.png)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/mast3r-p150
  • The image does not contain the weights. pull --with-weights puts them in your HF cache.
  • The server uses port 20000 (or the next free port). It is ready when the log shows Application startup complete.
  • POST /predict: image1, image2 (base64 PNG/JPEG, view 1 = reference camera; or images: [b64, b64]); optional output_format (npz | png, default npz), return_pose (false), intrinsics (3×3 K of view 2 in its original pixels), conf_pct (50).
  • Response: npz_b64 (a base64 .npz with pts3d1/pts3d2 float32 (512,512,3) and conf1/conf2 float16 (512,512), camera-1 frame), preprocess, summary, pose (R, t, focal_1, focal_2, or null), timing_ms. output_format: "png" returns 16-bit depth and 8-bit confidence PNGs.
  • GET /health, GET /info.
  • Newer code/ only (the container image does not have these yet):
    • POST /predict_npz: the same request. The body is the .npz file (application/x-npz). The header X-Mast3r-Meta holds the other fields as JSON.
    • Request field compress_npz (default false): true gives a deflated .npz. The default .npz is stored, which is larger but encodes faster.

Demo

Kitchen frames 00 / 03 (VGGT example scene) → both predicted pointmaps in the camera-1 frame, coloured by the source pixels (served npz output, code/make_demo.py).

Demo & Performances

Warm, batch 1, one 512×512 pair, fused traced graph (MAST3R_OPT=all), bench_breakdown.py with 30 iterations per run (median / min). All rows use the p150 configuration: dispatch on ETH cores, 1 command queue (CQ), 12×10 compute grid. The served path config is this configuration with uint8 input (the serve env of the newer code/).

Metric Performance
Model call model(img1, img2) (H2D + trace + D2H of both maps), served path config 27.6–27.8 ms median (2 runs) · 26.9 ms min
Device trace (ViT-L encoder ×2, decoder, two DPT heads), served path config 23.7–23.9 ms median · 23.1 ms min
Model call, float input 28.4 ms median (2 runs) · 27.9 ms min
Device trace, float input 23.3–23.4 ms median · 23.0 ms min
Device span (profiler cycles at 1.35 GHz, 3 replays), measured before the LoFi MLP change 23.47–23.49 ms · 688 programs
Host prep · H2D, served path config (uint8 pixels) 0.43 ms · 0.54 ms
D2H of both maps, served path config · float input 2.75 ms · 3.1–3.4 ms
return_pose forward on the real pair (sym_check.py): symmetric graph (sym) · two-pass 38.0–38.9 ms median · 53.9–54.9 ms median
Python model() call (mast3r_p150, timing_ms['forward'], 30 warm calls) 26.9–27.0 ms median · 26.5 ms min
HTTP server, newer code/ (timing_ms.forward · /predict npz total · /predict_npz total, 30 requests) 26.6–27.0 ms median · 80–83 ms · 73–75 ms

The measurement hardware is a Blackhole chip with a 12×10 compute grid and dispatch on ETH cores (patches/tt-metal-eth-dispatch.patch), with 1 command queue. This is the configuration of a single p150. The 2026-10-04 build runs the encoder and decoder MLPs at LoFi (MAST3R_OPT knobs lofie, lofid). An independent verifier measured the rows above again on 2026-10-05 in this configuration. The device span row is from the previous build. Accuracy against the torch fp32 reference on the real kitchen pair (served path config): pts3d PCC 0.99015 / 0.99097 (gate 0.989), conf PCC 0.99246 / 0.99519 (gate 0.99); synthetic pair test_mast3r.py end_to_end PCC 0.9985. Details: VERIFICATION_2026-10-03.md.

Before 2026-10-05, this card showed a served model call of 26.5 ms. That number used worker (Tensix) dispatch with 2 CQ, which does not give a 12×10 grid on a p150. With 1 CQ, the head-1 readback does not overlap the head-2 compute, so the model call is about 1.2 ms slower. The outputs are bit-identical.

A whole model(path, path) call takes 50–59 ms median, because it includes PNG decode and resize on the host. predict_pairs takes 27.3–28.1 ms per pair. The Python API gives arrays that are bit-identical to the server /predict and /predict_npz outputs, including the pose.

RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in PyTorch 2.11, batch 1, on the kitchen demo pair. The "incl. H2D/D2H" column compares with our model call (27.7 ms). The "forward only" column compares with our device trace (23.8 ms). Full table: GPU_COMPARISON.md.

RTX 5090 precision GPU incl. H2D/D2H vs current build (27.7 ms) GPU forward only vs ours (23.8 ms)
fp32 strict, eager 103.4 ms Blackhole 3.73× faster 102.3 ms: Blackhole 4.30× faster
tf32, eager 70.8 ms Blackhole 2.55× faster 70.0 ms: Blackhole 2.94× faster
bf16 autocast, eager 63.8 ms Blackhole 2.30× faster 62.6 ms: Blackhole 2.63× faster
fp16 autocast, eager 56.6 ms Blackhole 2.04× faster 55.5 ms: Blackhole 2.33× faster
bf16 weights, eager 42.3 ms Blackhole 1.53× faster 41.3 ms: Blackhole 1.73× faster
bf16 weights + SDPA, eager 35.3 ms Blackhole 1.28× faster 34.3 ms: Blackhole 1.44× faster
bf16 weights + torch.compile (CUDA graphs) 21.1 ms GPU 1.31× faster 20.1 ms: GPU 1.18× faster

The Blackhole lead comes from one metal trace that runs fused, block-sharded bf16 kernels on all 120 compute cores. The eager GPU rows pay for kernel launches and explicit softmax attention. The autocast rows also cast the fp32 weights again on each forward. torch.compile with CUDA graphs removes these costs, and then the GPU is about 1.2-1.3× faster. The newer code/ writes a stored npz. This change decreases the npz encode from about 220 ms to about 17 ms and the served /predict total from about 285 ms to about 80 ms.

Caveats

  • Does not scale to multiple p150a in a mesh configuration. The current build uses a 12x10 compute grid of Tensix cores. To get this grid, the dispatch functions move from 10 Tensix cores to ETH cores (patches/tt-metal-eth-dispatch.patch). Thus, this build assumes that you do not need chip-to-chip ethernet communication.
  • The newer code/ uses ETH dispatch and 1 CQ by default (MAST3R_DISPATCH=auto, MAST3R_CQS=1). MAST3R_CQS=2 is an opt-in for Galaxy only. It needs worker (Tensix) dispatch, which gives an 11×10 grid on a p150. The container image and its tt-model.yaml still set MAST3R_CQS=2 until the image is rebuilt. Without the ETH-dispatch patch, auto uses Tensix dispatch (11×10 on a p150) and logs a warning.
  • Fixed 512×512 input, exactly one image pair per request (batch 1). Non-square images are gray-padded to square (MAST3R_PREPROC=crop centre-crops instead). The estimated focal can be 5-8 % high, so pass intrinsics when you know them.
  • Numbers published before 2026-09-14 describe an older network with wrong decoder taps. This card does not repeat them.
  • DUSt3R backbone only: no MASt3R matcher / descriptor head and no N-view global alignment. The pose is single-pair PairViewer (estimated-focal PnP-RANSAC). With return_pose, one symmetric graph (sym) computes the pose maps. Its outputs are bit-identical to the two-pass path.
  • bf16 on device: on the kitchen pair, pts3d PCC vs fp32 is 0.990 / 0.991 per head and conf PCC is 0.992 / 0.995. Loosen thresholds that you tuned on the reference. Some optimizations change numerics (SDPA chunks, GELU polynomial, HiFi3 head convs, upsample kernel, LoFi MLP). OPT_REPORT.md lists each change.
  • The LoFi MLP change (2026-10-04) passes all gates. Some metrics without a gate decrease a little. The real-pair median depth error |dz|/z against fp32 goes from 1.17 / 1.10 % to 1.24 / 1.17 %. The synthetic uint8 head-1 PCC goes from 0.99733 to 0.99593. To get the previous numerics, set MAST3R_OPT=all,-lofie,-lofid.
  • Dense outputs come base64-encoded (.npz, or 16-bit/8-bit PNG).
  • Weights are NAVER's DUSt3R checkpoint under CC-BY-NC-SA-4.0: non-commercial use only, share-alike.
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.
  • p150a power was not measured, so no efficiency comparison is made.

Licensing

  • Weights: naver/DUSt3R_ViTLarge_BaseDecoder_512_dpt, CC-BY-NC-SA-4.0 (non-commercial); fetched by serve from NAVER's repo, not redistributed here.
  • Port and serving code (code/), from changh95/tt-mast3r: the DUSt3R port follows the weights' CC-BY-NC-SA-4.0 (share-alike, non-commercial); the serving layer (code/models/server/) is Apache-2.0 per its SPDX headers; tt-metal / tt-nn are Apache-2.0.

Provenance

These are the exact sources the container image was built from. code/ has since been updated (2026-10-04 build with the LoFi MLP, the faster npz encode and the mast3r_p150 Python package; 2026-10-05: ETH dispatch and 1 CQ are the default for the server and the harnesses; see OPT_REPORT.md and PYTHON.md) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:

component built from
tt-metal 8b98410e730bb504fea43a88609756e34821d91d
code/ digest (image) b45444934e1d76bf (sha256, first 16 hex digits; the current code/ differs)
built 2026-09-14T06:35:13+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for changh95/mast3r-p150

Finetuned
(3)
this model

Collection including changh95/mast3r-p150

Paper for changh95/mast3r-p150