tt-waypoint
Waypoint-1.5-1B is Overworld's interactive "world model". It is an autoregressive causal diffusion transformer that generates video one frame at a time, steered by mouse and scroll input.
This package runs it on Tenstorrent Blackhole via TTNN:
- What it is: a from-scratch bring-up. Nothing in tt-metal's
models.tt_ditcovers this architecture. - How you use it: a stateful, session-based HTTP API served from a small ASGI server.
- What it needs: one Blackhole chip (profile
P150, a 1x1 mesh). All hardware verification so far ran on one chip of a P300c board. - Maturity: experimental. Frames are coherent, but it runs roughly 1,400x below real time, and no performance work has been done yet.
Packaged and published with tt-model-manager as a v6 thin bundle (manifest schema 6). It is a pip/venv install, not a container image.
At a glance
| Architecture | 24-layer autoregressive causal diffusion transformer (per-layer ring-buffer frame cache, 512 tokens per frame), plus a small CNN VAE (ChunkedStreamingTAEHV) |
| Hardware | 1 Blackhole chip (profile p150, mesh P150, 1x1). Verified on one chip of a P300c; never run on a physical P150 |
| Input / output | Seed image: one, resized to 512x1024. Output: one generated 512x1024 frame per step, 4 denoising steps per frame. No text prompt |
| License | Apache-2.0 (weights and port code) |
| Status | Experimental: correct end to end per step and per component, not performance-tuned |
| Model CI v0 | not yet run |
Intended use
Direct use: interactive demos and research into world-model inference on Tenstorrent hardware. Seed a session with an image, then step it forward one frame at a time with directional (mouse) and zoom (scroll) controls.
Out-of-scope use:
- Real-time or interactive-rate play. It runs at about 0.04 FPS against the upstream 60 FPS target.
- Multiple concurrent users. There is one active session per server process.
- Keyboard or button control. The 256-wide button vector has no published semantics, so it is always sent as zeros.
- Text-prompted generation. This checkpoint has
prompt_conditioning: null. - Multi-chip meshes.
Quickstart
uv tool install tenstorrent # once: the Tenstorrent CLI, `tt`
tt model pull episod/tt-waypoint
tt serve episod/tt-waypoint
tt model pull downloads two things:
- This bundle: a small venv-install recipe. It is
install.sh/run.shplus two wheels;ttnn,torchand the base HTTP stack come from an index at install time. - The
Overworld/Waypoint-1.5-1Bweights: revision391f928, into the bundle's own HF cache. Weights come down by default, andtthas no--with-weightsflag. The pull fetches the whole upstream repo, about 11.4 GB. The server reads only three files from it.
serve starts the model's own HTTP server on port 20000, or the next free port if 20000 is busy. It is ready when it logs Application startup complete.
Measured on a fresh install of this bundle (P300c QuietBox, 2026-09-27):
| step | time |
|---|---|
pull --with-weights (venv plus weights) |
4 min 40 s |
| Launch to ready | 7 s |
First POST /v1/sessions |
about 39 s, which includes first-time kernel compile |
| Each later step | about 25 s |
Without tt-cli, tt-model alone does the whole job:
tt-model pull episod/tt-waypoint --with-weights
tt-model serve episod/tt-waypoint
Serve profiles
| profile | hardware | mesh | sessions per process |
|---|---|---|---|
p150 |
1 Blackhole chip. Verified on one chip of a P300c; a physical P150 has not been tested | P150 (1x1) |
1 |
No multi-chip profile exists. SUPPORTED_MESH_SHAPES is {(1, 1)}, and the server refuses any other WAYPOINT_MESH_SHAPE.
Using it
This is not an OpenAI-compatible API. /v1/models exists only as a stub for tooling. The real interface is a stateful session: seed once from an image, then step.
# Seed a session from a real starting image (base64-encoded PNG/JPEG; resized to 512x1024):
curl -s localhost:20000/v1/sessions -H 'Content-Type: application/json' \
-d "{\"image_b64\": \"$(base64 -w0 my_photo.png)\"}"
# -> {"session_id": "...", "frame_index": 1, "frame_b64": "<base64 PNG>"}
# Step the session forward one frame, steered by direction/zoom:
curl -s localhost:20000/v1/sessions/<session_id>/step \
-H 'Content-Type: application/json' -d '{"direction": "forward", "zoom": 0.0}'
# -> {"session_id": "...", "frame_index": 2, "frame_b64": "<base64 PNG>"}
# End the session:
curl -s -X DELETE localhost:20000/v1/sessions/<session_id>
| endpoint | purpose |
|---|---|
POST /v1/sessions |
Takes {"image_b64": str}. Seeds a new session and returns the decoded seed frame. Discards any session already in progress |
POST /v1/sessions/{id}/step |
Takes {"direction": str = "forward", "zoom": float = 0.0}. Generates one frame |
DELETE /v1/sessions/{id} |
Ends the session |
GET /health, GET /tt-liveness |
Readiness and liveness. Both return 503 until the model is loaded |
GET /v1/models |
Stub listing, for tooling |
direction: one offorward,back,left,right,forward_left,forward_right,back_left,back_right. These map to the model's[dx, dy]mouse input.zoom: the scroll scalar. Only its sign has a documented meaning.- Seeding: the API has no RNG seed parameter. Per-step noise is drawn inside the server.
- Errors:
503: the model is still starting.404: the session ID is unknown, or a newerPOST /v1/sessionssuperseded it.409: you stepped with no active session.400: the request was bad.
The GitHub repo also has a local Gradio UI (app.py), which is a pure HTTP client for this same server.
Expected performance
| metric | value | source |
|---|---|---|
| Per-step denoiser correctness vs HF reference (8 real denoising steps, correlation) | about 0.959 on average | BRINGUP_LOG.md (fp32-accumulation row; step-trace row). The range before the fp32-accumulation fix was 0.897-0.971 |
| 24-layer transformer, single forward pass (correlation vs reference) | 0.963 | BRINGUP_LOG.md (fp32-accumulation row) |
| Single decoder layer PCC (layer 0, frame-0 prefill) | 0.999 | BRINGUP_LOG.md (ttm-* skills row) |
| VAE encoder / decoder (correlation vs reference) | 0.9996 / > 0.99 on all 12 frames | BRINGUP_LOG.md (Stage 5 rows) |
| Generated frames from a real seed image | Visually coherent and plausible. Inspected by eye; no numeric perceptual metric | bring-up log; re-checked on both bundles below |
| Warm latency per generated frame, installed bundle (4 denoise steps + 1 cache commit) | median 24.95 s (min 24.64, max 25.35), N=15 | published bundle 1dd34b8, 2026-09-27 |
| Same, re-verified on this revision of the bundle | median 25.72 s (min 25.70, max 25.91), N=3 | this bundle, 2026-09-27. +3% vs the row above |
| Effective throughput | 0.040 FPS | 1 / median |
| Latency vs session length (steps 1 / 5 / 10 / 15 / 16) | 25.44 / 25.05 / 24.79 / 24.95 / 24.64 s | flat, no drift over 16 frames |
| First step after a cold start | 25.44 s. The first POST /v1/sessions on a cold process takes 38-39 s |
includes kernel compile |
Dev build, for comparison (tt-metal v0.78.0 from source, in-process benchmark.py) |
27.41 s per frame | PORT_PLAN.md |
| Upstream GPU reference, for scale | 56 FPS on an RTX 5090, 4-step unquantized | upstream model card |
Methodology: on one Blackhole chip of a P300c (1x1 mesh, ttnn==0.78.0 from PyPI, eager execution, batch 1, one session), a session seeded from ref_activations/seed_frame_ref.png was stepped with the default controls (forward, zoom 0), timing host HTTP wall-clock per POST .../step after one untimed warmup step: N=15 on the published bundle 1dd34b8 and N=3 on this bundle revision (weights 391f928), both on 2026-09-27.
How to read the accuracy numbers:
- Why a trajectory match is the wrong test. This model is iterative and self-referential: each frame's denoising runs on this port's own previous output. Any per-step difference, including ordinary bf16 rounding, therefore makes the generated trajectory diverge from any single reference run, the way two chaotic systems do. Comparing generated-frame pixels against one reference trajectory gives correlations as low as 0.05-0.24. The bring-up investigation found that this is the wrong bar, not a defect.
- What the claim rests on: correctness of each individual step, correctness of each component, and visual inspection of real-image sessions.
- What it does not claim: a bit-exact trajectory match.
The full investigation is in BRINGUP_LOG.md.
Limitations
- Far from real time. It takes about 25 s per frame (0.04 FPS) against the model's 60 FPS target. No tracing, kernel fusion, batching across sigma steps or quantization has been attempted.
- One session per process. The KV-cache state lives on the same objects as the loaded weights. Starting a new session silently ends the previous one, and a stale ID returns
404. - Single chip only. There is no multi-chip or tensor-parallel path. Every hardware run used one chip of a P300c board. The
p150profile states the requirement (one chip); it has not been tested on a physical P150 board. - Mouse and scroll control only. The 256-wide button input is always zero, because no control-scheme mapping has been published.
zoomhas sign-only semantics. - The 24-layer correlation (0.963) is below the 0.99 bar. The gap was diagnosed as bf16 compounding over a deep residual stack, with no localized logic bug. The in-repo
test_full_model.pyassert is left failing on purpose rather than weakened. - Resolution is fixed at 512x1024. Upstream's GPU numbers are at 720p, so the performance comparison above is not a controlled one.
- The pull over-fetches. The server loads only
transformer/config.json,transformer/diffusion_pytorch_model.safetensorsandvae/diffusion_pytorch_model.safetensors, at the pinned revision. The bundle's weight pull has no file filter, sopull --with-weightsalso downloads about 4 GB the server never reads: the rootmodel.safetensorsandassets/.
Risks and safety considerations
- Drift over long sessions. Per-step bf16 noise accumulates. How many frames stay usable has not been measured. In a 16-frame
forwardsession from the reference seed:- The view settled into a close-up grass texture by frame 4.
- It lost contrast by frame 16: pixel standard deviation 42-50 on frames 1-8, 14 on frame 16.
- Out-of-distribution seeds give noise. A synthetic or random-static seed image produces unstructured output. The upstream reference behaves the same way under that input. Seed with a real, natural image.
- Numerical deviation from the reference. ttnn SDPA has no fp32 mode. This port uses bf16 with fp32 accumulation (
fp32_dest_acc_en, HiFi4). Outputs are plausible but will not match GPU output pixel for pixel. - Generated video is synthetic. It is not a faithful simulation of any real place.
Licensing
- Weights:
Overworld/Waypoint-1.5-1B, Apache-2.0. - Port and serving code: the
waypoint_ttnnwheel, Apache-2.0 (repo LICENSE). - Vendored tt-metal code:
tt_waypoint_models_closure, which ismodels/common/lightweightmodule.pyfrom tt-metal v0.78.0. Apache-2.0.
Changelog
| date | change |
|---|---|
| 2026-09-27 | Repackaged with waypoint-ttnn 0.1.1. The server now loads weights at the pinned revision 391f928, overridable with TT_MODEL_WEIGHTS_REVISION, and fetches only the three files it reads. The manifest records that revision and tt_metal_version 0.78.0. run.sh exports the revision and picks chips from the bundle's device count. ttnn==0.78.0 is unchanged. Re-verified on hardware from a fresh install with an empty HF cache: 25.72 s per frame, coherent frames. |
| 2026-09-15 | Repackaged as a v6 thin bundle (pip/venv), replacing the v5.1 container image. The served code is unchanged. ttnn==0.78.0 comes from PyPI, and torch==2.14.0+cpu is pinned. Re-verified on hardware: mesh open, fresh weight fetch, seed and step, coherent frames. |
| 2026-09-10 | Benchmarked: 27.41 s per frame, flat across session length. Re-verified by a fresh pull and serve of the published package across several directions. |
| 2026-09-09 | First publish (v5.1 container). The card's port was corrected to 20000. session.py now resolves weights via snapshot_download instead of a host-only path. |
Related packages
None. This is the only Tenstorrent port of Waypoint-1.5-1B.
Feedback
- Questions or problems with this package: open a discussion at
https://huggingface.co/episod/tt-waypoint/discussions. That is the one channel that reaches the bundle's author. - Problems with the
tttooling itself: runtt report issue. It collects your environment and opens a prefilled issue against tenstorrent/tt-cli; it does not reach this package's author. - Product feedback: [email protected].
Provenance
The served-path code is split into two small wheels built for this bundle, both in wheels/:
| wheel | what it is |
|---|---|
waypoint_ttnn-0.1.1 |
The served-path closure of tsingletaryTT/tt-waypoint: session.py, the ASGI app, and the tt/ model classes. Bring-up scripts, tests and docs are not included; see the GitHub repo for those |
tt_waypoint_models_closure-0.78.0 |
The one tt-metal dependency this model has: models/common/lightweightmodule.py, pinned to tt-metal v0.78.0 |
| component | built from |
|---|---|
| tt-metal / ttnn | ttnn==0.78.0 (PyPI). The manifest tt_metal_version is 0.78.0 |
| model code | tsingletaryTT/tt-waypoint @ adc1b56. The wheel's .py files are identical to this commit |
| weights | Overworld/Waypoint-1.5-1B @ 391f92827075edcf4a8b3c8a2ddae010698f8636, pinned in the manifest and in session.py |
| Model CI v0 | not yet run |
| build | 2026-09-27 · tt-model-manager (producer tt_kernel_version 0.1.0) |
Model tree for episod/tt-waypoint
Base model
Overworld/Waypoint-1.5-1B