Anima Turbo for the Snapdragon 8 Elite neural unit

Context binaries (QAIRT 2.50.40, HTP v79, soc_model 69) that run circlestone-labs/Anima Turbo v1.1 entirely on the Snapdragon 8 Elite neural processing unit: prompt encoder, 28-block drawing model and picture decoder. 6 picture sizes (512 and 768 class; square, 9:16, 16:9), 4/6/8 steps, about 6โ€“7 s per 512ร—512 picture at 8 steps on an iQOO 13. Optional guide picture (layout/pose control) through kohya-ss/Anima-LLLite any-test-like-v2, built into the drawing model.

Non-commercial use only โ€” see LICENSE.md and NOTICE. Generated pictures may be used commercially. This is a modified version of Anima and is not an official CircleStone Labs release.

Files (models/, copy all of them into the app's files/models folder)

File What it is Precision
text.bin Qwen3-0.6B + Anima's 6-block adapter (prompt โ†’ 256ร—1024 context) 8-bit weights in blocks of 32, fp16 maths
part0.bin โ€ฆ part3.bin the drawing model, 7 blocks each 8-bit weights, fp16 maths
dec.bin tiny Wan 2.1 decoder (taew2_1), single picture fp16
xlsr.bin optional 4ร— upscaler (512 โ†’ 2048) fp16
qwen_embed_f16.raw Qwen token table, divided by 64 (the encoder's residual scale) fp16
t5_embed_f16.raw Anima adapter's T5 token table fp16
qwen_tokenizer.json, t5_tokenizer.json tokenizers โ€”

Low memory set (what the app's "Low memory" choice downloads, about 2.4 GB instead of 3.0 GB):

  • extra/draw_m4/part0-3.bin โ€” drawing model with feed-forward layers 4-bit, rest 8-bit (1.62 GB);
  • extra/low/text.bin โ€” 4-bit prompt encoder with a ConvRot/QuaRot-style rotation baked into the weights (regular Hadamard: residual stream 1024, attention heads 2ร—64; final un-rotation folded into the adapter);
  • extra/low/qwen_embed_f16.raw โ€” the matching rotated token table (it must be used with that encoder). Pictures vary more from the Standard ones.

manifest.json lists every file with its size and sha256 per set (the original 512-only set, kept for older app builds). manifest2.json is the current set: the multi-size, guide-capable files under ms/:

File What it is
ms/standard/part0-3.bin drawing model, 8-bit weights, all 6 sizes in one file per piece (weights stored once, one graph s<W>x<H> per size), with the guide adapter inside; extra inputs cemb (guide embedding) and st (strength, 0 = no guide)
ms/low/part0-3.bin the same with the feed-forward layers 4-bit
ms/dec.bin the tiny decoder for all 6 sizes
ms/guide.bin the guide reader (any-test-like-v2's conditioning network), all 6 sizes: grey picture 0..1 โ†’ cemb

Sizes (width ร— height): 512ร—512, 384ร—672, 672ร—384; 768ร—768, 576ร—1024, 1024ร—576. Every size has a patch count (w/16 ร— h/16) divisible by 16 โ€” other counts compile into programs about 4ร— slower on this chip. The guide works best for the first half of the steps (ComfyUI end_percent 0.5) at strength ~1.

manifest2.json also uses 8-bit word tables: models/qwen_embed_q8.raw and extra/low/qwen_embed_q8.raw (151936 ร— 1024 int8, then 151936 float32 row scales; value = int8 ร— scale). Half the size of the fp16 tables; pictures match the fp16 tables at 0.995-0.999 (final latent) on the phone.

These binaries only run on HTP v79 (Snapdragon 8 Elite). Other chips need the pieces recompiled.

How the pieces fit together

  • Prompt: Qwen ids (no special tokens) and T5 ids (+ end token 1) are looked up in the two tables on the CPU; text.bin takes them with padding masks and returns the context (rows past the prompt are zero; prompts up to 255 T5 tokens).
  • Drawing: each Euler step runs part0 โ†’ part3 on the 1024 patches (32ร—32 of a 64ร—64 latent); sigmas from ComfyUI's "simple" schedule with shift 3. The residual stream is divided by 256 inside the weights so it fits in fp16.
  • dec.bin turns the final latent (as is, no latent scaling) into the 512ร—512 picture.

The app, and building these files yourself

These files are made for the Anima app, which downloads them itself. The scripts that build them from the original model weights, with a step-by-step guide, are in build/ in this repo.

Downloads last month
29
GGUF
Model size
8B params
Architecture
bailingmoe3
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for sebekblessing/anima-qnn-8elite

Quantized
(48)
this model