Instructions to use sebekblessing/anima-qnn-8elite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use sebekblessing/anima-qnn-8elite with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf sebekblessing/anima-qnn-8elite:Q4_0 # Run inference directly in the terminal: llama cli -hf sebekblessing/anima-qnn-8elite:Q4_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf sebekblessing/anima-qnn-8elite:Q4_0 # Run inference directly in the terminal: llama cli -hf sebekblessing/anima-qnn-8elite:Q4_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf sebekblessing/anima-qnn-8elite:Q4_0 # Run inference directly in the terminal: ./llama-cli -hf sebekblessing/anima-qnn-8elite:Q4_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf sebekblessing/anima-qnn-8elite:Q4_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf sebekblessing/anima-qnn-8elite:Q4_0
Use Docker
docker model run hf.co/sebekblessing/anima-qnn-8elite:Q4_0
- LM Studio
- Jan
- Ollama
How to use sebekblessing/anima-qnn-8elite with Ollama:
ollama run hf.co/sebekblessing/anima-qnn-8elite:Q4_0
- Unsloth Desktop
- Pi
How to use sebekblessing/anima-qnn-8elite with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sebekblessing/anima-qnn-8elite:Q4_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "sebekblessing/anima-qnn-8elite:Q4_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use sebekblessing/anima-qnn-8elite with Docker Model Runner:
docker model run hf.co/sebekblessing/anima-qnn-8elite:Q4_0
- Lemonade
How to use sebekblessing/anima-qnn-8elite with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull sebekblessing/anima-qnn-8elite:Q4_0
Run and chat with the model
lemonade run user.anima-qnn-8elite-Q4_0
List all available models
lemonade list
- Hermes Agent
How to use sebekblessing/anima-qnn-8elite with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sebekblessing/anima-qnn-8elite:Q4_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default sebekblessing/anima-qnn-8elite:Q4_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use sebekblessing/anima-qnn-8elite with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sebekblessing/anima-qnn-8elite:Q4_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "sebekblessing/anima-qnn-8elite:Q4_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Anima Turbo for the Snapdragon 8 Elite neural unit
Context binaries (QAIRT 2.50.40, HTP v79, soc_model 69) that run circlestone-labs/Anima Turbo v1.1 entirely on the Snapdragon 8 Elite neural processing unit: prompt encoder, 28-block drawing model and picture decoder. 6 picture sizes (512 and 768 class; square, 9:16, 16:9), 4/6/8 steps, about 6โ7 s per 512ร512 picture at 8 steps on an iQOO 13. Optional guide picture (layout/pose control) through kohya-ss/Anima-LLLite any-test-like-v2, built into the drawing model.
Non-commercial use only โ see LICENSE.md and NOTICE. Generated pictures may be used commercially.
This is a modified version of Anima and is not an official CircleStone Labs release.
Files (models/, copy all of them into the app's files/models folder)
| File | What it is | Precision |
|---|---|---|
text.bin |
Qwen3-0.6B + Anima's 6-block adapter (prompt โ 256ร1024 context) | 8-bit weights in blocks of 32, fp16 maths |
part0.bin โฆ part3.bin |
the drawing model, 7 blocks each | 8-bit weights, fp16 maths |
dec.bin |
tiny Wan 2.1 decoder (taew2_1), single picture | fp16 |
xlsr.bin |
optional 4ร upscaler (512 โ 2048) | fp16 |
qwen_embed_f16.raw |
Qwen token table, divided by 64 (the encoder's residual scale) | fp16 |
t5_embed_f16.raw |
Anima adapter's T5 token table | fp16 |
qwen_tokenizer.json, t5_tokenizer.json |
tokenizers | โ |
Low memory set (what the app's "Low memory" choice downloads, about 2.4 GB instead of 3.0 GB):
extra/draw_m4/part0-3.binโ drawing model with feed-forward layers 4-bit, rest 8-bit (1.62 GB);extra/low/text.binโ 4-bit prompt encoder with a ConvRot/QuaRot-style rotation baked into the weights (regular Hadamard: residual stream 1024, attention heads 2ร64; final un-rotation folded into the adapter);extra/low/qwen_embed_f16.rawโ the matching rotated token table (it must be used with that encoder). Pictures vary more from the Standard ones.
manifest.json lists every file with its size and sha256 per set (the original 512-only set, kept for older
app builds). manifest2.json is the current set: the multi-size, guide-capable files under ms/:
| File | What it is |
|---|---|
ms/standard/part0-3.bin |
drawing model, 8-bit weights, all 6 sizes in one file per piece (weights stored once, one graph s<W>x<H> per size), with the guide adapter inside; extra inputs cemb (guide embedding) and st (strength, 0 = no guide) |
ms/low/part0-3.bin |
the same with the feed-forward layers 4-bit |
ms/dec.bin |
the tiny decoder for all 6 sizes |
ms/guide.bin |
the guide reader (any-test-like-v2's conditioning network), all 6 sizes: grey picture 0..1 โ cemb |
Sizes (width ร height): 512ร512, 384ร672, 672ร384; 768ร768, 576ร1024, 1024ร576. Every size has a patch count (w/16 ร h/16) divisible by 16 โ other counts compile into programs about 4ร slower on this chip. The guide works best for the first half of the steps (ComfyUI end_percent 0.5) at strength ~1.
manifest2.json also uses 8-bit word tables: models/qwen_embed_q8.raw and extra/low/qwen_embed_q8.raw
(151936 ร 1024 int8, then 151936 float32 row scales; value = int8 ร scale). Half the size of the fp16 tables;
pictures match the fp16 tables at 0.995-0.999 (final latent) on the phone.
These binaries only run on HTP v79 (Snapdragon 8 Elite). Other chips need the pieces recompiled.
How the pieces fit together
- Prompt: Qwen ids (no special tokens) and T5 ids (+ end token 1) are looked up in the two tables on the
CPU;
text.bintakes them with padding masks and returns the context (rows past the prompt are zero; prompts up to 255 T5 tokens). - Drawing: each Euler step runs part0 โ part3 on the 1024 patches (32ร32 of a 64ร64 latent); sigmas from ComfyUI's "simple" schedule with shift 3. The residual stream is divided by 256 inside the weights so it fits in fp16.
dec.binturns the final latent (as is, no latent scaling) into the 512ร512 picture.
The app, and building these files yourself
These files are made for the Anima app, which downloads them itself.
The scripts that build them from the original model weights, with a step-by-step guide, are in
build/ in this repo.
- Downloads last month
- 29
4-bit
8-bit
Model tree for sebekblessing/anima-qnn-8elite
Base model
nvidia/Cosmos-Predict2-2B-Text2Image