Instructions to use uark-cviu/QuPAINT-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use uark-cviu/QuPAINT-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="uark-cviu/QuPAINT-4B", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("uark-cviu/QuPAINT-4B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use uark-cviu/QuPAINT-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "uark-cviu/QuPAINT-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "uark-cviu/QuPAINT-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/uark-cviu/QuPAINT-4B
- SGLang
How to use uark-cviu/QuPAINT-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "uark-cviu/QuPAINT-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "uark-cviu/QuPAINT-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "uark-cviu/QuPAINT-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "uark-cviu/QuPAINT-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use uark-cviu/QuPAINT-4B with Docker Model Runner:
docker model run hf.co/uark-cviu/QuPAINT-4B
QuPAINT-4B
Physics-Aware Instruction Tuning Approach to Quantum Material Discovery IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026 — Findings Track
Project page · Paper (PDF) · Code · Data
QuPAINT is a vision-language model for analyzing two-dimensional (2D) quantum material flakes in optical micrographs. Given a micrograph and a question, it enumerates candidate flakes, reasons about their optical contrast against the substrate, and commits to a set of bounding boxes — for example, which flakes are monolayer candidates.
The model addresses the domain shift that makes 2D-material identification hard in practice: flake appearance depends not only on material and thickness but on substrate, illumination, camera response, focus, and noise, so models trained in one lab degrade in another. QuPAINT fuses visual embeddings with optical priors via Physics-Informed Attention and is instruction-tuned on QMat-Instruct, a physics-informed multimodal instruction dataset with reasoning traces grounded in observable optical evidence.
Model details
| Architecture | vision transformer encoder + MLP projector + 4B-parameter language model |
| Parameters | ~4B |
| Precision | bfloat16 |
| Input | optical micrograph (resized to a 1792×1344 canvas, tiled into 448px patches, up to 12 tiles + thumbnail) |
| Output | free-text analysis ending in a <CONCLUSION> span of <box> quadruples in percent coordinates |
| Developed by | CVIU Lab, University of Arkansas |
Usage
Install the requirements (torch, torchvision, transformers>=4.51, Pillow,
accelerate), then:
import torch
from transformers import AutoModel, AutoTokenizer
ckpt = "uark-cviu/QuPAINT-4B"
model = AutoModel.from_pretrained(
ckpt,
dtype=torch.bfloat16,
low_cpu_mem_usage=True,
trust_remote_code=True,
use_flash_attn=False, # set True if flash-attn is installed
).eval().cuda()
tokenizer = AutoTokenizer.from_pretrained(ckpt, trust_remote_code=True)
# preprocess.py ships with this repo; it applies the tuning-time canvas and tiling
from huggingface_hub import hf_hub_download
import importlib.util, sys
spec = importlib.util.spec_from_file_location(
"qupaint_preprocess", hf_hub_download(ckpt, "preprocess.py")
)
preprocess = importlib.util.module_from_spec(spec); spec.loader.exec_module(preprocess)
pixel_values = preprocess.load_image("micrograph.jpg").to(torch.bfloat16).cuda()
response = model.chat(
tokenizer,
pixel_values,
"<image>\nIdentify monolayer candidates and provide bounding boxes [x,y,w,h].",
dict(max_new_tokens=4096, do_sample=False),
)
print(response)
For a CLI, a Gradio demo, and box parsing/plotting helpers, use the GitHub repository:
git clone https://github.com/uark-cviu/QuPAINT && cd QuPAINT
pip install -r requirements.txt
python demo.py --checkpoint uark-cviu/QuPAINT-4B
Prompts
| Goal | Prompt |
|---|---|
| Detect everything | Identify all flakes and provide bounding boxes [x,y,w,h] for each |
| Monolayer candidates | Identify monolayer candidates and provide bounding boxes [x,y,w,h] |
| Counting | How many material flakes are there in the image? |
| Optical reasoning | Describe the optical contrast of the flakes relative to the substrate |
Output format
All flakes are collected first as <box>59.76, 6.88, 4.74, 4.44</box> <box>42.52, 50.68, 7.39, 9.08</box> ...
Monolayer flakes show lighter contrast and subtle color shift relative to the substrate.
<CONCLUSION>
The monolayer candidates are at: <box>42.52, 50.68, 7.39, 9.08</box>
</CONCLUSION>
Boxes are x, y, width, height as percent of image width/height with a top-left
origin, so they map back onto an image of any size. Flakes listed before the
<CONCLUSION> span are candidates the model considered; the span holds what it commits
to for the question asked.
Greedy decoding (do_sample=False) makes repeated runs on the same micrograph reproducible.
Intended use and limitations
Intended use. Non-commercial academic research on automated characterization of 2D quantum materials: flake detection, monolayer screening, counting, visual grounding, and optical-contrast reasoning in exfoliation workflows.
Limitations.
- Trained on optical micrographs of exfoliated flakes, largely on SiO₂/Si substrates. Very different substrates, magnifications, or modalities are out of distribution.
- Layer assignment is inferred from optical contrast alone. The model never observes Raman, AFM, or photoluminescence signals, and its judgments should be confirmed by those methods before being relied on.
- It is a generative model: it can miss faint flakes or propose spurious ones, and box coordinates are approximate. Treat outputs as candidate proposals for a human or a downstream measurement tool, not as ground-truth measurements.
- Counting and area figures stated in prose are not computed arithmetically; derive such numbers from the parsed boxes instead.
Training data
Instruction-tuned on QMat-Instruct, built from Synthia-generated synthetic microscopy images with layer-dependent optical behavior, supervised with annotation-conditioned reasoning traces restricted to observable optical cues. See the paper and the project page for details.
Evaluation
The paper evaluates QuPAINT on QF-Bench — 8,854 images, 280,526 annotated flakes across eight materials (BN, Graphene, MoS₂, MoSe₂, MoWSe₂, WS₂, WSe₂, WTe₂), labeled mono-layer (1L), few-layer (2–4L), and thick (5+L) — covering detection, counting, visual grounding, and image-specific reasoning, including generalization to a material excluded from training. QF-Bench is released separately; see the project page for status.
Citation
@InProceedings{nguyen2026qupaint,
author = {Nguyen, Xuan Bac and Nguyen, Hoang-Quan and Pandey, Sankalp and Faltermeier, Tim and Borys, Nicholas and Churchill, Hugh and Luu, Khoa},
title = {QuPAINT: Physics-Aware Instruction Tuning Approach to Quantum Material Discovery},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings},
month = {June},
year = {2026},
pages = {8684--8694}
}
Authors
Xuan Bac Nguyen¹, Hoang-Quan Nguyen¹, Sankalp Pandey¹, Tim Faltermeier², Nicholas Borys², Hugh Churchill³, Khoa Luu¹
¹ CVIU Lab, University of Arkansas · ² University of Utah · ³ Department of Physics, University of Arkansas
Acknowledgements
Partly supported by the MonArk NSF Quantum Foundry (DMR-1906383) and an NSF Quantum Award (2444042), with GPU resources from the Arkansas High-Performance Computing Center.
License
These weights are released for non-commercial academic research only (see LICENSE).
The inference code on GitHub is MIT licensed.
- Downloads last month
- 48