QuPAINT-4B

Physics-Aware Instruction Tuning Approach to Quantum Material Discovery IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026 — Findings Track

Project page · Paper (PDF) · Code · Data

QuPAINT is a vision-language model for analyzing two-dimensional (2D) quantum material flakes in optical micrographs. Given a micrograph and a question, it enumerates candidate flakes, reasons about their optical contrast against the substrate, and commits to a set of bounding boxes — for example, which flakes are monolayer candidates.

The model addresses the domain shift that makes 2D-material identification hard in practice: flake appearance depends not only on material and thickness but on substrate, illumination, camera response, focus, and noise, so models trained in one lab degrade in another. QuPAINT fuses visual embeddings with optical priors via Physics-Informed Attention and is instruction-tuned on QMat-Instruct, a physics-informed multimodal instruction dataset with reasoning traces grounded in observable optical evidence.

Model details

Architecture vision transformer encoder + MLP projector + 4B-parameter language model
Parameters ~4B
Precision bfloat16
Input optical micrograph (resized to a 1792×1344 canvas, tiled into 448px patches, up to 12 tiles + thumbnail)
Output free-text analysis ending in a <CONCLUSION> span of <box> quadruples in percent coordinates
Developed by CVIU Lab, University of Arkansas

Usage

Install the requirements (torch, torchvision, transformers>=4.51, Pillow, accelerate), then:

import torch
from transformers import AutoModel, AutoTokenizer

ckpt = "uark-cviu/QuPAINT-4B"
model = AutoModel.from_pretrained(
    ckpt,
    dtype=torch.bfloat16,
    low_cpu_mem_usage=True,
    trust_remote_code=True,
    use_flash_attn=False,        # set True if flash-attn is installed
).eval().cuda()
tokenizer = AutoTokenizer.from_pretrained(ckpt, trust_remote_code=True)

# preprocess.py ships with this repo; it applies the tuning-time canvas and tiling
from huggingface_hub import hf_hub_download
import importlib.util, sys
spec = importlib.util.spec_from_file_location(
    "qupaint_preprocess", hf_hub_download(ckpt, "preprocess.py")
)
preprocess = importlib.util.module_from_spec(spec); spec.loader.exec_module(preprocess)

pixel_values = preprocess.load_image("micrograph.jpg").to(torch.bfloat16).cuda()
response = model.chat(
    tokenizer,
    pixel_values,
    "<image>\nIdentify monolayer candidates and provide bounding boxes [x,y,w,h].",
    dict(max_new_tokens=4096, do_sample=False),
)
print(response)

For a CLI, a Gradio demo, and box parsing/plotting helpers, use the GitHub repository:

git clone https://github.com/uark-cviu/QuPAINT && cd QuPAINT
pip install -r requirements.txt
python demo.py --checkpoint uark-cviu/QuPAINT-4B

Prompts

Goal Prompt
Detect everything Identify all flakes and provide bounding boxes [x,y,w,h] for each
Monolayer candidates Identify monolayer candidates and provide bounding boxes [x,y,w,h]
Counting How many material flakes are there in the image?
Optical reasoning Describe the optical contrast of the flakes relative to the substrate

Output format

All flakes are collected first as <box>59.76, 6.88, 4.74, 4.44</box> <box>42.52, 50.68, 7.39, 9.08</box> ...
Monolayer flakes show lighter contrast and subtle color shift relative to the substrate.
<CONCLUSION>
The monolayer candidates are at: <box>42.52, 50.68, 7.39, 9.08</box>
</CONCLUSION>

Boxes are x, y, width, height as percent of image width/height with a top-left origin, so they map back onto an image of any size. Flakes listed before the <CONCLUSION> span are candidates the model considered; the span holds what it commits to for the question asked.

Greedy decoding (do_sample=False) makes repeated runs on the same micrograph reproducible.

Intended use and limitations

Intended use. Non-commercial academic research on automated characterization of 2D quantum materials: flake detection, monolayer screening, counting, visual grounding, and optical-contrast reasoning in exfoliation workflows.

Limitations.

  • Trained on optical micrographs of exfoliated flakes, largely on SiO₂/Si substrates. Very different substrates, magnifications, or modalities are out of distribution.
  • Layer assignment is inferred from optical contrast alone. The model never observes Raman, AFM, or photoluminescence signals, and its judgments should be confirmed by those methods before being relied on.
  • It is a generative model: it can miss faint flakes or propose spurious ones, and box coordinates are approximate. Treat outputs as candidate proposals for a human or a downstream measurement tool, not as ground-truth measurements.
  • Counting and area figures stated in prose are not computed arithmetically; derive such numbers from the parsed boxes instead.

Training data

Instruction-tuned on QMat-Instruct, built from Synthia-generated synthetic microscopy images with layer-dependent optical behavior, supervised with annotation-conditioned reasoning traces restricted to observable optical cues. See the paper and the project page for details.

Evaluation

The paper evaluates QuPAINT on QF-Bench — 8,854 images, 280,526 annotated flakes across eight materials (BN, Graphene, MoS₂, MoSe₂, MoWSe₂, WS₂, WSe₂, WTe₂), labeled mono-layer (1L), few-layer (2–4L), and thick (5+L) — covering detection, counting, visual grounding, and image-specific reasoning, including generalization to a material excluded from training. QF-Bench is released separately; see the project page for status.

Citation

@InProceedings{nguyen2026qupaint,
    author    = {Nguyen, Xuan Bac and Nguyen, Hoang-Quan and Pandey, Sankalp and Faltermeier, Tim and Borys, Nicholas and Churchill, Hugh and Luu, Khoa},
    title     = {QuPAINT: Physics-Aware Instruction Tuning Approach to Quantum Material Discovery},
    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings},
    month     = {June},
    year      = {2026},
    pages     = {8684--8694}
}

Authors

Xuan Bac Nguyen¹, Hoang-Quan Nguyen¹, Sankalp Pandey¹, Tim Faltermeier², Nicholas Borys², Hugh Churchill³, Khoa Luu¹

¹ CVIU Lab, University of Arkansas · ² University of Utah · ³ Department of Physics, University of Arkansas

Acknowledgements

Partly supported by the MonArk NSF Quantum Foundry (DMR-1906383) and an NSF Quantum Award (2444042), with GPU resources from the Arkansas High-Performance Computing Center.

License

These weights are released for non-commercial academic research only (see LICENSE). The inference code on GitHub is MIT licensed.

Downloads last month
48
Safetensors
Model size
590k params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support