Mobile-O β€” GRPO checkpoint (v2mcptf-grpo-500)

The project's main GRPO checkpoint: 500 steps of GRPO from the v2mcptf-mixed SFT init, with a unified OCR + GenEval reward. This is the baseline every later experiment was measured against.

For the best overall checkpoint see ahmedheakl/rand-mobile (soup3-targets), which meets 3 of 6 targets to this one's 2. This checkpoint is nonetheless part of that soup's lineage β€” it is one of the two task vectors inside ta-dpg-s1.0, which is one of the three soup members.

Benchmarks

Best setting is cfg 3.0, 20 DPM-Solver++ steps (2 of 6 targets met):

benchmark cfg 3.0 (best) cfg 1.5 (protocol) target status @ cfg 3.0
GenEval 0.9284 0.9116 β‰₯ 0.90 met
ImageReward (MJHQ-30K) 0.9250 0.7633 β‰₯ 0.90 met
DPG-Bench 81.727 80.679 β‰₯ 85 not met
FID (MJHQ-30K) 15.398 13.508 ≀ 8 not met
ImgEdit (local Qwen2.5-VL-72B judge) 3.020 3.093 β‰₯ 3.5 not met
GEdit (EN, local Qwen2.5-VL-72B judge) 6.620 6.520 β‰₯ 6.7 not met

Other measured settings on this checkpoint: cfg 4.5 gives GenEval 0.9237 / DPG 82.331 / ImageReward 0.9727 (its best) but FID 17.165 and GEdit 6.420. Interval guidance (CFG applied only for t ∈ [0, 0.9]) at cfg 4.5 gives DPG 82.979 β€” this checkpoint's best DPG β€” at FID 15.959.

Note the FID/alignment trade: cfg 1.5 has the best FID (13.508) and the worst ImageReward (0.7633); cfg 4.5 inverts it. No setting of this checkpoint clears more than two targets.

What is in this repo

Only the trained head: the SANA DiT and the diffusion connector, 602 tensors (model.dit.* 548, model.diffusion_connector.* 54). The vision-language model is frozen during training and is not included here.

To run this you also need:

  • openbmb/MiniCPM-V-4_6 β€” the frozen VLM encoder
  • Efficient-Large-Model/Sana_600M_512px_diffusers β€” for the DC-AE (f32, 32-channel, 16Γ—16 latent at 512px) and the scheduler config

Connector type is mcptf, fusing 1 VLM layer (derive from fusion.layer_weights.shape[0]; a vlm_num_layers field elsewhere in this project can incorrectly read 4 β€” the weights are authoritative).

Inference settings used for every number above

scheduler DPM-Solver++, solver_order=2, flow_shift=3
steps 20
guidance cfg 3.0 for the best row, cfg 1.5 for the protocol row
null condition the empty prompt through the same VLM+connector path (not a zero vector)
resolution 512Γ—512

For editing, the null condition is the instruction rather than the empty prompt.

Training

GRPO from SFT init Mobile-O-0.5B-SFT-minicpm-v2mcptf-mixed, 500 steps.

reward unified OCR + GenEval
group size 12
learning rate 6e-5
KL beta 0.04
clip range 1e-4
advantage clip 5.0
rollout 10 denoise steps, cfg 1.0, shift 3.0, noise level 0.7

Caveats

  • Four of six targets are not met (DPG, FID, ImgEdit, GEdit). This is the project baseline, not its best result.
  • GEdit and ImgEdit are scored by a local Qwen2.5-VL-72B judge, not the GPT-4o leaderboard scale; the two are not comparable.
  • The head alone is not a runnable model β€” see What is in this repo.
Downloads last month
18
Safetensors
Model size
0.6B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support