Instructions to use ahmedheakl/rand-mobile-grpo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ahmedheakl/rand-mobile-grpo with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ahmedheakl/rand-mobile-grpo", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Mobile-O β GRPO checkpoint (v2mcptf-grpo-500)
The project's main GRPO checkpoint: 500 steps of GRPO from the v2mcptf-mixed SFT init, with a
unified OCR + GenEval reward. This is the baseline every later experiment was measured against.
For the best overall checkpoint see ahmedheakl/rand-mobile
(soup3-targets), which meets 3 of 6 targets to this one's 2. This checkpoint is nonetheless part of
that soup's lineage β it is one of the two task vectors inside ta-dpg-s1.0, which is one of the
three soup members.
Benchmarks
Best setting is cfg 3.0, 20 DPM-Solver++ steps (2 of 6 targets met):
| benchmark | cfg 3.0 (best) | cfg 1.5 (protocol) | target | status @ cfg 3.0 |
|---|---|---|---|---|
| GenEval | 0.9284 | 0.9116 | β₯ 0.90 | met |
| ImageReward (MJHQ-30K) | 0.9250 | 0.7633 | β₯ 0.90 | met |
| DPG-Bench | 81.727 | 80.679 | β₯ 85 | not met |
| FID (MJHQ-30K) | 15.398 | 13.508 | β€ 8 | not met |
| ImgEdit (local Qwen2.5-VL-72B judge) | 3.020 | 3.093 | β₯ 3.5 | not met |
| GEdit (EN, local Qwen2.5-VL-72B judge) | 6.620 | 6.520 | β₯ 6.7 | not met |
Other measured settings on this checkpoint: cfg 4.5 gives GenEval 0.9237 / DPG 82.331 / ImageReward
0.9727 (its best) but FID 17.165 and GEdit 6.420. Interval guidance (CFG applied only for
t β [0, 0.9]) at cfg 4.5 gives DPG 82.979 β this checkpoint's best DPG β at FID 15.959.
Note the FID/alignment trade: cfg 1.5 has the best FID (13.508) and the worst ImageReward (0.7633); cfg 4.5 inverts it. No setting of this checkpoint clears more than two targets.
What is in this repo
Only the trained head: the SANA DiT and the diffusion connector, 602 tensors
(model.dit.* 548, model.diffusion_connector.* 54). The vision-language model is frozen during
training and is not included here.
To run this you also need:
openbmb/MiniCPM-V-4_6β the frozen VLM encoderEfficient-Large-Model/Sana_600M_512px_diffusersβ for the DC-AE (f32, 32-channel, 16Γ16 latent at 512px) and the scheduler config
Connector type is mcptf, fusing 1 VLM layer (derive from fusion.layer_weights.shape[0]; a
vlm_num_layers field elsewhere in this project can incorrectly read 4 β the weights are
authoritative).
Inference settings used for every number above
| scheduler | DPM-Solver++, solver_order=2, flow_shift=3 |
| steps | 20 |
| guidance | cfg 3.0 for the best row, cfg 1.5 for the protocol row |
| null condition | the empty prompt through the same VLM+connector path (not a zero vector) |
| resolution | 512Γ512 |
For editing, the null condition is the instruction rather than the empty prompt.
Training
GRPO from SFT init Mobile-O-0.5B-SFT-minicpm-v2mcptf-mixed, 500 steps.
| reward | unified OCR + GenEval |
| group size | 12 |
| learning rate | 6e-5 |
| KL beta | 0.04 |
| clip range | 1e-4 |
| advantage clip | 5.0 |
| rollout | 10 denoise steps, cfg 1.0, shift 3.0, noise level 0.7 |
Caveats
- Four of six targets are not met (DPG, FID, ImgEdit, GEdit). This is the project baseline, not its best result.
- GEdit and ImgEdit are scored by a local Qwen2.5-VL-72B judge, not the GPT-4o leaderboard scale; the two are not comparable.
- The head alone is not a runnable model β see What is in this repo.
- Downloads last month
- 18