# Any suggestions for speeding up model training (vae and solo) on M2 mac

**URL:** <https://discourse.scverse.org/t/any-suggestions-for-speeding-up-model-training-vae-and-solo-on-m2-mac/2371>\
**Category:** scvi-tools\
**Created:** [July 6, 2024, 10:11pm UTC](https://discourse.scverse.org/t/any-suggestions-for-speeding-up-model-training-vae-and-solo-on-m2-mac/2371 "2024-07-06T22:11:21Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![pointwave](https://avatars.discourse-cdn.com/v4/letter/p/ed655f/32.png) [@pointwave](https://discourse.scverse.org/u/pointwave)\
**Post date:** [July 6, 2024, 10:11pm UTC](https://discourse.scverse.org/t/any-suggestions-for-speeding-up-model-training-vae-and-solo-on-m2-mac/2371/1 "2024-07-06T22:11:21Z")

</div>

Hello scverse!  
first time asking a question so let me know what I can improve  
I’ve been trying to train the vae and solo models but using my MPS gpu throws the same error mentioned in this post ([Error when training model on M3 Max MPS](https://discourse.scverse.org/t/error-when-training-model-on-m3-max-mps/1896)) so I’ve been going cpu only. It is absolutely slow because of the dataset size (700000 x 35000), but I was wondering if you all had any suggestions for things I could do to make sure this is going at the max possible speed.

I’ve been running this code

```python
scvi.settings.dl_num_workers = 11
scvi.settings.batch_size = 2048
scvi.settings.num_threads = 10

scvi.model.SCVI.setup_anndata(adata)
vae = scvi.model.SCVI(adata)
vae.train()

```

if it helps, here is the startup output of the code above

```python
GPU available: True (mps), used: False
TPU available: False, using: 0 TPU cores
IPU available: False, using: 0 IPUs
HPU available: False, using: 0 HPUs
/opt/miniconda3/envs/scanpy_env/lib/python3.9/site-packages/lightning/pytorch/trainer/setup.py:187: GPU available but not used. You can set it by doing `Trainer(accelerator='gpu')`.
/opt/miniconda3/envs/scanpy_env/lib/python3.9/site-packages/lightning/pytorch/trainer/connectors/data_connector.py:436: Consider setting `persistent_workers=True` in 'train_dataloader' to speed up the dataloader worker initialization.

```

Specs: M2 Max, 94gb ram,  
cpu usage during training: 50-65%  
ram usage during training: 40~gb

---

<div class="post-metadata">

**Author:** ![martinkim0](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/martinkim0/32/881_2.png) [@martinkim0](https://discourse.scverse.org/u/martinkim0)\
**Post date:** [July 8, 2024, 5:34pm UTC](https://discourse.scverse.org/t/any-suggestions-for-speeding-up-model-training-vae-and-solo-on-m2-mac/2371/2 "2024-07-08T17:34:29Z")

</div>

I haven’t tried this recently so not sure if it’s stable, but you can [install](https://developer.apple.com/metal/pytorch/) an MPS-supported version of PyTorch and then run use the MPS backend by passing in:

```auto
vae.train(accelerator="mps")

```
