# SCVI and SCANVI for label transfer how to assess accuracy?

**URL:** <https://discourse.scverse.org/t/scvi-and-scanvi-for-label-transfer-how-to-assess-accuracy/2456>\
**Category:** scvi-tools\
**Tags:** scanvi\
**Created:** [August 18, 2024, 2:08pm UTC](https://discourse.scverse.org/t/scvi-and-scanvi-for-label-transfer-how-to-assess-accuracy/2456 "2024-08-18T14:08:05Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Flu09](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/flu09/32/1140_2.png) [@Flu09](https://discourse.scverse.org/u/Flu09)\
**Post date:** [August 18, 2024, 2:08pm UTC](https://discourse.scverse.org/t/scvi-and-scanvi-for-label-transfer-how-to-assess-accuracy/2456/1 "2024-08-18T14:08:05Z")

</div>

Hello, I followed this tutorial [sanbomics\_scripts/scvi\_label\_transfer.ipynb at main · mousepixels/sanbomics\_scripts · GitHub](https://github.com/mousepixels/sanbomics_scripts/blob/main/scvi_label_transfer.ipynb)

where it first concatenated the sample(s) of unknown labels with the reference. Now ref and my samples have different batches then I trained vae=scvi.model.SCVI(adata) without any extra arguments and it ran for 6 epoch only ( data is about 1.4 million) so I wonder if 6 epoch is a good number and how to assess that the model is accurate??

then lvae = scvi.model.SCANVI.from\_scvi\_model(vae, adata = adata, unlabeled\_category = ‘Unknown’,  
labels\_key = ‘cell\_ontology\_class’)

lvae.train(max\_epochs=20, n\_samples\_per\_label=100)

was done to predict the labels of my samples which have unknown. I want to ask what does the n\_samples\_per\_label mean ? what I understand is that it takes representative cells for each label in this case 100 cells. those representative cells from the unknown cells? or what?

I would appreciate it if you help regarding this method

---

<div class="post-metadata">

**Author:** ![cane11](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/cane11/32/241_2.png) [@cane11](https://discourse.scverse.org/u/cane11)\
**Post date:** [August 18, 2024, 3:12pm UTC](https://discourse.scverse.org/t/scvi-and-scanvi-for-label-transfer-how-to-assess-accuracy/2456/2 "2024-08-18T15:12:23Z")

</div>

I tend to train for at least 20 epochs. However, this is more an experience based thing. You should check elbo\_validation and elbo\_train afterwards. You can increase batch\_size to 1024 (increases runtime by a factor of 8). Yes it takes 1000 representative cells for each celltype (or if there are less than 100 cells of a celltype all cells of this type). The classifier doesn’t have balanced weights and this helps with balancing.

---

<div class="post-metadata">

**Author:** ![Flu09](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/flu09/32/1140_2.png) [@Flu09](https://discourse.scverse.org/u/Flu09)\
**Post date:** [August 20, 2024, 12:13pm UTC](https://discourse.scverse.org/t/scvi-and-scanvi-for-label-transfer-how-to-assess-accuracy/2456/3 "2024-08-20T12:13:24Z")

</div>

Thank you.  
strangely I do not have elbo\_validation. May be I need to add another argument?

> > > vae.history.keys()  
> > > dict\_keys([‘kl\_weight’, ‘train\_loss\_step’, ‘train\_loss\_epoch’, ‘elbo\_train’, ‘reconstruction\_loss\_train’, ‘kl\_local\_train’, ‘kl\_global\_train’])

> > > vae  
> > > SCVI model with the following parameters:  
> > > n\_hidden: 128, n\_latent: 10, n\_layers: 1, dropout\_rate: 0.1, dispersion: gene,  
> > > gene\_likelihood: zinb, latent\_distribution: normal.  
> > > Training status: Trained  
> > > Model’s adata is minified?: False

> > > lvae.history.keys()  
> > > dict\_keys([‘train\_loss\_step’, ‘train\_loss\_epoch’, ‘elbo\_train’, ‘reconstruction\_loss\_train’, ‘kl\_local\_train’, ‘kl\_global\_train’, ‘train\_classification\_loss’, ‘train\_accuracy’, ‘train\_f1\_score’, ‘train\_calibration\_error’])

> > > lvae  
> > > ScanVI Model with the following params:  
> > > unlabeled\_category: Unknown, n\_hidden: 128, n\_latent: 10, n\_layers: 1,  
> > > dropout\_rate: 0.1, dispersion: gene, gene\_likelihood: zinb  
> > > Training status: Trained  
> > > Model’s adata is minified?: False

---

<div class="post-metadata">

**Author:** ![cane11](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/cane11/32/241_2.png) [@cane11](https://discourse.scverse.org/u/cane11)\
**Post date:** [August 20, 2024, 5:25pm UTC](https://discourse.scverse.org/t/scvi-and-scanvi-for-label-transfer-how-to-assess-accuracy/2456/4 "2024-08-20T17:25:37Z")

</div>

Can you add check\_val\_every\_n\_epoch=1 to train? This enables tracking validation losses.
