# Label transfer with SCVI-SCANVI pipeline changes (predicts wrong) labels in ref data

**URL:** <https://discourse.scverse.org/t/label-transfer-with-scvi-scanvi-pipeline-changes-predicts-wrong-labels-in-ref-data/945>\
**Category:** scvi-tools\
**Tags:** scanvi, scvi\
**Created:** [November 21, 2022, 4:42pm UTC](https://discourse.scverse.org/t/label-transfer-with-scvi-scanvi-pipeline-changes-predicts-wrong-labels-in-ref-data/945 "2022-11-21T16:42:05Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Avaptel18](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/avaptel18/32/448_2.png) [@Avaptel18](https://discourse.scverse.org/u/Avaptel18)\
**Post date:** [November 21, 2022, 4:42pm UTC](https://discourse.scverse.org/t/label-transfer-with-scvi-scanvi-pipeline-changes-predicts-wrong-labels-in-ref-data/945/1 "2022-11-21T16:42:05Z")

</div>

Hi, I am following this tutorial " Integration and label transfer with Tabula Muris" [Integration and label transfer with Tabula Muris - scvi-tools](https://docs.scvi-tools.org/en/latest/tutorials/notebooks/tabula_muris.html) on SCVI-tools docs page and everything works fine and I save the predicted labels in the new metadata column

adata.obs[“C\_scANVI”] = lvae.predict(adata) #saving predicted labels in new column

However when I check the ref labels, some of them are predicted differently than what it was before I trained SCVI model (when I concatenated with my query data).  
I don’t understand why this is?  
My understanding is that I am training the SCVI model on ref labelled data and then using SCANVI to transfer the labels on ‘unknown’ labels in query dataset.  
Why is it predicting some of the ref data labels wrongly?  
Any advise please. Am I doing something wrong here?  
Thanks!

As you can see in the screenshot the ref data cell label ‘Cell cycle\_TCGGTCTGTGAGAGGG-1\_32\_1-1’ is changed from Trm-c to CTL-c.

 ![Screen Shot 2022-11-21 at 15.50.24](https://canada1.discourse-cdn.com/flex035/uploads/forum11/original/1X/56da9c73ccefb67d61aa218a38083470d72f5978.png)

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [December 2, 2022, 5:34am UTC](https://discourse.scverse.org/t/label-transfer-with-scvi-scanvi-pipeline-changes-predicts-wrong-labels-in-ref-data/945/2 "2022-12-02T05:34:32Z")

</div>

Can you describe how many of the training labels are wrong?

The predict function makes a prediction for each cell, including the reference data, for which by default 90% is train and 10% is a validation set. Either scanvi is getting it wrong because there’s something systematically off and/or there is noise in the training labels.

---

<div class="post-metadata">

**Author:** ![Avaptel18](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/avaptel18/32/448_2.png) [@Avaptel18](https://discourse.scverse.org/u/Avaptel18)\
**Post date:** [December 6, 2022, 5:33pm UTC](https://discourse.scverse.org/t/label-transfer-with-scvi-scanvi-pipeline-changes-predicts-wrong-labels-in-ref-data/945/3 "2022-12-06T17:33:13Z")

</div>

Hi Adam,  
SCANVI predicted 35.2% labels wrong in the reference dataset.

Here is some of my code, probably not enough for you to figure out what’s wrong but providing here anyway, just in case if there is any obvious mistake.

#################  
adata.layers[“counts”] = adata.X.copy() #preserve count layer  
sc.pp.normalize\_total(adata, target\_sum=1e4)  
sc.pp.log1p(adata)  
adata.raw = adata

sc.pp.highly\_variable\_genes(adata,  
flavor = ‘seurat\_v3’,  
n\_top\_genes=3000, #3000 hvg selected  
layer = “counts”,  
batch\_key=“batch”,  
subset = True)

scvi.model.SCVI.setup\_anndata(adata, layer=“counts”,  
batch\_key=“batch”)  
vae = scvi.model.SCVI(adata, n\_layers=4, n\_latent=30)  
vae.train()

vae  
 SCVI Model with the following params:  
n\_hidden: 128, n\_latent: 30, n\_layers: 4, dropout\_rate: 0.1, dispersion: gene,  
gene\_likelihood: zinb, latent\_distribution: normal  
Training status: Trained

## Transfer of annotations with scANVI

adata.obs[“celltype\_scanvi”] = ‘Unknown’  
ss2\_idx = adata.obs[‘batch’] == “1”  
adata.obs[“celltype\_scanvi”][ss2\_idx] = adata.obs.Ident2[ss2\_idx]

scvi.model.SCANVI.setup\_anndata(adata,  
layer=“counts”,  
batch\_key=“batch”,  
labels\_key=“celltype\_scanvi”,  
unlabeled\_category=“Unknown”)

lvae = scvi.model.SCANVI.from\_scvi\_model(vae, “Unknown”,  
adata=adata,  
labels\_key=“celltype\_scanvi”)

lvae.train(max\_epochs=20, n\_samples\_per\_label=100)

lave  
 ScanVI Model with the following params:  
unlabeled\_category: Unknown, n\_hidden: 128, n\_latent: 30, n\_layers: 4, dropout\_rate: 0.1,  
dispersion: gene, gene\_likelihood: zinb  
Training status: Trained

#################

Am I missing something here?  
Please help!  
Thank you.

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [December 11, 2022, 1:13am UTC](https://discourse.scverse.org/t/label-transfer-with-scvi-scanvi-pipeline-changes-predicts-wrong-labels-in-ref-data/945/4 "2022-12-11T01:13:22Z")

</div>

It’s hard for me to diagnose without understanding what kinds of mistakes it’s making. Is it predicting random or related cell types?

Can you try training the scanvi part for longer?

---

<div class="post-metadata">

**Author:** ![Avaptel18](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/avaptel18/32/448_2.png) [@Avaptel18](https://discourse.scverse.org/u/Avaptel18)\
**Post date:** [December 11, 2022, 12:00pm UTC](https://discourse.scverse.org/t/label-transfer-with-scvi-scanvi-pipeline-changes-predicts-wrong-labels-in-ref-data/945/5 "2022-12-11T12:00:31Z")

</div>

Hi Adam, It’s predicting all related cell types. I can try training scanvi for longer and see if it improves the prediction.  
Ideally what percentage of correct prediction I should get?  
Thanks!

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [December 12, 2022, 6:15am UTC](https://discourse.scverse.org/t/label-transfer-with-scvi-scanvi-pipeline-changes-predicts-wrong-labels-in-ref-data/945/6 "2022-12-12T06:15:01Z")

</div>

the accuracy on the labeled data should be near 100%. We should be able to expose the training accuracy in the model history to make this easier to check.

---

<div class="post-metadata">

**Author:** ![Avaptel18](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/avaptel18/32/448_2.png) [@Avaptel18](https://discourse.scverse.org/u/Avaptel18)\
**Post date:** [December 12, 2022, 3:42pm UTC](https://discourse.scverse.org/t/label-transfer-with-scvi-scanvi-pipeline-changes-predicts-wrong-labels-in-ref-data/945/7 "2022-12-12T15:42:07Z")

</div>

Hi Adam, that would be great! thanks.  
I have tried increasing max\_epochs for SCANVI and it decreased wrong prediction from 35 to 26% and I can try further increase but there seems to be something missing as I am training SCANVI with 36 label types and it only predicts about 18 now. I don’t understand why its omitting half cell types?  
I am specifying this option (n\_samples\_per\_label=100) but all ref labels have over 140 cells.  
Can I specify specific parameter so it predicts all cell types?  
Thanks!

I have 145000 cells in ref dataset with 36 immune cell types (skin cd45+ cells: [https://www.science.org/doi/10.1126/sciimmunol.abl9165?url\_ver=Z39.88-2003&rfr\_id=ori:rid:crossref.org&rfr\_dat=cr\_pub%20%200pubmed](https://www.science.org/doi/10.1126/sciimmunol.abl9165?url_ver=Z39.88-2003&rfr_id=ori:rid:crossref.org&rfr_dat=cr_pub%20%200pubmed)).  
My query dataset is also skin cd45+ 102000 cells and I think I should have all 36 cell types present in my dataset.

---

<div class="post-metadata">

**Author:** ![miyang](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/miyang/32/755_2.png) [@miyang](https://discourse.scverse.org/u/miyang)\
**Post date:** [July 31, 2023, 2:09pm UTC](https://discourse.scverse.org/t/label-transfer-with-scvi-scanvi-pipeline-changes-predicts-wrong-labels-in-ref-data/945/8 "2023-07-31T14:09:14Z")

</div>

hello,  
Is there a sort of pvalue or any kind of significance level for the predicted cell types ?  
Thanks !

---

<div class="post-metadata">

**Author:** ![martinkim0](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/martinkim0/32/881_2.png) [@martinkim0](https://discourse.scverse.org/u/martinkim0)\
**Post date:** [July 31, 2023, 3:47pm UTC](https://discourse.scverse.org/t/label-transfer-with-scvi-scanvi-pipeline-changes-predicts-wrong-labels-in-ref-data/945/9 "2023-07-31T15:47:39Z")

</div>

Hi, passing in `soft=True` to SCANVI’s [predict](https://docs.scvi-tools.org/en/stable/api/reference/scvi.model.SCANVI.html#scvi.model.SCANVI.predict) method returns prediction probabilities from the cell type classifier. However, these shouldn’t be interpreted as p-values or significance levels, nor are they typically well calibrated (_i.e._ the classifier is confident even when giving wrong predictions).
