# scANVI relables known cells with known types incorrectly

**URL:** <https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98>\
**Category:** scvi-tools\
**Tags:** scanvi\
**Created:** [May 27, 2021, 12:21pm UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98 "2021-05-27T12:21:15Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![KevinMenden](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/kevinmenden/32/41_2.png) [@KevinMenden](https://discourse.scverse.org/u/KevinMenden)\
**Post date:** [May 27, 2021, 12:21pm UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/1 "2021-05-27T12:21:15Z")

</div>

Hi scvi-tools Team,

I have been trying out scVI and scANVI and am evaluating how good label transfer works. For my current test setup, I have 4 different datasets (human liver), which I have manually labeled. To test scANVI, I overwrite the labels for one dataset with “Unknown”. I pretty much just follow your tutorial for scANVI.

The label transfer works pretty nice for the dataset with the “Unknown” cell types. However, when I use the `SCANVI.predict()` function, it generates wrong labels for actually known cell types (so not marked as “Unknown”). And bad ones at that.

Now my question, is that expected, i.e. is scANVI supposed to predict labels for these cells as well? And the more difficult question, any idea why it behaves this way? I’m still assuming that I’m just doing something wrong but I can’t figure out what. I have tried both starting from a pre-trained scVI model and training a scANVI model from scratch. Any help would be highly appreciated!

I’ll put code and a figure below. It’s pretty clear when looking at the NKT cluster in `celltype_scanvi` and then comparing to the same cluster in `C_scANVI`. This cluster doesn’t contain unlabeled cells.

Cheers,  
Kevin

```python
adata.obs["celltype_scanvi"] = 'Unknown'
# Get the labels for datasets 0, 1, 2

batch_idx = adata.obs['batch'] == "0"
adata.obs["celltype_scanvi"][batch_idx] = adata.obs.celltype[batch_idx]

batch_idx = adata.obs['batch'] == "1"
adata.obs["celltype_scanvi"][batch_idx] = adata.obs.celltype[batch_idx]

batch_idx = adata.obs['batch'] == "2"
adata.obs["celltype_scanvi"][batch_idx] = adata.obs.celltype[batch_idx]

adata.obs['celltype_scanvi'] = adata.obs['celltype_scanvi'].astype("str")

np.unique(adata.obs["celltype_scanvi"], return_counts=True)

scvi.data.setup_anndata(
    adata,
    layer="counts",
    batch_key="batch",
    labels_key="celltype_scanvi",
)

lvae = scvi.model.SCANVI(adata, "Unknown", n_latent=30, n_layers=2)

lvae.train(n_samples_per_label=100)

adata.obs["C_scANVI"] = lvae.predict(adata)
adata.obsm["X_scANVI"] = lvae.get_latent_representation(adata)
sc.pp.neighbors(adata, use_rep="X_scANVI")
sc.tl.umap(adata)

sc.pl.umap(adata, color=["celltype_scanvi", "C_scANVI", "batch"], ncols=1, frameon=False)

```

 ![image](https://canada1.discourse-cdn.com/flex035/uploads/forum11/original/1X/3693b4636f0a0476baee54985c1f48ad8406c098.jpeg)

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [May 27, 2021, 4:10pm UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/2 "2021-05-27T16:10:07Z")

</div>

I’ll have to look at this again more closely later, but a few quick comments:

1. What version of scvi-tools are you using? In the latest version, the workflow is to now do some pre training with an `SCVI` model, which was done implicitly before, but we separated it out for code reasons.

```python
vae = scvi.model.SCVI(adata, n_layers=2, n_latent=30)
vae.train()
scanvi_model = scvi.model.SCANVI.from_scvi_model(vae, 'Unknown')
scanvi_model.train(25)

```

though I see now [this](https://docs.scvi-tools.org/en/stable/user_guide/notebooks/harmonization.html) tutorial was not properly updated with this workflow.

1. In your workflow, how many epochs is scanvi trained for?

> [@KevinMenden](#):
>
> Now my question, is that expected, i.e. is scANVI supposed to predict labels for these cells as well?

Yes, though the accuracy should be higher.

---

<div class="post-metadata">

**Author:** ![KevinMenden](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/kevinmenden/32/41_2.png) [@KevinMenden](https://discourse.scverse.org/u/KevinMenden)\
**Post date:** [May 28, 2021, 10:44am UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/3 "2021-05-28T10:44:55Z")

</div>

Thanks for the answer!

I’m using version 0.10.0

I have tried both, training an `SCVI` model before and then starting from that, or training a SCANVI model directly.  
In the former case, it trains for ~60 epochs SCVI and then ~7 epochs SCANVI. In the latter case just ~60 epochs SCANVI.

Okay given that I wouldn’t want to re-label cells which I have already labeled, I could of course extract the predictions and just label the unknown cells manually. Maybe having this as an option would be helpful? (i.e. fix labels of cells with known labels).  
Of course this still doesn’t solve the question why it behaves so weird for some clusters 🤔

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [May 28, 2021, 3:52pm UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/4 "2021-05-28T15:52:51Z")

</div>

> [@KevinMenden](#):
>
> I could of course extract the predictions and just label the unknown cells manually. Maybe having this as an option would be helpful? (i.e. fix labels of cells with known labels).

This should probably be a default. We will make a note of it.

> [@KevinMenden](#):
>
> In the former case, it trains for ~60 epochs SCVI and then ~7 epochs SCANVI. In the latter case just ~60 epochs SCANVI.

This seems like a small number of epochs. How many cells do you have?

---

<div class="post-metadata">

**Author:** ![KevinMenden](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/kevinmenden/32/41_2.png) [@KevinMenden](https://discourse.scverse.org/u/KevinMenden)\
**Post date:** [May 29, 2021, 7:43am UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/5 "2021-05-29T07:43:53Z")

</div>

Ist about 140k cells. I will simply try more epochs then. I’ll let you know if it helped.

---

<div class="post-metadata">

**Author:** ![galenxing](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/galenxing/32/53_2.png) [@galenxing](https://discourse.scverse.org/u/galenxing)\
**Post date:** [May 31, 2021, 8:53pm UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/6 "2021-05-31T20:53:55Z")

</div>

Hey Kevin,

scANVI predicting known celltypes incorrectly is something I’ve also observed – but haven’t extensively tested.

A few more suggestions to potentially improve results:

1. If the frequency of your smallest celltype size is greater than 100, I would set the `n_samples_per_label` arg in `lvae.train()` to that number. (This way you’ll train on more cells each epoch)
2. I agree with Adam in increasing the number of scANVI epochs. I would even train for like 50 epochs since with the `n_samples_per_label` param, you’re subsampling the train set.

Note, you’ll need to install the latest version of scvi-tools off of master. I just fixed a bug in max\_epochs for scANVI. [fix scANVI max\_epochs bug when pretrained by galenxing · Pull Request #1079 · YosefLab/scvi-tools · GitHub](https://github.com/YosefLab/scvi-tools/pull/1079)

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [June 7, 2021, 4:13pm UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/7 "2021-06-07T16:13:11Z")

</div>

@KevinMenden We’d like to further troubleshoot this. Are you able to share your data with us?

---

<div class="post-metadata">

**Author:** ![KevinMenden](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/kevinmenden/32/41_2.png) [@KevinMenden](https://discourse.scverse.org/u/KevinMenden)\
**Post date:** [June 7, 2021, 5:18pm UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/8 "2021-06-07T17:18:50Z")

</div>

Hi both,

sorry for not responding, I was on vacation.

I will try out your ideas. Yes the data are public so I can share them with you. I can basically send you the datasets as processed by me and the scripts I use.

Any preference about how to share the data with you?

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [June 7, 2021, 7:21pm UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/9 "2021-06-07T19:21:45Z")

</div>

I think a Google colab notebook (like our tutorials) that reproduces the issue is easiest, but even sharing the data and script is sufficient (dropbox, google, etc.)

Thanks!

---

<div class="post-metadata">

**Author:** ![KevinMenden](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/kevinmenden/32/41_2.png) [@KevinMenden](https://discourse.scverse.org/u/KevinMenden)\
**Post date:** [June 8, 2021, 7:07am UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/10 "2021-06-08T07:07:43Z")

</div>

Alright, I’ll send you something tomorrow!

---

<div class="post-metadata">

**Author:** ![KevinMenden](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/kevinmenden/32/41_2.png) [@KevinMenden](https://discourse.scverse.org/u/KevinMenden)\
**Post date:** [June 9, 2021, 6:02am UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/11 "2021-06-09T06:02:19Z")

</div>

Okay I’ve uploaded the labeled datasets (in .h5ad format) and the script I used here:

> **[scanvi\_liver\_test](https://onedrive.live.com/redir?resid=11B3DD8A19EBF12E%21225012&authkey=%21AGvoMHYhEcChMYM&e=80I0GE)**
>
> Folder

You should be able to just run the notebook from within that folder. I’ll install the patched scANVI version now and try to set `max_epochs` higher. Didn’t work with the current version.

---

<div class="post-metadata">

**Author:** ![KevinMenden](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/kevinmenden/32/41_2.png) [@KevinMenden](https://discourse.scverse.org/u/KevinMenden)\
**Post date:** [June 9, 2021, 8:24am UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/12 "2021-06-09T08:24:29Z")

</div>

Quick update from my side:

- increasing the scANVI epochs to 50 didn’t really help
- additionally removing the subsampling did help

Without the subsampling and training scANVI for 50 epochs, it looks much better now and basically all cells are labeled correctly. A few labels have changed but those probably make sense.

---

<div class="post-metadata">

**Author:** ![nrclaudio](https://avatars.discourse-cdn.com/v4/letter/n/34f0e0/32.png) [@nrclaudio](https://discourse.scverse.org/u/nrclaudio)\
**Post date:** [March 28, 2023, 9:29am UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/13 "2023-03-28T09:29:50Z")

</div>

Hey,

Having the exact same problem here. I’m wondering what Kevin means by removing the subsampling? Just setting n\_samples\_per\_label to the total amount of least frequent cell type?

Thanks

---

<div class="post-metadata">

**Author:** ![martinkim0](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/martinkim0/32/881_2.png) [@martinkim0](https://discourse.scverse.org/u/martinkim0)\
**Post date:** [April 18, 2023, 8:50pm UTC](https://discourse.scverse.org/t/scanvi-relables-known-cells-with-known-types-incorrectly/98/14 "2023-04-18T20:50:00Z")

</div>

Hi @nrclaudio, sorry for the late reply. This would mean setting `n_samples_per_label=None`, which is the default option.
