# Get\_normalized\_expression function arguments

**URL:** <https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140>\
**Category:** scvi-tools\
**Tags:** totalvi\
**Created:** [July 8, 2021, 10:59pm UTC](https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140 "2021-07-08T22:59:08Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![andrewjkwok](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/andrewjkwok/32/79_2.png) [@andrewjkwok](https://discourse.scverse.org/u/andrewjkwok)\
**Post date:** [July 8, 2021, 10:59pm UTC](https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140/1 "2021-07-08T22:59:08Z")

</div>

Hi,

I had a few questions about the `get_normalized_expression` function that I hoped to get some clarification for. From the totalVI tutorial ([https://docs.scvi-tools.org/en/stable/user\_guide/notebooks/totalVI.html](https://docs.scvi-tools.org/en/stable/user_guide/notebooks/totalVI.html)), `n_samples` is set to 25 and `transform_batch` is given the list of both datasets.

First, does `n_samples` refer to the number of cells (surely it can’t be number of biological samples, as the example dataset only has 2 individuals?), and if yes, why is the default 1 / why is the suggested number in the tutorial 25 / what might be a recommended number to set this as?

Second, how exactly should the `transform_batch` argument be used? I understand from the documentation and github ([how to get corrected expression matrix after batch removal · Issue #786 · YosefLab/scvi-tools · GitHub](https://github.com/YosefLab/scvi-tools/issues/786)) that it is about which batch to condition over. Intuitively, it seems to be that it would make the most sense to condition over all the batches as is also done in the tutorial, but would there be any situation where that might not be recommended?

Many thanks in advance.

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [July 9, 2021, 4:30am UTC](https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140/2 "2021-07-09T04:30:58Z")

</div>

> [@andrewjkwok](#):
>
> First, does `n_samples` refer to the number of cells (surely it can’t be number of biological samples, as the example dataset only has 2 individuals?), and if yes, why is the default 1 / why is the suggested number in the tutorial 25 / what might be a recommended number to set this as?

`n_samples` refers to Monte Carlo sampling for each cell. The normalized expression is a random variable, and we return the average over 25 samples in this case. It’s an unbiased estimate of the expectation, but you need a LOT more samples to reduce the variance of this estimate. So empirically, 25 just seemed to work well.

> [@andrewjkwok](#):
>
> Intuitively, it seems to be that it would make the most sense to condition over all the batches as is also done in the tutorial, but would there be any situation where that might not be recommended?

Generally, you would take all the batches, but you could have the case where one cell type is only seen in one batch. In this case, you’d want to call the function separately for that cell type, and not use the transform batch param in that case.

---

<div class="post-metadata">

**Author:** ![andrewjkwok](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/andrewjkwok/32/79_2.png) [@andrewjkwok](https://discourse.scverse.org/u/andrewjkwok)\
**Post date:** [July 9, 2021, 9:27am UTC](https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140/3 "2021-07-09T09:27:43Z")

</div>

That’s super useful to know - thank you!

---

<div class="post-metadata">

**Author:** ![andrewjkwok](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/andrewjkwok/32/79_2.png) [@andrewjkwok](https://discourse.scverse.org/u/andrewjkwok)\
**Post date:** [August 28, 2021, 12:24pm UTC](https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140/4 "2021-08-28T12:24:23Z")

</div>

Hello - just wanted to briefly follow up on this function. Is it possible to only get the denoised/normalised protein expression data matrix, and not the RNA one, or vice versa?

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [August 30, 2021, 5:36pm UTC](https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140/5 "2021-08-30T17:36:58Z")

</div>

The method returns both, you can ignore the RNA denoised expression. It wouldn’t save any time really to reimplement in a way that only returns protein, so feel free to just ignore the RNA part for now.

---

<div class="post-metadata">

**Author:** ![andrewjkwok](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/andrewjkwok/32/79_2.png) [@andrewjkwok](https://discourse.scverse.org/u/andrewjkwok)\
**Post date:** [September 1, 2021, 2:23pm UTC](https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140/6 "2021-09-01T14:23:35Z")

</div>

I see. The problem I’m running into is actually that I keep running out of memory, and was wondering whether returning only the protein expression might reduce the memory? If not I suppose then I don’t have options other than increase memory or forgo the denoised expression values?

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [September 1, 2021, 2:58pm UTC](https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140/7 "2021-09-01T14:58:16Z")

</div>

You can do two things:

1. Reduce `batch_size`
2. Pass an argument to `gene_list` (e.g., a list of two genes, so only two genes are used)

Both of these steps will save you memory.

---

<div class="post-metadata">

**Author:** ![andrewjkwok](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/andrewjkwok/32/79_2.png) [@andrewjkwok](https://discourse.scverse.org/u/andrewjkwok)\
**Post date:** [September 4, 2021, 10:23pm UTC](https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140/8 "2021-09-04T22:23:51Z")

</div>

Ah, number 2 makes a lot of sense - thank you!

---

<div class="post-metadata">

**Author:** ![andrewjkwok](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/andrewjkwok/32/79_2.png) [@andrewjkwok](https://discourse.scverse.org/u/andrewjkwok)\
**Post date:** [September 9, 2021, 12:25pm UTC](https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140/9 "2021-09-09T12:25:22Z")

</div>

I tried the strategy of passing an argument to `gene_list`, but get an error:

`ValueError: Value passed for key 'denoised_rna' is of incorrect shape. Values of layers must match dimensions (0, 1) of parent. Value had shape (92009, 3) while it should have had (92009, 4000).`

Do you have any idea how this could be fixed?

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [September 11, 2021, 2:46am UTC](https://discourse.scverse.org/t/get-normalized-expression-function-arguments/140/10 "2021-09-11T02:46:03Z")

</div>

Probably a bug, would you be able to make an issue on GitHub? Thanks!
