# scVI imputation confusion

**URL:** <https://discourse.scverse.org/t/scvi-imputation-confusion/111>\
**Category:** scvi-tools\
**Tags:** scvi, imputation\
**Created:** [June 16, 2021, 12:07pm UTC](https://discourse.scverse.org/t/scvi-imputation-confusion/111 "2021-06-16T12:07:46Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![gdewael](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/gdewael/32/64_2.png) [@gdewael](https://discourse.scverse.org/u/gdewael)\
**Post date:** [June 16, 2021, 12:07pm UTC](https://discourse.scverse.org/t/scvi-imputation-confusion/111/1 "2021-06-16T12:07:47Z")

</div>

Hi scvi-tools team,

Great piece of software.  
I’m looking to benchmark some models/dataset w.r.t. imputation performance.

In your documentation, it is not immediately clear how to properly impute gene expression values using scVI.  
From the scVI paper:

> This mapping goes through intermediate values `ρ^n_g`, which provide a batch-corrected, normalized estimate of the percentage of transcripts in each cell `n` that originate from each gene `g` . We used these estimates for differential expression analysis and its scaled version (multiplying `ρ^n_g` by the estimated library size `ℓ_n`) for imputation.

I have surmised that `ρ^n` and `ℓ_n` can be obtained through the functions `get_normalized_expression` and `get_latent_representation`

My question is in regards to the `library_size` argument of the former function. In your [user guide](https://docs.scvi-tools.org/en/stable/user_guide/notebooks/api_overview.html), you use a common library size. Hence my question: to benchmark imputation performance, should expression frequencing be scaled to latent library sizes or a common library size?

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [June 16, 2021, 4:56pm UTC](https://discourse.scverse.org/t/scvi-imputation-confusion/111/2 "2021-06-16T16:56:46Z")

</div>

> [@gdewael](#):
>
> I have surmised that `ρ^n` and `ℓ_n` can be obtained through the functions `get_normalized_expression` and `get_latent_representation`

So with this function you get \ell\_n\rho\_n, where you can use the parameters to set \ell\_n=1 (default I believe).

> [@gdewael](#):
>
> My question is in regards to the `library_size` argument of the former function. In your [user guide](https://docs.scvi-tools.org/en/stable/user_guide/notebooks/api_overview.html), you use a common library size. Hence my question: to benchmark imputation performance, should expression frequencing be scaled to latent library sizes or a common library size?

It’s not easy to give a straightforward answer to this as it depends on the setup of your benchmarking experiment. Most people would tend to look at the data on a common library scale, I assume.
