# Input to scvi encoder

**URL:** <https://discourse.scverse.org/t/input-to-scvi-encoder/233>\
**Category:** scvi-tools\
**Tags:** scvi\
**Created:** [December 6, 2021, 6:35pm UTC](https://discourse.scverse.org/t/input-to-scvi-encoder/233 "2021-12-06T18:35:13Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![willtownes](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/willtownes/32/114_2.png) [@willtownes](https://discourse.scverse.org/u/willtownes)\
**Post date:** [December 6, 2021, 6:35pm UTC](https://discourse.scverse.org/t/input-to-scvi-encoder/233/1 "2021-12-06T18:35:13Z")

</div>

I have a vague memory that in the original version of scvi, the counts were normalized somehow before being passed into the encoder, perhaps as log(1+CPM). My recollection is this was more numerically stable than passing raw counts directly. Is that still the case? I couldn’t find anything in the documentation or in the code itself.

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [December 6, 2021, 10:14pm UTC](https://discourse.scverse.org/t/input-to-scvi-encoder/233/2 "2021-12-06T22:14:24Z")

</div>

Just a log(1+x) transform. It’s here:

> <https://github.com/YosefLab/scvi-tools/blob/c01cb292cbb4f2974bd4f6403df3665fd9fac800/scvi/module/_vae.py#L256-L266>

And yes, more numerically stable!

---

<div class="post-metadata">

**Author:** ![willtownes](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/willtownes/32/114_2.png) [@willtownes](https://discourse.scverse.org/u/willtownes)\
**Post date:** [December 7, 2021, 8:08pm UTC](https://discourse.scverse.org/t/input-to-scvi-encoder/233/3 "2021-12-07T20:08:18Z")

</div>

Excellent, thanks a ton Adam!

---

<div class="post-metadata">

**Author:** ![yxy92](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/yxy92/32/851_2.png) [@yxy92](https://discourse.scverse.org/u/yxy92)\
**Post date:** [November 3, 2023, 2:18pm UTC](https://discourse.scverse.org/t/input-to-scvi-encoder/233/4 "2023-11-03T14:18:05Z")

</div>

Hi Adam, I can see that log(1+X) keep partial info of the raw data while increasing the numerical stability as you mentioned. Is there a specific reason that scVI does not use more common log2(1+CPM) or log2(1+TPM)?

---

<div class="post-metadata">

**Author:** ![martinkim0](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/martinkim0/32/881_2.png) [@martinkim0](https://discourse.scverse.org/u/martinkim0)\
**Post date:** [November 14, 2023, 7:31pm UTC](https://discourse.scverse.org/t/input-to-scvi-encoder/233/5 "2023-11-14T19:31:20Z")

</div>

Hi, we just use `log(1 + x)` for simplicity as we generally only care about numerical stability of the model.
