# Differences in library sizes between reference and query

**URL:** <https://discourse.scverse.org/t/differences-in-library-sizes-between-reference-and-query/2399>\
**Category:** scvi-tools\
**Created:** [July 19, 2024, 9:04pm UTC](https://discourse.scverse.org/t/differences-in-library-sizes-between-reference-and-query/2399 "2024-07-19T21:04:22Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![emdann](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/emdann/32/254_2.png) [@emdann](https://discourse.scverse.org/u/emdann)\
**Post date:** [July 19, 2024, 9:04pm UTC](https://discourse.scverse.org/t/differences-in-library-sizes-between-reference-and-query/2399/1 "2024-07-19T21:04:22Z")

</div>

Is anyone aware of benchmarks that have tested how large differences in library size (i.e. mean total counts per cell) affect integration with scVI and query mapping with scArches? I am working with an scVI model trained on a reference with significantly lower counts per cell compared to the query (where reference mean total count ~ 1000, query mean total count ~ 5000). After query mapping, I see a relationship between total counts for a cell and similarity to the reference. I’m curious to see if other people have looked at this, before going into doing downsampling experiments, since in my case a bunch of other biological factors are correlated with total counts per cell.

From the user guide I get that the scVI model default is to use sum of counts for library size

> the recent default for scVI is to treat library size as observed, equal to the total RNA UMI count of a cell.

Is training the reference model with `use_observed_lib_size=False` likely to make a difference here?

Thanks a lot!

---

<div class="post-metadata">

**Author:** ![cane11](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/cane11/32/241_2.png) [@cane11](https://discourse.scverse.org/u/cane11)\
**Post date:** [July 20, 2024, 6:59am UTC](https://discourse.scverse.org/t/differences-in-library-sizes-between-reference-and-query/2399/2 "2024-07-20T06:59:16Z")

</div>

Hi, it’s usually not a good idea in my hands. Purely for the latent space, it might be fine but I wouldn’t trust downstream methods like get\_normalized\_counts or differential expression. In general, I would suggest training from scratch and check whether the same structure shows up. It doesn’t necessarily mean that you are interested in this feature and additionally adding total\_counts as continuous covariate key might then help.

---

<div class="post-metadata">

**Author:** ![emdann](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/emdann/32/254_2.png) [@emdann](https://discourse.scverse.org/u/emdann)\
**Post date:** [July 22, 2024, 3:24pm UTC](https://discourse.scverse.org/t/differences-in-library-sizes-between-reference-and-query/2399/3 "2024-07-22T15:24:18Z")

</div>

Hi @cane11, thanks for the tips. By training from scratch do you mean training an scVI model on concatenated reference and query, instead of using query mapping?

---

<div class="post-metadata">

**Author:** ![cane11](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/cane11/32/241_2.png) [@cane11](https://discourse.scverse.org/u/cane11)\
**Post date:** [July 22, 2024, 3:55pm UTC](https://discourse.scverse.org/t/differences-in-library-sizes-between-reference-and-query/2399/4 "2024-07-22T15:55:41Z")

</div>

Hi Emma, yes after concatenation or just use the query data if you are not interested in the reference data.
