# Recommendation for transform\_batch / categorical\_covariate\_keys to obtain "batch corrected" counts

**URL:** <https://discourse.scverse.org/t/recommendation-for-transform-batch-categorical-covariate-keys-to-obtain-batch-corrected-counts/2346>\
**Category:** scvi-tools\
**Tags:** integration\
**Created:** [June 22, 2024, 3:24pm UTC](https://discourse.scverse.org/t/recommendation-for-transform-batch-categorical-covariate-keys-to-obtain-batch-corrected-counts/2346 "2024-06-22T15:24:34Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![wgao688](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/wgao688/32/1087_2.png) [@wgao688](https://discourse.scverse.org/u/wgao688)\
**Post date:** [June 22, 2024, 3:24pm UTC](https://discourse.scverse.org/t/recommendation-for-transform-batch-categorical-covariate-keys-to-obtain-batch-corrected-counts/2346/1 "2024-06-22T15:24:34Z")

</div>

Hi,

Thanks for making this great package! I’m working with a multi-batch dataset with several donors, generated from multiple studies with their own batch effects (technology and varying sequencing depth). I am interested in generating a “batch-corrected” count matrix for downstream analysis. I see that in `scvi.model.SCVI.setup_anndata()` there are options for `categorical_covariate_keys` and `continuous_covariate_keys`. So then I could use the various batch effects like “technology” and “donor” as categorical covariates, for example.

I also see that once I build the model, there is a function model.get\_normalized\_expression, which can take a `transform_batch` argument, but how should I combine this with the categorical\_covariates\_keys above?

I saw a similar question here but didn’t see a specific recommendation: [link](https://github.com/scverse/scvi-tools/issues/1613).

Additionally, I see that the transform\_batch requires one to specify a specific batch to treat each sample as it if came from as noted here: [link](https://github.com/scverse/scvi-tools/issues/786). In that case, would it make more sense to average over all larger batch effects such as “technology” but not the individual to individual batches like “donor”? Would that be reasonable?

Overall, I want to remove technical artifacts such as sequencing depth and cell vs. nuclei effects from my datasets, without removing biological states such as tissue location or sex. Thanks!

---

<div class="post-metadata">

**Author:** ![cane11](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/cane11/32/241_2.png) [@cane11](https://discourse.scverse.org/u/cane11)\
**Post date:** [July 3, 2024, 10:45pm UTC](https://discourse.scverse.org/t/recommendation-for-transform-batch-categorical-covariate-keys-to-obtain-batch-corrected-counts/2346/2 "2024-07-03T22:45:02Z")

</div>

Hi, to use transform\_batch you have to use batch\_key instead of categorical\_covariate\_key. You can provide a list of batches and the output is the average over these batches (so.e.g. All samples for a specific technology).
