# Batch key and categorical variables for get\_normalized\_expression()

**URL:** <https://discourse.scverse.org/t/batch-key-and-categorical-variables-for-get-normalized-expression/3841>\
**Category:** scvi-tools\
**Created:** [October 29, 2025, 2:03pm UTC](https://discourse.scverse.org/t/batch-key-and-categorical-variables-for-get-normalized-expression/3841 "2025-10-29T14:03:40Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![munta](https://avatars.discourse-cdn.com/v4/letter/m/e19b73/32.png) [@munta](https://discourse.scverse.org/u/munta)\
**Post date:** [October 29, 2025, 2:03pm UTC](https://discourse.scverse.org/t/batch-key-and-categorical-variables-for-get-normalized-expression/3841/1 "2025-10-29T14:03:41Z")

</div>

Hi.

many thanks for the nice tool.  
I have combined multiple public scRNA-seq data and used scvi-tool for integration.

```auto
scvi.model.SCVI.setup_anndata(adata, layer="counts", batch_key="Project",
                              categorical_covariate_keys=['Patient','Sex'])
vae = scvi.model.SCVI(adata, n_layers=2, n_latent=30, gene_likelihood="nb")
vae.train(accelerator='gpu')
vae.save('../data/scvi_models/scvi_integration_model_project_sample_sex')

```

Then, I clustered the data based on scvi latent space

```auto
adata.obsm["X_scVI"] = vae.get_latent_representation()
sc.pp.neighbors(adata, use_rep="X_scVI")
sc.tl.umap(adata)
sc.tl.leiden(adata,resolution=0.3)

```

Now, my aim is to go with downstream analysis such as defining markers, DEGs (using different techniques like MAST in R), infercnvpy, (liana+) cell-cell communication analysis, pertpy (MiloR, Augur, scCODA) etc..  
Which normalized counts should I use for that? And how to determine the categorical covariates key if I use `get_normalized_expressionz()`.

best,

---

<div class="post-metadata">

**Author:** ![cane11](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/cane11/32/241_2.png) [@cane11](https://discourse.scverse.org/u/cane11)\
**Post date:** [October 29, 2025, 7:27pm UTC](https://discourse.scverse.org/t/batch-key-and-categorical-variables-for-get-normalized-expression/3841/2 "2025-10-29T19:27:51Z")

</div>

Hi. For none of the mentioned tools you would want to use normalized expression.

---

<div class="post-metadata">

**Author:** ![munta](https://avatars.discourse-cdn.com/v4/letter/m/e19b73/32.png) [@munta](https://discourse.scverse.org/u/munta)\
**Post date:** [November 4, 2025, 11:01am UTC](https://discourse.scverse.org/t/batch-key-and-categorical-variables-for-get-normalized-expression/3841/4 "2025-11-04T11:01:55Z")

</div>

Thank you for your answer. But, MAST, cellchat, nichenet they all need normalized counts …  
Anyway, my question is, in order to get the batch-corrected counts, I should set the transform\_batch parameter in vae.get\_normalized\_expression. Is it ok to use the largest batch I have in the data? Otherwise, I am not sure how to choose the “preferable” batch! thanks in advance
