# How to retain gene only present/mapped in one of the dataset before integration?

**URL:** <https://discourse.scverse.org/t/how-to-retain-gene-only-present-mapped-in-one-of-the-dataset-before-integration/1653>\
**Category:** scanpy\
**Created:** [August 1, 2023, 2:48pm UTC](https://discourse.scverse.org/t/how-to-retain-gene-only-present-mapped-in-one-of-the-dataset-before-integration/1653 "2023-08-01T14:48:58Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![dub2s](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/dub2s/32/361_2.png) [@dub2s](https://discourse.scverse.org/u/dub2s)\
**Post date:** [August 1, 2023, 2:48pm UTC](https://discourse.scverse.org/t/how-to-retain-gene-only-present-mapped-in-one-of-the-dataset-before-integration/1653/1 "2023-08-01T14:48:58Z")

</div>

Hii all

I am integrating 4-5 datasets, all belonging to mouse. One of the datasets has a gene, “eGFP”, which is not present in other datasets. By the term “not present”, I mean to say that only one of the datasets was aligned with genome containing eGFP as a gene while other datasets have been aligned with a different genome (which doesn’t contain eGFP even though it has all the other genes from mus musculus assembly).

Is there a way I could retain this gene while integrating? Intuitively, I would like to just define a new gene “eGFP” in other datasets and keep its count value to zero for all the cells. However, I am not sure how to define a new gene ?
