# Merging identical genes from 10x fixed scRNA

**URL:** <https://discourse.scverse.org/t/merging-identical-genes-from-10x-fixed-scrna/2142>\
**Category:** anndata\
**Created:** [March 6, 2024, 5:13pm UTC](https://discourse.scverse.org/t/merging-identical-genes-from-10x-fixed-scrna/2142 "2024-03-06T17:13:49Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![pakiessling](https://avatars.discourse-cdn.com/v4/letter/p/73ab20/32.png) [@pakiessling](https://discourse.scverse.org/u/pakiessling)\
**Post date:** [March 6, 2024, 5:13pm UTC](https://discourse.scverse.org/t/merging-identical-genes-from-10x-fixed-scrna/2142/1 "2024-03-06T17:13:50Z")

</div>

Hi,

I have a perhaps unusual use-case.  
I am working with the ouput of cellranger multi for the new probe-based fixed single cell kit.  
The unfiltered `raw_feature_bc_matrix.h5` which I want to utilize with Cellbender and Co contains probes and not transcript species.

That means when I load it in with `sc.read_10x_h5()` there will be duplicate entries in adata.var and in the columns of adata.X as some genes have multiple probes targeting them. These entries have identical var\_names

What would be the most graceful way to merge these entries?

---

<div class="post-metadata">

**Author:** ![ivirshup](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/ivirshup/32/160_2.png) [@ivirshup](https://discourse.scverse.org/u/ivirshup)\
**Post date:** [March 6, 2024, 5:55pm UTC](https://discourse.scverse.org/t/merging-identical-genes-from-10x-fixed-scrna/2142/2 "2024-03-06T17:55:11Z")

</div>

Do you have both the same targets and the same expression levels?

Otherwise, maybe you’d want the “better” probe? I don’t know how you’d decide that though.

---

<div class="post-metadata">

**Author:** ![pakiessling](https://avatars.discourse-cdn.com/v4/letter/p/73ab20/32.png) [@pakiessling](https://discourse.scverse.org/u/pakiessling)\
**Post date:** [March 6, 2024, 6:17pm UTC](https://discourse.scverse.org/t/merging-identical-genes-from-10x-fixed-scrna/2142/3 "2024-03-06T18:17:34Z")

</div>

Hi Isaac, the expression levels of each probe is different.  
They seem to be targeting different forms of the genes.

However, I see now that 10x simply removes blacklisted probes which results in unique .var in the end.

![grafik](https://canada1.discourse-cdn.com/flex035/uploads/forum11/original/1X/f66be39ae3f124b3e1d598e8d9e838625c485a74.png)

So not relevant after all.

Out of curiosity how would one merge genes? Extract the index number and then operate on anndata.X?

---

<div class="post-metadata">

**Author:** ![ivirshup](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/ivirshup/32/160_2.png) [@ivirshup](https://discourse.scverse.org/u/ivirshup)\
**Post date:** [March 6, 2024, 6:24pm UTC](https://discourse.scverse.org/t/merging-identical-genes-from-10x-fixed-scrna/2142/4 "2024-03-06T18:24:21Z")

</div>

When I’ve done this with microarray data in the past was something like:

- DataFrame where each row is a gene
- Group by probe target
- Some aggregation (`max`, `mean`, etc)

If you are okay with densify the matrix, this should be straight forward. Maybe with `flox` or `numpy-groupies`.

For sparse, it’s a little more complicated. But this would be a good extension of the new `sc.get.aggregate` and I’ve opened an issue to track it:

> <https://github.com/scverse/scanpy/issues/2898>
>
> \### What kind of feature would you like to request?
> 
> Additional function param…eters / changed functionality / changed defaults?
> 
> \### Please describe your wishes
> 
> @Intron7, found a use case 😆
> 
> It could be nice for \`sc.get.aggregate\` to be able to return sparse matrices where we don't expect the aggregation to return very dense data.
> 
> Previously discussed in:
> 
> \* https://github.com/scverse/scanpy/issues/2892
> 
> Usecases include:
> 
> \* Taking the \`max\` for multiple reports of a genes (\`sc.get.aggregate(adata, "probe\_target", "max")\`, e.g. https://discourse.scverse.org/t/merging-identical-genes-from-10x-fixed-scrna/2142)
> \* \*(note: max is not currently implemented)\*
> \* Small aggregations, e.g. only summing neighbors
> 
> This would require both api design choices for what the argument is called, and efficient implementations for both dense and sparse results (\`python-graphblas\` could be useful here)
