# Sc.tl.ingest over representing rare cell types in spatial transcriptomics data

**URL:** <https://discourse.scverse.org/t/sc-tl-ingest-over-representing-rare-cell-types-in-spatial-transcriptomics-data/3868>\
**Category:** scanpy\
**Created:** [November 18, 2025, 8:46am UTC](https://discourse.scverse.org/t/sc-tl-ingest-over-representing-rare-cell-types-in-spatial-transcriptomics-data/3868 "2025-11-18T08:46:43Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![spatts14](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/spatts14/32/1034_2.png) [@spatts14](https://discourse.scverse.org/u/spatts14)\
**Post date:** [November 18, 2025, 8:46am UTC](https://discourse.scverse.org/t/sc-tl-ingest-over-representing-rare-cell-types-in-spatial-transcriptomics-data/3868/1 "2025-11-18T08:46:43Z")

</div>

I am trying to use sc.tl.ingest to predict cell types for my Xenium 5k spatial transcriptomics data.

I have subsetted the reference dataset to only contain the disease types and tissue types that are present in my STx data. However, when I use sc.tl.ingest, I am getting an over representation of rare cell types so I know it is not working correctly. I realize this is likely because my STx has a reduced number of genes compared to scRNAseq, less transcripts per cell, and fewer genes per cell. However, is there way to improve the prediction accuracy? Any other suggestions?

Thanks!
