# TCR and BCR genes and HVGs selection

**URL:** <https://discourse.scverse.org/t/tcr-and-bcr-genes-and-hvgs-selection/2139>\
**Category:** scanpy\
**Tags:** scvi\
**Created:** [March 5, 2024, 2:00pm UTC](https://discourse.scverse.org/t/tcr-and-bcr-genes-and-hvgs-selection/2139 "2024-03-05T14:00:42Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![pmarzano97](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/pmarzano97/32/554_2.png) [@pmarzano97](https://discourse.scverse.org/u/pmarzano97)\
**Post date:** [March 5, 2024, 2:00pm UTC](https://discourse.scverse.org/t/tcr-and-bcr-genes-and-hvgs-selection/2139/1 "2024-03-05T14:00:42Z")

</div>

Hello world!  
I’ve read in many papers that when performing a re-clustering of some populations, like T cells or B cells, prior to the step of integration and so on, they re-calculate the HVGs but excluding the TCR- or BCR-related genes, because they are donor-specific, especially when talking about BCR.

Can you help me how to remove the TCR- or BCR-related genes before computing the HVGs selection, but without removing them from the .var of the anndata, since I want to evaluate their expression during the step of cell annotation?

The code that I use to calculate the HVGs is the following:  
sc.pp.highly\_variable\_genes(adata,  
n\_top\_genes = 4000, flavor = “seurat\_v3”,  
layer = “raw”, batch\_key = ‘sample\_id’,  
subset = False)

Thanks a lot!  
Paolo

---

<div class="post-metadata">

**Author:** ![ivirshup](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/ivirshup/32/160_2.png) [@ivirshup](https://discourse.scverse.org/u/ivirshup)\
**Post date:** [March 6, 2024, 2:35pm UTC](https://discourse.scverse.org/t/tcr-and-bcr-genes-and-hvgs-selection/2139/2 "2024-03-06T14:35:09Z")

</div>

I _think_ you could do something like:

```python
sub = adata[:, ~tcr_bcr_genes]
hvg_df = sc.pp.highly_variable_genes(sub, inplace=False)
hvg_df.index = sub.var_names # only necessary in scanpy <1.10

```

Then grab the info out of the returned dataframe.

```python
adata.var["highly_variable"] = hvg_df["highly_variable"].reindex(adata.var_names, fill_value=False)

```

---

<div class="post-metadata">

**Author:** ![pmarzano97](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/pmarzano97/32/554_2.png) [@pmarzano97](https://discourse.scverse.org/u/pmarzano97)\
**Post date:** [April 2, 2024, 7:07am UTC](https://discourse.scverse.org/t/tcr-and-bcr-genes-and-hvgs-selection/2139/3 "2024-04-02T07:07:09Z")

</div>

Thank you so much @ivirshup!!!  
Do you know if there’s a repository or a document or anything else where the BCR/TCR related genes are listed (like for the ones related to the cell cycle in the guide)?

Thank you!!

---

<div class="post-metadata">

**Author:** ![ivirshup](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/ivirshup/32/160_2.png) [@ivirshup](https://discourse.scverse.org/u/ivirshup)\
**Post date:** [April 2, 2024, 8:51am UTC](https://discourse.scverse.org/t/tcr-and-bcr-genes-and-hvgs-selection/2139/4 "2024-04-02T08:51:54Z")

</div>

No problem.

For gene lists I don’t, but maybe @grst does?

---

<div class="post-metadata">

**Author:** ![grst](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/grst/32/17_2.png) [@grst](https://discourse.scverse.org/u/grst)\
**Post date:** [April 2, 2024, 9:41am UTC](https://discourse.scverse.org/t/tcr-and-bcr-genes-and-hvgs-selection/2139/5 "2024-04-02T09:41:42Z")

</div>

The Ensembl gene annotation provides this information in the “Transcript type” column which you can retrieve from Biomart. Other genome annotations should have similar annotations.

More speficially, BCR genes are those with `IG_[VDJDC]_(gene|pseudogene)` and TCR genes those with `TR_[VDJDC]_(gene|pseudogene)` transcript type.

---

<div class="post-metadata">

**Author:** ![cane11](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/cane11/32/241_2.png) [@cane11](https://discourse.scverse.org/u/cane11)\
**Post date:** [April 3, 2024, 1:02am UTC](https://discourse.scverse.org/t/tcr-and-bcr-genes-and-hvgs-selection/2139/6 "2024-04-03T01:02:07Z")

</div>

As I used it yesterday:

```auto
dataset = pybiomart.Dataset(name='hsapiens_gene_ensembl', host='http://www.ensembl.org')
table = dataset.query(attributes=['ensembl_gene_id', 'external_gene_name', 'chromosome_name', 'transcript_biotype'], use_attr_names=True)
merged = pd.merge(
    adata.var,
    table,
    left_on='gene_ids',
    right_on='ensembl_gene_id'
)

merged['transcript_biotype'].value_counts()
adata.var['gene_name'] = adata.var_names
gene_subset = pd.merge(
    adata.var,
    table[table['transcript_biotype'].isin(['protein_coding', 'lncRNA', 'IG_C_gene', 'TR_C_gene'])], # replace here with your filter.
    left_on='gene_ids',
    right_on='ensembl_gene_id',
)
gene_subset.index = gene_subset['gene_name']
gene_subset.index.value_counts()

```
