# Error in highly variable gene selection

**URL:** <https://discourse.scverse.org/t/error-in-highly-variable-gene-selection/276>\
**Category:** scanpy\
**Tags:** scrna-seq, gene-selection\
**Created:** [February 13, 2022, 10:43am UTC](https://discourse.scverse.org/t/error-in-highly-variable-gene-selection/276 "2022-02-13T10:43:49Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![win](https://avatars.discourse-cdn.com/v4/letter/w/65b543/32.png) [@win](https://discourse.scverse.org/u/win)\
**Post date:** [February 13, 2022, 10:43am UTC](https://discourse.scverse.org/t/error-in-highly-variable-gene-selection/276/1 "2022-02-13T10:43:49Z")

</div>

Hi

I am trying to use SCVI - tools for batch correction. I followed as in tutorial. I got the error in the step below.

```auto
sc.pp.highly_variable_genes(
    adata,
    n_top_genes=2000,
    subset=True,
    layer="counts",
    flavor="seurat_v3",
    batch_key='patient_id'
)

```

The error I got is

```auto
If you pass `n_top_genes`, all cutoffs are ignored.
extracting highly variable genes
---------------------------------------------------------------------------
ValueError Traceback (most recent call last)
<ipython-input-11-e00c5f508a6a> in <module>
----> 1 sc.pp.highly_variable_genes(
      2 adata,
      3 n_top_genes=2000,
      4 subset=True,
      5 layer="counts",

/opt/conda/lib/python3.8/site-packages/scanpy/preprocessing/_highly_variable_genes.py in highly_variable_genes(adata, layer, n_top_genes, min_disp, max_disp, min_mean, max_mean, span, n_bins, flavor, subset, inplace, batch_key)
    413

 
    414 if flavor == 'seurat_v3':
--> 415 return _highly_variable_genes_seurat_v3(
    416 adata,
    417 layer=layer,

/opt/conda/lib/python3.8/site-packages/scanpy/preprocessing/_highly_variable_genes.py in _highly_variable_genes_seurat_v3(adata, layer, n_top_genes, batch_key, span, subset, inplace)
     82 x = np.log10(mean[not_const])
     83 model = loess(x, y, span=span, degree=2)
---> 84 model.fit()
     85 estimat_var[not_const] = model.outputs.fitted_values
     86 reg_std = np.sqrt(10 ** estimat_var)

_loess.pyx in _loess.loess.fit()

ValueError: b'reciprocal condition number 1.0638e-15\n'

```

I am not sure what the issue and what cause this error. Could you please advise?  
Thanks

---

<div class="post-metadata">

**Author:** ![valehvpa](https://avatars.discourse-cdn.com/v4/letter/v/8c91f0/32.png) [@valehvpa](https://discourse.scverse.org/u/valehvpa)\
**Post date:** [February 13, 2022, 5:32pm UTC](https://discourse.scverse.org/t/error-in-highly-variable-gene-selection/276/2 "2022-02-13T17:32:06Z")

</div>

The exception stems from the scanpy module. Could you open an issue in [scanpy](https://github.com/theislab/scanpy/issues)?

---

<div class="post-metadata">

**Author:** ![win](https://avatars.discourse-cdn.com/v4/letter/w/65b543/32.png) [@win](https://discourse.scverse.org/u/win)\
**Post date:** [February 13, 2022, 7:09pm UTC](https://discourse.scverse.org/t/error-in-highly-variable-gene-selection/276/3 "2022-02-13T19:09:19Z")

</div>

ok, I have posted on scanpy discussion. Hope someone can help me there. Thanks

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [February 14, 2022, 9:09pm UTC](https://discourse.scverse.org/t/error-in-highly-variable-gene-selection/276/4 "2022-02-14T21:09:42Z")

</div>

Also this is usually due to having many many lowly expressed (almost 0 genes). It will likely work if you run `filter_genes` first

---

<div class="post-metadata">

**Author:** ![win](https://avatars.discourse-cdn.com/v4/letter/w/65b543/32.png) [@win](https://discourse.scverse.org/u/win)\
**Post date:** [February 15, 2022, 9:22am UTC](https://discourse.scverse.org/t/error-in-highly-variable-gene-selection/276/5 "2022-02-15T09:22:12Z")

</div>

Hi @adamgayoso thanks so much for advice. I tried `sc.pp.filter_genes(adata, min_counts=3)` but it still shows the error. The strange thing is when I remove the `batch_key='patient_id'` argument in `sc.pp.highly_variable_genes`, it works. I am a bit confused. Any suggestion?

thanks

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [February 15, 2022, 4:49pm UTC](https://discourse.scverse.org/t/error-in-highly-variable-gene-selection/276/6 "2022-02-15T16:49:30Z")

</div>

It could be that within one batch category you have many all zero genes, if I had to guess.

---

<div class="post-metadata">

**Author:** ![win](https://avatars.discourse-cdn.com/v4/letter/w/65b543/32.png) [@win](https://discourse.scverse.org/u/win)\
**Post date:** [February 15, 2022, 7:50pm UTC](https://discourse.scverse.org/t/error-in-highly-variable-gene-selection/276/7 "2022-02-15T19:50:23Z")

</div>

Hi @adamgayoso , thanks for advice. May I ask an alternative. Does SCVI really need to run exactly this line

```auto
    adata,
    n_top_genes=1200,
    subset=True,
    layer="counts",
    flavor="seurat_v3",
    batch_key="cell_source"
)

```

Will it cause any effect on scvi downstream analysis if I remove `batch_key = 'cell_source`

Or can I just run the routine scanpy highvar `sc.pl.highly_variable_genes(adata)`

Thanks

---

<div class="post-metadata">

**Author:** ![Valentine\_Svensson](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/valentine_svensson/32/10_2.png) [@Valentine\_Svensson](https://discourse.scverse.org/u/Valentine_Svensson)\
**Post date:** [March 20, 2022, 4:55am UTC](https://discourse.scverse.org/t/error-in-highly-variable-gene-selection/276/8 "2022-03-20T04:55:01Z")

</div>

Hi,

You can select highly variably genes with any procedure. If you are selecting a small number of genes, it is of course important that you are obtaining genes that vary due to the processes you are interested in within your data.

More succinctly: it will all work to do what you propose. Alternative HVG selection methods might give ‘cleaner’ clusters, but you can try other methods if you are disappointed with the results.

/Valentine

---

<div class="post-metadata">

**Author:** ![adamgayoso](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/adamgayoso/32/100_2.png) [@adamgayoso](https://discourse.scverse.org/u/adamgayoso)\
**Post date:** [March 21, 2022, 5:07am UTC](https://discourse.scverse.org/t/error-in-highly-variable-gene-selection/276/9 "2022-03-21T05:07:31Z")

</div>

> [@Valentine\_Svensson](#):
>
> You can select highly variably genes with any procedure.

Exactly! Though important to check what the expected input layer is… e.g., seurat v3 wants counts, the others want log normalized
