# Sparse matrix error using totalVI integration

**URL:** <https://discourse.scverse.org/t/sparse-matrix-error-using-totalvi-integration/1663>\
**Category:** scvi-tools\
**Tags:** integration\
**Created:** [August 4, 2023, 8:41am UTC](https://discourse.scverse.org/t/sparse-matrix-error-using-totalvi-integration/1663 "2023-08-04T08:41:21Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![pouria](https://avatars.discourse-cdn.com/v4/letter/p/e56c9b/32.png) [@pouria](https://discourse.scverse.org/u/pouria)\
**Post date:** [August 4, 2023, 8:41am UTC](https://discourse.scverse.org/t/sparse-matrix-error-using-totalvi-integration/1663/1 "2023-08-04T08:41:21Z")

</div>

Hi all,  
I have a CITE-seq dataset from 5 different donor that I’m trying to integrate using muon and totalVI. I started by concatenating the ADT and RNA data for each donor separately using the concatenate function in scanpy. After that i created a muon object using mu.MuData({“rna”: RNAdata, “protein”: Protdata}).

When I start the training of my model I directly get the following error message:

" INFO Computing empirical prior initialization for protein background.

--------------------------------------------------------------------------- TypeError Traceback (most recent call last) Cell In[208], line 1 ----\> 1 vae = scvi.model.TOTALVI(posdata1) File [~/anaconda3/envs/Scanpy2/lib/python3.11/site-packages/scvi/model/\_totalvi.py:142](https://file+.vscode-resource.vscode-cdn.net/Users/poumom/Library/CloudStorage/OneDrive-KarolinskaInstitutet/Mac/Documents/ADT%20test%20stuff/~/anaconda3/envs/Scanpy2/lib/python3.11/site-packages/scvi/model/_totalvi.py:142), in TOTALVI. **init** (self, adata, n\_latent, gene\_dispersion, protein\_dispersion, gene\_likelihood, latent\_distribution, empirical\_protein\_background\_prior, override\_missing\_proteins, \*\*model\_kwargs) 136 emp\_prior = ( 137 empirical\_protein\_background\_prior 138 if empirical\_protein\_background\_prior is not None 139 else (self.summary\_stats.n\_proteins \> 10) 140 ) 141 if emp\_prior: → 142 prior\_mean, prior\_scale = self.\_get\_totalvi\_protein\_priors(adata) 143 else: 144 prior\_mean, prior\_scale = None, None File [~/anaconda3/envs/Scanpy2/lib/python3.11/site-packages/scvi/model/\_totalvi.py:1163](https://file+.vscode-resource.vscode-cdn.net/Users/poumom/Library/CloudStorage/OneDrive-KarolinskaInstitutet/Mac/Documents/ADT%20test%20stuff/~/anaconda3/envs/Scanpy2/lib/python3.11/site-packages/scvi/model/_totalvi.py:1163), in TOTALVI.\_get\_totalvi\_protein\_priors(self, adata, n\_cells) 1161 for c in batch\_pro\_exp: 1162 try: → 1163 gmm.fit(np.log1p(c.reshape(-1, 1))) 1164 # when cell is all 0 1165 except ConvergenceWarning: File [~/anaconda3/envs/Scanpy2/lib/python3.11/site-packages/sklearn/mixture/\_base.py:181](https://file+.vscode-resource.vscode-cdn.net/Users/poumom/Library/CloudStorage/OneDrive-KarolinskaInstitutet/Mac/Documents/ADT%20test%20stuff/~/anaconda3/envs/Scanpy2/lib/python3.11/site-packages/sklearn/mixture/_base.py:181), in BaseMixture.fit(self, X, y) 155 “”"Estimate model parameters with the EM algorithm.

…

538 ) 539 elif isinstance(accept\_sparse, (list, tuple)): 540 if len(accept\_sparse) == 0: TypeError: A sparse matrix was passed, but dense data is required. Use X.toarray() to convert to a dense numpy array."

Can any of you make any sense of this? I’ve tried converting my matrixes to dense but i still get the same error message.

Very thankful for any response

---

<div class="post-metadata">

**Author:** ![pouria](https://avatars.discourse-cdn.com/v4/letter/p/e56c9b/32.png) [@pouria](https://discourse.scverse.org/u/pouria)\
**Post date:** [August 4, 2023, 8:50am UTC](https://discourse.scverse.org/t/sparse-matrix-error-using-totalvi-integration/1663/2 "2023-08-04T08:50:24Z")

</div>

I upload the ADT data in this way:

```auto
    # Load the ADT data
    adt_counts_sparse = mmread(adt_counts_file).tocsr()

    adt_barcodes = pd.read_csv(adt_barcodes_file, compression='gzip', sep='\t', header=None).values.flatten()
    adt_barcodes = [barcode + "-1" for barcode in adt_barcodes]
    adt_tags = pd.read_csv(adt_tags_file, compression='gzip', header=None).values.flatten()

    # Make sure the counts, barcodes, and tags align
    if adt_counts_sparse.shape != (len(adt_barcodes), len(adt_tags)):
        adt_counts_sparse = adt_counts_sparse.transpose()

    # Create the ADT AnnData object
    globals()[donor_name + '_adt'] = sc.AnnData(X=adt_counts_sparse, 
                                                obs=pd.DataFrame(index=adt_barcodes), 
                                                var=pd.DataFrame(index=adt_tags))

    # Change the matrix format in the uploaded ADT
    globals()[donor_name + '_adt'].X = csr_matrix(globals()[donor_name + '_adt'].X)
    
    return globals()[donor_name + '_adt']

```
