# Build a large anndata object column by colum

**URL:** <https://discourse.scverse.org/t/build-a-large-anndata-object-column-by-colum/785>\
**Category:** anndata\
**Created:** [September 29, 2022, 8:32am UTC](https://discourse.scverse.org/t/build-a-large-anndata-object-column-by-colum/785 "2022-09-29T08:32:14Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![SilasK](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/silask/32/390_2.png) [@SilasK](https://discourse.scverse.org/u/SilasK)\
**Post date:** [September 29, 2022, 8:32am UTC](https://discourse.scverse.org/t/build-a-large-anndata-object-column-by-colum/785/1 "2022-09-29T08:32:14Z")

</div>

Hello, I have mappin results for \>1K samples and \>1B genes. when I load a pandas data frame into memory it is about 1GB per sample. Hence to load all samples into memory will not be possible.

Is there a way to build a sparse matrix sample by sample (column by column) ideally backed to the disk?

Will I be able to apply algorithms on the combined data like total sum normalization. Or anyway do I need to find another distributed way to handle this data?

---

<div class="post-metadata">

**Author:** ![ivirshup](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.scverse.org/ivirshup/32/160_2.png) [@ivirshup](https://discourse.scverse.org/u/ivirshup)\
**Post date:** [September 29, 2022, 11:09am UTC](https://discourse.scverse.org/t/build-a-large-anndata-object-column-by-colum/785/2 "2022-09-29T11:09:08Z")

</div>

Hey @SilasK,

Yeah, that should be possible either in memory or on disk. I would note that with scverse (and the python ecosystem more generally) samples are rows instead of columns.

I don’t know what format your data is in right now, but the code would look roughly like:

```python
import numpy as np
from scipy import sparse

samples = ...

indptr = np.zeros(N_SAMPLES + 1, dtype=np.int64)
indices_list = []
data_list = []

for i, sample in enumerate(samples):
    row_dense = read_sample(sample)
    row_indices = np.nonzero(row_dense)
    row_data = row_dense[row_indices]
    indptr[i + 1] = indptr[i] + len(row_indices)
    indices_list.append(row_indices)
    data_list.append(row_data)

X = sparse.csr_matrix(
    (
        np.concatenate(data_list),
        np.concatenate(indices_list),
        indptr
    ),
    shape=(len(samples), N_VARIABLES)
)

```
