All benchmarks

Predict Modality

Release v0.1.0 v0.1.1-rc2rc v0.1.1-rc1rc

Predicting the profiles of one modality (e.g. protein abundance) from another (e.g. mRNA expression).

12 methods
4 control methods
8 datasets
8 metrics
3 releases
Task repository MIT v0.1.1-rc2

Experimental techniques to measure multiple modalities within the same single cell are increasingly becoming available. The demand for these measurements is driven by the promise to provide a deeper insight into the state of a cell. Yet, the modalities are also intrinsically linked. We know that DNA must be accessible (ATAC data) to produce mRNA (expression data), and mRNA in turn is used as a template to produce protein (protein abundance). These processes are regulated often by the same molecules that they produce: for example, a protein may bind DNA to prevent the production of more mRNA. Understanding these regulatory processes would be transformative for synthetic biology and drug target discovery. Any method that can predict a modality from another must have accounted for these regulatory processes, but the demand for multi-modal data shows that this is not trivial.

Contributors

  • Alejandro Granados
    author
  • Alex Tong
    author
  • Bastian Rieck
    author
  • Christopher Lance
    author
  • Daniel Burkhardt
    author
  • Kai Waldrant
    contributor
  • Kaiwen Deng
    contributor
  • Louise Deconinck
    author
  • Robrecht Cannoodt
    authormaintainer
  • Xueer Chen
    contributor
  • Jiwei Liu
    contributor
  • Marius Lange
    contributor

Leaderboard

Methods ranked by scaled overall mean. Each cell encodes a score from 0 to 1 by size and intensity.

QC: Normalisation Visualisation 8 plots

Per metric: points placed by control-anchored scaled score (x); dashed lines mark scaled 0 and 1 (worst/best control); the lower axis shows the raw score. Points beyond [-0.2, 1.2] are clamped to the edge as triangles. Hover a dot or line to highlight it and read details.

methodcontrol
  • MAElower better
    SolutionSolutionscButterflyscButterflyGuanlab-dengkwGuanlab-dengkwKNNR (Py)KNNR (Py)KNNR (R)KNNR (R)Linear ModelLinear ModelCellMapper+PCA/CCACellMapper+PCA/CCASimple MLPSimple MLPNovelNovelMean per geneMean per geneRandom predictionsRandom predictionsZerosZerosBABELBABELSS-OPMSS-OPM1.0190.7640.510.255000.250.50.751rawscaled
  • Mean pearson per cellhigher better
    SolutionSolutionNovelNovelKNNR (Py)KNNR (Py)Linear ModelLinear ModelCellMapper+PCA/CCACellMapper+PCA/CCASimple MLPSimple MLPKNNR (R)KNNR (R)SS-OPMSS-OPMBABELBABELMean per geneMean per geneGuanlab-dengkwGuanlab-dengkwRandom predictionsRandom predictionsscButterflyscButterflyZerosZeros00.250.50.75100.250.50.751rawscaled
  • Mean pearson per genehigher better
    SolutionSolutionNovelNovelKNNR (Py)KNNR (Py)Simple MLPSimple MLPLinear ModelLinear ModelCellMapper+PCA/CCACellMapper+PCA/CCAKNNR (R)KNNR (R)SS-OPMSS-OPMBABELBABELGuanlab-dengkwGuanlab-dengkwscButterflyscButterflyMean per geneMean per geneZerosZerosRandom predictionsRandom predictions-6.1e-30.2450.4970.748100.250.50.751rawscaled
  • Mean spearman per cellhigher better
    SolutionSolutionNovelNovelKNNR (Py)KNNR (Py)CellMapper+PCA/CCACellMapper+PCA/CCASimple MLPSimple MLPKNNR (R)KNNR (R)Linear ModelLinear ModelBABELBABELMean per geneMean per geneGuanlab-dengkwGuanlab-dengkwSS-OPMSS-OPMRandom predictionsRandom predictionsscButterflyscButterflyZerosZeros00.250.50.75100.250.50.751rawscaled
  • Mean spearman per genehigher better
    SolutionSolutionNovelNovelKNNR (Py)KNNR (Py)Linear ModelLinear ModelSimple MLPSimple MLPCellMapper+PCA/CCACellMapper+PCA/CCAKNNR (R)KNNR (R)Guanlab-dengkwGuanlab-dengkwBABELBABELscButterflyscButterflySS-OPMSS-OPMMean per geneMean per geneZerosZerosRandom predictionsRandom predictions-4.3e-30.2470.4980.749100.250.50.751rawscaled
  • Overall pearsonhigher better
    SolutionSolutionNovelNovelKNNR (Py)KNNR (Py)Linear ModelLinear ModelSimple MLPSimple MLPCellMapper+PCA/CCACellMapper+PCA/CCAKNNR (R)KNNR (R)BABELBABELSS-OPMSS-OPMMean per geneMean per geneGuanlab-dengkwGuanlab-dengkwRandom predictionsRandom predictionsscButterflyscButterflyZerosZeros00.250.50.75100.250.50.751rawscaled
  • Overall spearmanhigher better
    SolutionSolutionNovelNovelKNNR (Py)KNNR (Py)Simple MLPSimple MLPLinear ModelLinear ModelCellMapper+PCA/CCACellMapper+PCA/CCAKNNR (R)KNNR (R)BABELBABELSS-OPMSS-OPMMean per geneMean per geneGuanlab-dengkwGuanlab-dengkwRandom predictionsRandom predictionsscButterflyscButterflyZerosZeros00.250.50.75100.250.50.751rawscaled
  • RMSElower better
    SolutionSolutionKNNR (Py)KNNR (Py)NovelNovelLinear ModelLinear ModelSimple MLPSimple MLPCellMapper+PCA/CCACellMapper+PCA/CCAKNNR (R)KNNR (R)Mean per geneMean per geneGuanlab-dengkwGuanlab-dengkwscButterflyscButterflyBABELBABELZerosZerosRandom predictionsRandom predictionsSS-OPMSS-OPM1.341.0050.670.335000.250.50.751rawscaled
QC: Indicator table 31 errors11 warnings

Automated checks on the benchmark run and its results: missing values, score scaling, metric ranges and similar. Errors are high-severity issues that usually need a maintainer's attention; warnings are lower-severity signals. Findings that are expected for this task are listed separately as silenced.

31 high-severity issues need review. 303 of 345 checks passed.

  • error Raw results Task number of results

    Number of results should be equal to #datasets × #methods × #metrics Task: predict_modality Number of results: 712 Number of datasets: 8 Number of methods: 16 Number of metrics: 8 Expected number of results: 1024

  • error Raw results Dataset 'openproblems_neurips2022/pbmc_multiome/swap' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_multiome/swap Number of results: 64 Expected number of results: 128 Percentage missing: 50%

  • error Raw results Dataset 'openproblems_neurips2021/bmmc_cite/normal' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_cite/normal Number of results: 88 Expected number of results: 128 Percentage missing: 31%

  • error Raw results Dataset 'openproblems_neurips2021/bmmc_cite/swap' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_cite/swap Number of results: 88 Expected number of results: 128 Percentage missing: 31%

  • error Raw results Dataset 'openproblems_neurips2022/pbmc_multiome/normal' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_multiome/normal Number of results: 88 Expected number of results: 128 Percentage missing: 31%

  • error Raw results Dataset 'openproblems_neurips2022/pbmc_cite/normal' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_cite/normal Number of results: 88 Expected number of results: 128 Percentage missing: 31%

  • error Raw results Dataset 'openproblems_neurips2022/pbmc_cite/swap' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_cite/swap Number of results: 88 Expected number of results: 128 Percentage missing: 31%

  • error Raw results Method 'cellmapper_scvi' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Method: cellmapper_scvi Number of results: 0 Expected number of results: 64 Percentage missing: 100%

  • error Raw results Method 'guanlab_dengkw_pm' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Method: guanlab_dengkw_pm Number of results: 16 Expected number of results: 64 Percentage missing: 75%

  • error Raw results Method 'babel' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Method: babel Number of results: 8 Expected number of results: 64 Percentage missing: 88%

  • error Raw results Method 'senkin_tmp' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Method: senkin_tmp Number of results: 0 Expected number of results: 64 Percentage missing: 100%

  • error Raw results Method 'scbutterfly' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Method: scbutterfly Number of results: 16 Expected number of results: 64 Percentage missing: 75%

  • error Raw results Metric 'mean_pearson_per_cell' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Metric: mean_pearson_per_cell Number of results: 89 Expected number of results: 128 Percentage missing: 30%

  • error Raw results Metric 'mean_spearman_per_cell' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Metric: mean_spearman_per_cell Number of results: 89 Expected number of results: 128 Percentage missing: 30%

  • error Raw results Metric 'mean_pearson_per_gene' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Metric: mean_pearson_per_gene Number of results: 89 Expected number of results: 128 Percentage missing: 30%

  • error Raw results Metric 'mean_spearman_per_gene' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Metric: mean_spearman_per_gene Number of results: 89 Expected number of results: 128 Percentage missing: 30%

  • error Raw results Metric 'overall_pearson' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Metric: overall_pearson Number of results: 89 Expected number of results: 128 Percentage missing: 30%

  • error Raw results Metric 'overall_spearman' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Metric: overall_spearman Number of results: 89 Expected number of results: 128 Percentage missing: 30%

  • error Raw results Metric 'rmse' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Metric: rmse Number of results: 89 Expected number of results: 128 Percentage missing: 30%

  • error Raw results Metric 'mae' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Metric: mae Number of results: 89 Expected number of results: 128 Percentage missing: 30%

  • error Raw results Dataset 'openproblems_neurips2022/pbmc_multiome/swap' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_multiome/swap Succeeded processes: 8 Attempted processes: 16 Percentage failed: 50%

  • error Raw results Dataset 'openproblems_neurips2021/bmmc_cite/normal' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_cite/normal Succeeded processes: 11 Attempted processes: 16 Percentage failed: 31%

  • error Raw results Dataset 'openproblems_neurips2021/bmmc_cite/swap' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_cite/swap Succeeded processes: 11 Attempted processes: 16 Percentage failed: 31%

  • error Raw results Dataset 'openproblems_neurips2022/pbmc_multiome/normal' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_multiome/normal Succeeded processes: 11 Attempted processes: 16 Percentage failed: 31%

  • error Raw results Dataset 'openproblems_neurips2022/pbmc_cite/normal' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_cite/normal Succeeded processes: 11 Attempted processes: 16 Percentage failed: 31%

  • error Raw results Dataset 'openproblems_neurips2022/pbmc_cite/swap' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_cite/swap Succeeded processes: 11 Attempted processes: 16 Percentage failed: 31%

  • error Raw results Method 'cellmapper_scvi' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Method: cellmapper_scvi Succeeded processes: 0 Attempted processes: 8 Percentage failed: 100%

  • error Raw results Method 'guanlab_dengkw_pm' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Method: guanlab_dengkw_pm Succeeded processes: 2 Attempted processes: 8 Percentage failed: 75%

  • error Raw results Method 'babel' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Method: babel Succeeded processes: 1 Attempted processes: 8 Percentage failed: 88%

  • error Raw results Method 'senkin_tmp' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Method: senkin_tmp Succeeded processes: 0 Attempted processes: 8 Percentage failed: 100%

  • error Raw results Method 'scbutterfly' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Method: scbutterfly Succeeded processes: 2 Attempted processes: 8 Percentage failed: 75%

Show 11 warnings
  • warning Raw results Method 'novel' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Method: novel Number of results: 48 Expected number of results: 64 Percentage missing: 25%

  • warning Raw results Method 'novel' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Method: novel Succeeded processes: 6 Attempted processes: 8 Percentage failed: 25%

  • warning Raw results Dataset 'openproblems_neurips2021/bmmc_multiome/normal' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_multiome/normal Number of results: 104 Expected number of results: 128 Percentage missing: 19%

  • warning Raw results Dataset 'openproblems_neurips2021/bmmc_multiome/swap' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_multiome/swap Number of results: 104 Expected number of results: 128 Percentage missing: 19%

  • warning Raw results Method 'simple_mlp' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Method: simple_mlp Number of results: 56 Expected number of results: 64 Percentage missing: 12%

  • warning Raw results Method 'ss_opm' % missing

    Percentage of missing results should be less than 10% Task: predict_modality Method: ss_opm Number of results: 56 Expected number of results: 64 Percentage missing: 12%

  • warning Raw results Task number of successful processes

    Number of successful processes should be equal to the number of attempted processes Task: predict_modality Succeeded processes: 267 Attempted processes: 306 Percentage failed: 13%

  • warning Raw results Dataset 'openproblems_neurips2021/bmmc_multiome/normal' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_multiome/normal Succeeded processes: 13 Attempted processes: 16 Percentage failed: 19%

  • warning Raw results Dataset 'openproblems_neurips2021/bmmc_multiome/swap' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_multiome/swap Succeeded processes: 13 Attempted processes: 16 Percentage failed: 19%

  • warning Raw results Method 'simple_mlp' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Method: simple_mlp Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

  • warning Raw results Method 'ss_opm' % failed

    Percentage of failed processes should be less than 10% Task: predict_modality Method: ss_opm Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%

Method info 12

Cross-modal autoencoder predicting scRNA-seq expression from scATAC-seq accessibility (Wu et al. 2021)

BABEL is a deep autoencoder that learns modality-specific encoders and decoders for paired scRNA-seq and scATAC-seq measurements. RNA is encoded/decoded with dense layers; ATAC accessibility is encoded with a chromosome-split architecture (per-chromosome sub-networks, matching genomic locality) that groups peaks by chromosome parsed from their genomic-coordinate names (e.g. "chr1:1000-2000"). The full architecture is bidirectional (RNA->RNA, RNA->ATAC, ATAC->RNA, ATAC->ATAC), trained jointly with a negative binomial reconstruction loss for RNA counts and a binary cross-entropy loss for binarized ATAC accessibility, but this component only exposes the ATAC->RNA (GEX) prediction direction. The RNA->ATAC direction was found to collapse to a per-peak base-rate prediction that ignores the RNA input -- a known failure mode of unweighted binary cross-entropy under the extreme class imbalance typical of ATAC accessibility data (~3% positive) -- confirmed via near-zero discrimination between accessible/inaccessible peaks even after training to convergence, and is intentionally disabled rather than silently returned.

Modality prediction in a PCA/CCA space using CellMapper

CellMapper is a general framework for k-NN based mapping tasks in single-cell and spatial genomics. This variant uses CellMapper to project modalities from a reference dataset (train) onto a query dataset (test) in a PCA/CCA latent space.

Modality prediction in an scVI latent space using CellMapper

CellMapper is a general framework for k-NN based mapping tasks in single-cell and spatial genomics. This variant uses CellMapper to project modalities from a reference dataset (train) onto a query dataset (test) in a modality-specific latent space computed with suitable scvi-tools models. For gene expression data, we use the scVI model on raw counts (nb likelihood), for ADT data, we use the scVI models on normalized counts (gaussian likelihood), and for ATAC data, we use the PeakVI model on raw counts. The actual CellMapper pipeline is modality-agnostic.

A kernel ridge regression method with RBF kernel.

This is a solution developed by Team Guanlab - dengkw in the Neurips 2021 competition to predict one modality from another using kernel ridge regression (KRR) with RBF kernel. Truncated SVD is applied on the combined training and test data from modality 1 followed by row-wise z-score normalization on the reduced matrix. The truncated SVD of modality 2 is predicted by training a KRR model on the normalized training matrix of modality 1. Predictions on the normalized test matrix are then re-mapped to the modality 2 feature space via the right singular vectors.

K-nearest neighbor regression in Python.

K-nearest neighbor regression in R.

Linear model regression.

A linear model regression method.

A method using encoder-decoder MLP model

This method trains an encoder-decoder MLP model with one output neuron per component in the target. As an input, the encoders use representations obtained from ATAC and GEX data via LSI transform and raw ADT data. The hyperparameters of the models were found via broad hyperparameter search using the Optuna framework.

Dual-VAE adversarial translator for paired single-cell multi-omics (GEX<->ATAC)

scButterfly (Basic variant, scButterfly-B) is a dual variational autoencoder with an adversarial translator that learns to convert between paired single-cell modalities. For the predict-modality task it is trained on paired Multiome cells and used to translate the held-out test modality. Chromosome grouping for the ATAC branch is parsed from peak coordinates in the feature names. Only Multiome GEX<->ATAC is supported; CITE-seq (ADT) datasets are not handled.

LightGBM + bidirectional GRU ensemble for CITE-seq protein prediction (OpenProblems 2022 2nd place)

Two-stage method from the OpenProblems NeurIPS 2021 competition. Stage 1 trains four LightGBM models on different RNA feature representations (log-normalized, CLR-TSVD, custom sqrt-normalized, and raw counts). Stage 2 refines predictions with two neural network architectures: a bidirectional GRU with cosine-similarity loss and a dense bidirectional GRU with MSE loss. Final predictions are a weighted blend (55% cosine, 45% MSE) of per-fold averaged outputs.

Ensemble of MLPs trained on different sites (team AXX)

This folder contains the AXX solution to the OpenProblems-NeurIPS2021 Single-Cell Multimodal Data Integration. Team took the 4th place of the modality prediction task in terms of overall ranking of 4 subtasks: namely GEX to ADT, ADT to GEX, GEX to ATAC and ATAC to GEX. Specifically, our methods ranked 3rd in GEX to ATAC and 4th in GEX to ADT. More details about the task can be found in the competition webpage.

1st place solution of the Kaggle Open Problems Multimodal Single-Cell Integration challenge.

Encoder-decoder MLP method using SVD-based dimensionality reduction for both inputs and targets, followed by batch-median correction. The encoder maps (optionally augmented) cell embeddings to a latent space; multiple decoder blocks predict target expression in the SVD-compressed space. The method was the winning solution of the NeurIPS 2021 Open Problems Multimodal Single-Cell Integration Kaggle competition.

Control method info 4
Mean per gene

Returns the mean expression value per gene.

Random predictions

Returns random training profiles.

Solution

Returns the ground-truth solution.

Zeros

Returns a prediction consisting of all zeros.

Metric info 8
MAElower is betterChai & Draxler, 2014

The mean absolute error.

The average difference between the expression values and the predicted expression values.

Mean pearson per cellhigher is betterPearson, 1895

The mean of the pearson values of per-cell expression value vectors.

The mean of the pearson values of per-cell expression value vectors. A per-cell vector with zero variance contributes a correlation of 0 rather than NA.

Mean pearson per genehigher is betterPearson, 1895

The mean of the pearson values of per-gene expression value vectors.

The mean of the pearson values of per-gene expression value vectors. A per-gene vector with zero variance contributes a correlation of 0 rather than NA.

Mean spearman per cellhigher is betterKENDALL, 1938

The mean of the spearman values of per-cell expression value vectors.

The mean of the spearman values of per-cell expression value vectors. A per-cell vector with zero variance contributes a correlation of 0 rather than NA.

Mean spearman per genehigher is betterKENDALL, 1938

The mean of the spearman values of per-gene expression value vectors.

The mean of the spearman values of per-gene expression value vectors. A per-gene vector with zero variance contributes a correlation of 0 rather than NA.

Overall pearsonhigher is betterPearson, 1895

The pearson correlation of the vectorized expression matrices.

The pearson correlation between the solution and the prediction, with both matrices flattened into a single vector. A constant matrix scores 0 rather than NA.

Overall spearmanhigher is betterKENDALL, 1938

The spearman correlation of the vectorized expression matrices.

The spearman correlation between the solution and the prediction, with both matrices flattened into a single vector. A constant matrix scores 0 rather than NA.

RMSElower is betterChai & Draxler, 2014

The root mean squared error.

The square root of the mean of the square of all of the error.

Dataset info 8
NeurIPS2021 CITE-Seq (ADT2GEX) unlinked

Single-cell CITE-Seq (GEX+ADT) data collected from bone marrow mononuclear cells of 12 healthy human donors.

Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X 3 prime Single-Cell Gene Expression kit with Feature Barcoding in combination with the BioLegend TotalSeq B Universal Human Panel v1.0. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2021. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.

NeurIPS2021 CITE-Seq (GEX2ADT) unlinked

Single-cell CITE-Seq (GEX+ADT) data collected from bone marrow mononuclear cells of 12 healthy human donors.

Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X 3 prime Single-Cell Gene Expression kit with Feature Barcoding in combination with the BioLegend TotalSeq B Universal Human Panel v1.0. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2021. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.

NeurIPS2021 Multiome (ATAC2GEX) unlinked

Single-cell Multiome (GEX+ATAC) data collected from bone marrow mononuclear cells of 12 healthy human donors.

Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X Multiome Gene Expression and Chromatin Accessibility kit. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2021. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.

NeurIPS2021 Multiome (GEX2ATAC) unlinked

Single-cell Multiome (GEX+ATAC) data collected from bone marrow mononuclear cells of 12 healthy human donors.

Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X Multiome Gene Expression and Chromatin Accessibility kit. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2021. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.

OpenProblems NeurIPS2022 CITE-Seq (ADT2GEX) unlinked

Single-cell CITE-Seq (GEX+ADT) data collected from bone marrow mononuclear cells of 12 healthy human donors.

Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X 3 prime Single-Cell Gene Expression kit with Feature Barcoding in combination with the BioLegend TotalSeq B Universal Human Panel v1.0. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2022. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.

OpenProblems NeurIPS2022 CITE-Seq (GEX2ADT) unlinked

Single-cell CITE-Seq (GEX+ADT) data collected from bone marrow mononuclear cells of 12 healthy human donors.

Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X 3 prime Single-Cell Gene Expression kit with Feature Barcoding in combination with the BioLegend TotalSeq B Universal Human Panel v1.0. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2022. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.

OpenProblems NeurIPS2022 Multiome (ATAC2GEX) unlinked

Single-cell Multiome (GEX+ATAC) data collected from bone marrow mononuclear cells of 12 healthy human donors.

Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X Multiome Gene Expression and Chromatin Accessibility kit. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2022. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.

OpenProblems NeurIPS2022 Multiome (GEX2ATAC) unlinked

Single-cell Multiome (GEX+ATAC) data collected from bone marrow mononuclear cells of 12 healthy human donors.

Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X Multiome Gene Expression and Chromatin Accessibility kit. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2022. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.

References

  1. ... (2024). Predicting cellular profiles across modalities in longitudinal single-cell data: An Open Problems competition. In Preparation.
  2. https://doi.org/10.1073/pnas.202307011
  3. Cao, Y., Zhao, X., Tang, S., Jiang, Q., Li, S., Li, S., & Chen, S. (2024). scButterfly: a versatile single-cell cross-modality translation method via dual-aligned variational autoencoders. 10.1038/s41467-024-47418-x ↗
  4. Chai, T., & Draxler, R. R. (2014). Root mean square error (RMSE) or mean absolute error (MAE)? 10.5194/gmdd-7-1525-2014 ↗
  5. Fix, E., & Hodges, J. L. (1989). Discriminatory Analysis. Nonparametric Discrimination: Consistency Properties. International Statistical Review / Revue Internationale de Statistique, 57(3), 238. 10.2307/1403797 ↗
  6. KENDALL, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1–2), 81–93. 10.1093/biomet/30.1-2.81 ↗
  7. Lance, C., Luecken, M. D., Burkhardt, D. B., Cannoodt, R., Rautenstrauch, P., Laddach, A., Ubingazhibov, A., Cao, Z.-J., Deng, K., Khan, S., Liu, Q., Russkikh, N., Ryazantsev, G., Ohler, U., Pisco, A. O., Bloom, J., Krishnaswamy, S., & Theis, F. J. (2022). Multimodal single cell data integration challenge: results and lessons learned. bioRxiv. 10.1101/2022.04.11.487796 ↗
  8. Lance, C., Shitov, V. A., Wen, H., Ji, Y., Holderrieth, P., Wu, Y., Liu, R., Cannoodt, R., Tang, W., Waldrant, K., DeMeo, B., Cortes, M., Kotlarz, D., Tang, J., Xie, Y., Theis, F. J., Burkhardt, D. B., & Luecken, M. D. (2026). Longitudinal modality prediction learns gene regulatory patterns: insights from a single-cell competition. 10.64898/2026.02.24.707614 ↗
  9. Lange, M. (2025). quadbio/cellmapper: v0.2.2. 10.5281/ZENODO.15683594 ↗
  10. Luecken, M., Burkhardt, D., Cannoodt, R., Lance, C., Agrawal, A., Aliee, H., Chen, A., Deconinck, L., Detweiler, A., Granados, A., Huynh, S., Isacco, L., Kim, Y., Klein, D., DE KUMAR, B., Kuppasani, S., Lickert, H., McGeever, A., Melgarejo, J., … Bloom, J. M. (2021). A sandbox for prediction and integration of DNA, RNA, and proteins in single cells. In J. Vanschoren & S. Yeung (Eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (Vol. 1). Curran. link ↗
  11. Luecken, M. D., Burkhardt, D. B., Cannoodt, R., Lance, C., Agrawal, A., Aliee, H., Chen, A. T., Deconinck, L., Detweiler, A. M., Granados, A. A., & others. (2021). A sandbox for prediction and integration of DNA, RNA, and proteins in single cells. Thirty-Fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  12. Pearson, K. (1895). VII. Note on regression and inheritance in the case of two parents. Proceedings of the Royal Society of London, 58(347–352), 240–242. 10.1098/rspl.1895.0041 ↗
  13. Wilkinson, G. N., & Rogers, C. E. (1973). Symbolic Description of Factorial Models for Analysis of Variance. Applied Statistics, 22(3), 392. 10.2307/2346786 ↗