Discover our new platform: Learn more

Five Common Mistakes in Spatial Transcriptomics Analysis – Part I

BioTuring Science Team
BioTuring Science Team
August 14, 2026

Spatial transcriptomics experiments require substantial investments in tissue preparation, specialized instrumentation, reagents, and computational analysis. Yet some of the most consequential errors occur after the data has been generated, during data analysis and interpretation.

Unlike experimental failures, analytical mistakes rarely produce obvious warnings. Clusters still form, cell types can be assigned, and differential expression or cell–cell communication analyses still generate convincing results. However, a computational result is not necessarily a biological discovery. Small analytical decisions made early in the workflow can silently introduce bias that propagates through every downstream analysis.

Producing reliable biological insights requires more than running computational pipelines. Researchers must continuously evaluate whether quality control preserves meaningful biology, whether cell segmentation accurately reflects tissue architecture, whether cell type annotations are well supported, and whether downstream analyses are biologically plausible.

In this two-part series, we discuss common analytical pitfalls in spatial transcriptomics. Part I focuses on three critical early-stage steps—quality control, cell segmentation, and cell type annotation—while Part II explores challenges in differential expression, spatial neighborhood analysis, and cell–cell communication.

Mistake 1: Over-Filtering During Quality Control

Quality control (QC) aims to remove technical artifacts such as damaged cells, background signals, and low-quality measurements. However, overly stringent filtering can also remove biologically meaningful cells, reducing the accuracy of every downstream analysis.

One of the most common mistakes is applying QC thresholds from another study without considering differences in tissue type, biological context, sequencing depth, or spatial platform. A threshold that performs well for one dataset may unintentionally remove rare or low-transcript cell populations in another.

Because QC is the first analytical step, its effects propagate throughout the entire workflow. Cells removed during filtering never appear in clustering, cell type annotation, differential expression, or spatial interaction analyses. The dataset may appear cleaner, but important biology may already have been lost.

Real-World Impact

Consider a Visium HD colon tissue dataset available in BioTuring SpatialX. Figure 1 compares two mitochondrial RNA (MT%) filtering thresholds: 10% and 25%.

At first glance, the tissue images appear nearly identical after filtering. However, the downstream UMAP tells a different story. Using the more stringent threshold removes more than 10,000 cells, substantially changing cluster structure and eliminating several cell populations from subsequent analyses.

This illustrates an important principle: a cleaner dataset is not always a better dataset. Strict QC thresholds can unintentionally remove biologically relevant populations, including immune cells or other naturally low-transcript cell types that contain fewer RNA molecules than highly transcriptionally active cells 1. Once these cells are removed, they cannot contribute to downstream analyses, potentially leading to incomplete or misleading biological conclusions.

Best Practices

Rather than relying on fixed thresholds, QC decisions should be guided by both data quality and biological context.

  • Examine QC distributions before selecting thresholds.
  • Review cells near filtering boundaries rather than removing them automatically.
  • Determine whether excluded cells represent technical artifacts or biologically meaningful populations.
  • Consider expected tissue biology when interpreting QC metrics.
  • Use platform-specific QC recommendations rather than universal cutoffs.
  • Revisit QC thresholds iteratively as biological results emerge.

The goal of QC is not to remove as many cells as possible, but to remove unreliable measurements while preserving meaningful biological information.

Figure 1: Effect of mitochondrial RNA (MT%) filtering thresholds on a Visium HD colon dataset. Applying a more stringent MT threshold (10%) removes more than 10,000 cells compared with a 25% threshold. Although the tissue images appear similar, the resulting UMAPs show substantial changes in cluster structure and cell composition, illustrating how QC decisions can influence downstream biological interpretation.

Mistake 2: Treating Cell Segmentation as a Black Box

In imaging-based spatial transcriptomics, cell segmentation determines how detected transcripts are assigned to individual cells. Because nearly every downstream analysis operates at the cell level, segmentation is one of the most critical steps in the analysis workflow 2.

Unlike quality control, which determines which cells are included in the analysis, segmentation determines how transcripts are assigned to those cells. Errors at this stage therefore change the biological identity of the cells themselves, influencing every downstream analysis—from cell type annotation to differential expression and spatial interaction analysis.

Segmentation errors generally occur in two forms:

  • Under-segmentation: Neighboring cells are merged into one artificial profile.
  • Over-segmentation: One cell is incorrectly divided into multiple objects.

Both errors distort gene expression profiles and can propagate through clustering, cell type annotation, differential expression, spatial neighborhood analysis, and cell–cell communication analyses.

A Real-World Example

Figure 2 shows an ovarian FFPE sample profiled with 10x Genomics Xenium. When segmentation incorporates morphological information (left), neighboring epithelial and immune cells are accurately separated. Without morphological guidance (right), several immune cells are merged with adjacent epithelial cells (highlighted in the light green circle), producing incorrect cell boundaries.

As a result, transcripts originating from multiple cells are assigned to a single segmented object. The resulting expression profile may simultaneously contain epithelial markers such as EPCAM and T-cell markers such as CD3D/CD3E (or whichever marker is actually shown in your figure), creating the appearance of a hybrid cell population.

Although clustering algorithms may identify this population as a distinct cluster, it is not a novel biological state—it is a technical artifact caused by inaccurate segmentation 2.

Why It Is Dangerous

Segmentation errors can create convincing but misleading biological signals. Under-segmentation may generate artificial hybrid cell states, whereas over-segmentation can fragment individual cells, increasing the apparent cellular heterogeneity.

Because these artifacts are derived from genuine transcript measurements, they often pass standard computational quality checks and can only be identified through careful inspection of the tissue image.

Best Practices

Rather than accepting segmentation results as a black box:

  • Overlay segmentation boundaries on the raw morphology or DAPI image.
  • Inspect representative tissue regions manually.
  • Examine unusually large or small segmented cells.
  • Check whether transcript assignments match realistic cell morphology.
  • Validate unexpected cell populations using known marker genes and spatial context.

Ultimately, a computational cell should correspond as closely as possible to a biological cell, not simply an object generated by a segmentation algorithm.

Figure 2: Cell segmentation of an ovarian FFPE sample profiled with 10x Genomics Xenium. Incorporating morphological information during segmentation (left) improves separation of adjacent epithelial and immune cells. Without morphology-guided segmentation (right), multiple neighboring cells are merged into a single segmented object (light green circle), leading to incorrect transcript assignment and potentially misleading downstream analyses.

Mistake 3: Trusting a Single, Unvalidated Cell-Type Annotation

Cell-type annotation is one of the most important interpretation steps in spatial transcriptomics because it forms the basis for nearly every downstream analysis. Whether researchers are comparing cell populations, identifying differentially expressed genes, or studying spatial interactions, the biological conclusions depend on accurate cell identities.

Numerous computational methods have been developed for cell-type annotation, but no single method consistently outperforms all others across tissues, diseases, and experimental platforms. Some rely on reference atlases to match expression profiles to previously characterized cell types, whereas others use marker gene enrichment or scoring methods. Each approach has its own strengths and limitations. Reference-based methods depend on the quality and relevance of the reference dataset, while marker-based methods require carefully selected and specific marker genes. Neither approach performs optimally for every tissue or biological condition.

This limitation becomes particularly important in disease tissues, developmental samples, and highly perturbed biological systems, where cell states may differ substantially from those represented in existing references.

A Real-World Example

Figure 3 illustrates how annotation results can change even within the same dataset. Using the Visium HD colon sample introduced in Mistake 1, a cell that passed quality control under both filtering strategies received different annotation results. Under one QC setting, the cell remained unassigned, whereas under another it was automatically labeled as a CD4 T cell, despite both analyses retaining the same cell (MT% <10%).

This example highlights an important point: cell-type annotation is influenced not only by the annotation algorithm itself but also by upstream analytical decisions such as quality control and data preprocessing. A confident label should therefore not be interpreted as definitive biological evidence.

Why It Matters

Annotation errors rarely remain isolated. Misclassified cells propagate through downstream analyses, influencing differential expression, spatial neighborhood analysis, and cell–cell communication. More importantly, disagreement between annotation methods is not always a problem—it can be biologically informative. Conflicting predictions may indicate rare cell populations, transitional states, disease-associated phenotypes, or cells that require additional investigation.

Best Practices

Rather than relying on a single annotation method, combine multiple sources of evidence whenever possible. Reference-based annotation, marker gene enrichment, and knowledge-assisted approaches each capture different aspects of cell identity, and agreement among them generally provides greater confidence than any single method alone.

  • Compare multiple annotation strategies, such as reference-based and marker-based methods.
  • Validate predicted cell identities using canonical marker genes.
  • Examine the spatial location of annotated cells to determine whether they are biologically plausible.
  • Investigate clusters where different methods produce conflicting labels instead of forcing immediate classification.
  • Treat uncertain annotations as candidates for further validation rather than definitive cell identities.

Ultimately, cell-type annotation is an interpretation step, not simply a labeling step. Computational predictions provide valuable guidance, but reliable biological conclusions require integrating algorithmic results with marker gene expression, tissue context, and biological knowledge.

Figure 3: Cell-type annotation of a Visium HD colon dataset following two different quality control strategies. Although the same cell was retained in both analyses (MT% <10%), it remained unassigned under one workflow but was automatically annotated as a CD4 T cell under the other. This example illustrates how upstream analytical decisions can influence downstream cell-type annotation and highlights the importance of validating automated predictions.

Want to see what these three steps look like done right, on your own dataset? Request a demo and see QC, segmentation, and annotation checked visually with SpatialX.

Request a Demo →

References

1. Subramanian, A., Alperovich, M., Yang, Y. & Li, B. Biology-inspired data-driven quality control for scientific discovery in single-cell transcriptomics. Genome Biol. 23, 267 (2022).

2. Mitchel, J., Gao, T., Petukhov, V., Cole, E. & Kharchenko, P. V. Impact and correction of segmentation errors in spatial transcriptomics. Nat. Genet. 58, 434–444 (2026).

0 comments