Discover our new platform: Learn more

Five Common Mistakes in Spatial Transcriptomics Analysis – Part II

BioTuring Science Team
BioTuring Science Team
August 20, 2026

In Part I of this series, we discussed how early analytical decisions—including quality control, cell segmentation, and cell-type annotation—establish the foundation for reliable spatial transcriptomics analysis. Even with these steps carefully optimized, important challenges remain during biological interpretation. In this article, we focus on two major pitfalls in downstream analysis: treating differential expression results as biological conclusions, and treating spatial associations or computational predictions as biological proof.

Mistake 4: Treating Gene Lists as Biological Conclusions

Differential expression (DE) analysis is one of the most common ways to interpret spatial transcriptomics data. However, not all DE analyses ask the same biological question.

One common use is marker identification: finding genes that distinguish one cell population or cluster from others. This can help characterize clusters or support cell-type annotation.

A different question is condition-level differential expression: asking whether the same cell population differs between biological conditions—for example, whether fibroblasts in tumor tissue differ from fibroblasts in normal tissue, or whether a cell population changes after treatment. Here, the goal is to determine whether expression differences are consistent across biological samples.

These two analyses should not be interpreted in the same way. A ranked list of marker genes can help describe a cell population, but it does not by itself demonstrate that a gene is consistently different between biological conditions.

Why It Matters

For marker identification, genes are typically ranked according to their enrichment in one cell population or cluster relative to others. This can reveal characteristic markers and help define cellular populations.

However, when comparing biological conditions, treating individual cells as independent replicates can lead to pseudoreplication, because many cells originate from the same biological sample. Thousands of cells from one patient, for example, do not represent thousands of independent biological replicates.

For multi-sample comparisons, one common approach is pseudobulk analysis, in which expression counts are aggregated for a defined cell population within each biological sample before comparing conditions. Other statistical models can also account for the hierarchical structure of cells nested within biological samples.

The key question is therefore not simply: “Which genes are different?” but “What biological question am I asking, and what is the appropriate unit of comparison?

From Marker Genes to Biological Interpretation

Even when a DE analysis is statistically appropriate, interpretation should not stop at the resulting gene list.

A single highly ranked gene may highlight one aspect of a cellular state without explaining the broader biological process. Examining the spatial distribution of important genes and evaluating coordinated changes across pathways or gene sets can provide additional context.

For example, several genes involved in extracellular matrix organization, inflammatory signaling, or cell–cell interactions may change together. Pathway-level analysis can therefore help move from individual gene differences toward a broader interpretation of the biological processes associated with the observed changes.

Real-World Example

A useful workflow is to move through several levels of evidence:

Cell population → Differential expression → Spatial localization → Pathway interpretation

For example, in the colon Visium HD dataset introduced in Part I, two CD4⁺ T-cell populations were selected from distinct spatial regions of the tissue (Figure 1A). CD4 and other T-cell markers supported their cell-type identity, while their spatial locations indicated that they occupied different tissue environments. 

Differential expression analysis between the two T-cell populations identified a number of significantly different genes (Figure 1B). However, stopping at the gene level can make biological interpretation difficult. For example, CD48, one of the highly ranked genes, showed higher expression in one region than the other. CD48 is expressed across multiple hematopoietic cell types and has been implicated in immune-cell activation and tumor-associated immune regulation 1,2. Its higher expression therefore does not, by itself, establish a specific functional state of these CD4⁺ T cells.

Rather than interpreting individual genes in isolation, pathway-level analysis provides a broader view of the biological programs that distinguish the two spatial populations (Figure 1C). The enrichment profiles revealed differences in pathways associated with epithelial–mesenchymal transition, IL6–JAK–STAT3 signaling, TNF-α signaling, hypoxia, and other cellular processes. Together with their distinct spatial locations, these patterns suggest that the two CD4⁺ T-cell populations are associated with different tissue microenvironments and may exhibit differences in their molecular states.

This illustrates an important principle: a DEG list tells you which genes differ, but pathway analysis can help explain what those differences may represent biologically. Spatial localization adds another layer of evidence by showing where those molecular programs occur within the tissue.

The goal is therefore not to replace gene-level analysis with pathway analysis, but to integrate both levels of evidence. Individual genes can highlight interesting candidates, while coordinated changes across gene sets and their spatial distribution provide stronger support for biological interpretation.

How to Catch It

  • First define whether the analysis is marker identification or condition-level differential expression.
  • For condition-level comparisons, account for biological replicates rather than treating individual cells as independent samples.
  • Consider pseudobulk or other statistical approaches that appropriately model the sample structure.
  • Consider effect size alongside statistical significance.
  • Examine the spatial distribution of important DE genes.
  • Avoid using a single marker as the explanation for a complex biological state.
  • Evaluate gene sets and pathways to identify coordinated biological processes.
  • Check whether molecular changes are consistent with tissue architecture and known biology.

Differential expression identifies molecular differences; biological interpretation requires understanding what those differences mean, where they occur, and whether they are reproducible across biological samples.

Figure 1: From gene-level differences to biological interpretation in spatial transcriptomics. (A) Two spatially distinct CD4⁺ T-cell populations selected from a Visium HD colon dataset, with CD4 marker expression shown spatially. (B) Differentially expressed genes between the two T-cell populations, ranked by statistical significance. (C) Pathway enrichment analysis reveals coordinated biological programs that distinguish the spatial populations, providing additional context beyond individual gene-level differences.

Mistake 5: Treating Spatial Associations and Computational Predictions as Biological Proof

Spatial transcriptomics enables researchers to investigate where cells are located, which populations tend to occur together, and which molecular interactions may occur between them. Neighborhood analysis and cell–cell communication analysis can therefore provide valuable clues about tissue organization and potential signaling mechanisms.

However, these analyses generate hypotheses rather than direct proof of biological mechanisms.

Spatial proximity does not necessarily mean biological interaction. Likewise, detecting ligand and receptor transcripts does not demonstrate that the corresponding proteins are present, physically interacting, or functionally influencing neighboring cells.

The key mistake is treating a computational prediction as if it were experimental evidence.

Why It Matters

Spatial relationships can arise for many reasons. Two cell populations may be located close to each other because they:

  • share a local tissue environment,
  • respond to common biological signals,
  • occupy the same anatomical region,
  • or have correlated molecular states.

Similarly, ligand–receptor analyses typically identify potential interactions based on the expression of ligand and receptor genes. Expression alone does not establish that signaling is occurring.

This is particularly important because computational cell–cell communication methods can produce many statistically ranked interactions. Incorporating spatial constraints can help distinguish biologically plausible interactions from associations that may arise simply from expression patterns or shared tissue environments.

Illustrative Example

Consider a hypothetical cell–cell communication analysis that predicts a CXCL12–CXCR4 interaction between podocytes and T cells in a kidney biopsy.

A researcher might conclude that podocytes are actively recruiting T cells through chemokine signaling. However, mapping the predicted interaction back onto tissue coordinates could reveal that podocytes are confined within the glomerulus while T cells are distributed elsewhere. The computational result may therefore represent an interesting hypothesis, but the spatial arrangement would not support a direct local interaction.

Importantly, the CXCL12–CXCR4 axis itself has established biological relevance in the kidney. Previous work demonstrated CXCL12 expression by podocytes and CXCR4 expression by adjacent glomerular endothelial cells, with functional evidence linking this signaling axis to renal vascular development 3.

The example above is illustrative rather than a reported podocyte–T-cell interaction. Its purpose is to show why a biologically plausible ligand–receptor pair still needs to be evaluated in its actual spatial and cellular context.

How to Catch It

Before interpreting a predicted spatial interaction as a biological mechanism:

  • Map the interacting populations back onto tissue coordinates.
  • Check whether the populations are physically close enough for the proposed interaction to be plausible.
  • Examine ligand and receptor expression in the relevant cells or spatial regions.
  • Consider whether the interaction is consistent with tissue architecture and known biology.
  • Compare the prediction with pathway-level evidence and relevant literature.
  • Distinguish clearly between a predicted interaction and an experimentally demonstrated mechanism.
  • Where possible, validate important hypotheses using orthogonal approaches, such as protein-level measurements or functional experiments.

The goal is not to discard computational predictions. Rather, they should be used to prioritize biologically testable hypotheses.

Spatial analysis can tell you where to look for a potential biological interaction; it does not, by itself, prove that the interaction occurs.

The Common Thread: Validate Before You Interpret

Figure 5: An iterative workflow for validating spatial analysis results. Computational predictions provide starting points for biological interpretation, but reliable conclusions require iterative evaluation of spatial context, molecular evidence, biological knowledge, and independent validation. Findings can be revisited and refined as new evidence emerges. At every step, ask: “Is this result reliable, and how can it be verified?”

Having the data and the computational pipeline is not enough. Meaningful interpretation also requires biological background knowledge and a clear understanding of the question being asked.

Most importantly, throughout the analysis, researchers should continually ask:

Is this result reliable, and how can I verify it?

That question—rather than any particular computational method—should remain at the center of scientific research.

Making Spatial Validation Easier With SpatialX

Reviewing QC decisions, checking segmentation quality, comparing annotation methods, interpreting gene expression changes, and validating spatial relationships often requires moving between multiple analysis tools.

BioTuring’s SpatialX brings many of these validation steps into a single interactive workflow, allowing researchers to inspect quality metrics, evaluate segmentation, compare annotations, and connect computational findings back to tissue context.

Supporting major spatial transcriptomics technologies—including 10x Genomics Visium, Visium HD, Xenium, Bruker CosMx, Vizgen MERSCOPE, and STOmics Stereo-seq—SpatialX helps researchers move beyond automated outputs and build spatial transcriptomics results they can interpret with confidence.

Ready to go beyond gene lists and p-values? 

See how SpatialX helps you validate spatial findings with confidence. 

Request a Demo →

References 

1.        Shi, G. et al. CD48 is a novel immune checkpoint on tumour‐associated macrophages in hepatocellular carcinoma. Gut gutjnl-2025-336744 (2026) doi:10.1136/gutjnl-2025-336744.

2.        McArdel, S. L., Terhorst, C. & Sharpe, A. H. Roles of CD48 in regulating immunity and tolerance. Clinical Immunology 164, 10–20 (2016).

3.        Takabatake, Y. et al. The CXCL12 (SDF-1)/CXCR4 Axis Is Essential for the Development of Renal Vasculature. Journal of the American Society of Nephrology 20, 1714–1723 (2009).

0 comments