
==== Front
Bull Math Biol
Bull Math Biol
Bulletin of Mathematical Biology
0092-8240
1522-9602
Springer US New York

39287883
1353
10.1007/s11538-024-01353-6
Original Article
Relational Persistent Homology for Multispecies Data with Application to the Tumor Microenvironment
Stolz Bernadette J. 12
Dhesi Jagdeep 2
Bull Joshua A. 2
Harrington Heather A. 23
Byrne Helen M. 24
http://orcid.org/0000-0003-0638-9195
Yoon Iris H. R. hyoon@wesleyan.edu

25
1 grid.5333.6 0000000121839049 Laboratory for Topology and Neuroscience, EPFL, Station 8, Lausanne, 1015 Switzerland
2 https://ror.org/052gg0110 grid.4991.5 0000 0004 1936 8948 Mathematical Institute, University of Oxford, Andrew Wiles Building, Woodstock Rd, Oxford, OX2 6GG UK
3 grid.4991.5 0000 0004 1936 8948 Wellcome Centre for Human Genetics, University of Oxford, Roosevelt Dr, Headington, Headington, Oxford, OX3 7BN UK
4 grid.4991.5 0000 0004 1936 8948 Ludwig Institute for Cancer Research, University of Oxford, Old Road Campus Research Build, Roosevelt Dr, Headington, Oxford, OX3 7DQ UK
5 https://ror.org/05h7xva58 grid.268117.b 0000 0001 2293 7601 Department of Mathematics and Computer Science, Wesleyan University, 265 Church Street, Middletown, 06459 USA
17 9 2024
17 9 2024
2024
86 11 12817 8 2023
29 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
Topological data analysis (TDA) is an active field of mathematics for quantifying shape in complex data. Standard methods in TDA such as persistent homology (PH) are typically focused on the analysis of data consisting of a single entity (e.g., cells or molecular species). However, state-of-the-art data collection techniques now generate exquisitely detailed multispecies data, prompting a need for methods that can examine and quantify the relations among them. Such heterogeneous data types arise in many contexts, ranging from biomedical imaging, geospatial analysis, to species ecology. Here, we propose two methods for encoding spatial relations among different data types that are based on Dowker complexes and Witness complexes. We apply the methods to synthetic multispecies data of a tumor microenvironment and analyze topological features that capture relations between different cell types, e.g., blood vessels, macrophages, tumor cells, and necrotic cells. We demonstrate that relational topological features can extract biological insight, including the dominant immune cell phenotype (an important predictor of patient prognosis) and the parameter regimes of a data-generating model. The methods provide a quantitative perspective on the relational analysis of multispecies spatial data, overcome the limits of traditional PH, and are readily computable.

Supplementary Information

The online version contains supplementary material available at 10.1007/s11538-024-01353-6.

Keywords

Persistent homology
Dowker complex
Witness complex
Multiplex imaging
Spatial relations
Mathematics Subject Classification

62R40
55N31
62P10
EPSRCEP/R018472/1 EP/R018472/1, EP/K041096/1, EP/R005125/1, EP/T001968/1 Stolz Bernadette J. Harrington Heather A. EPSRCEP/R018472/1 EP/R018472/1 Byrne Helen M. Yoon Iris H. R. L’Oreal-UNSECO UK and Ireland for Women in Science Rising Talent Programhttp://dx.doi.org/10.13039/501100000288 Royal Society RGF EA 201074, UF150238 Harrington Heather A. Leverhulme Trust and Emerson CollectiveCancer Research UKCTRQQR-2021/100002 Dhesi Jagdeep http://dx.doi.org/10.13039/100014599 Mark Foundation For Cancer Research issue-copyright-statement© The Author(s) under exclusive licence to Society for Mathematical Biology 2024
==== Body
pmcIntroduction

Topological data analysis (TDA) is a field of mathematics that develops topological tools for detecting the shape of data. A prominent tool in TDA, persistent homology (PH) (Ghrist 2008; Edelsbrunner and Harer 2008; Carlsson 2009; Edelsbrunner et al 2000), constructs a nested sequence of topological scaffolds of shapes from data, called a filtration of simplicial complexes. PH examines the evolution of topological features such as connected components (dimension 0) and loops (dimension 1) across the filtration. The filtration is constructed from meaningful aspects of the data at multiple scales such as distances (Vietoris 1927; Edelsbrunner 1993), function values (Chazal et al 2013; Günther et al 2012), and densities (Carlsson and Zomorodian 2007; Botnan and Lesnick 2023; Vipond et al 2021). One possible input to PH is point cloud data, and the output is a persistence diagram, which can be vectorized and integrated with statistics and machine learning methods (Ali et al 2022). PH provides an automatic, robust, and interpretable method for analyzing data arising in many fields of biology and medicine, including cancer biology (Lawson et al 2019; Singh et al 2014; Chittajallu et al 2018; Aukerman et al 2020; Nicolau et al 2011; Bhaskar et al 2021; Nardini et al 2021; Stolz et al 2020; Yang et al 2023), neuroscience (Gardner et al 2022; Curto and Itskov 2008; Dabaghian et al 2012; Giusti et al 2015), and genomics (Masoomy et al 2021; Cámara 2017; Benjamin et al 2023; Emmett et al 2015; Chan et al 2013).

Most existing PH applications are limited to the study of data relating to one species. Advanced data collection techniques now generate multispecies data in which distinct species may interact. Data of this nature are ubiquitous in science, ranging from cancer biology and ecology to geospatial analysis. By studying the spatial relationships among species, we can glean insights that would otherwise be missed in non-spatial analyses. Extracting spatial relationship information from such data, therefore, requires the development of novel analysis techniques. Recently, two topological methods have been proposed to study multispecies data (Bhaskar et al 2022; di Montesano et al 2022). The first approach concatenates topological features from different cell types in cancer images (Bhaskar et al 2022) but does not capture spatial relations between the different cell types. Another method, the chromatic alpha complex (di Montesano et al 2022), encompasses relations among species by constructing a multispecies version of the Delaunay triangulation; its computational implementation and interpretation are still under development.

Here, we present two topological approaches for encoding spatial relations among different species directly at the input level for PH. We implement and showcase these methods on synthetic multispecies data generated by an agent-based model (ABM) of the tumor microenvironment. We present two example pipelines incorporating the two topological encodings of spatial relations. The pipelines can be adjusted by making different choices for the topological approaches and their combination with vectorizations. We show that topological relations encode biological insight by predicting the dominant immune cell phenotype and by clustering the parameter regimes of the data-generating model using the relational topological features.

Mathematically, the multispecies data we consider can be viewed as a labeled point cloud P=⋃i=0mPi that consists of m+1 different species whose spatial distributions may be related to one another. Each point p∈P is in R2. Note that both topological methods can be applied to point clouds in Rn for n≥2. We generated synthetic multispecies spatial data from an ABM that simulates the behavior of different cell types in a tumor microenvironment (Bull and Byrne 2023). The proposed topological methods are built on Dowker complexes (Dowker 1952) and witness complexes (de Silva and Carlsson 2004). These relational PH methods, which we refer to as Dowker PH and multispecies witness PH, use one species, e.g., P0, as the potential vertex set for a simplicial complex and use another species to create a filtration.

Dowker PH (Chowdhury and Mémoli 2018) is based on a Dowker complex (Dowker 1952), which is a simplicial complex that represents relations between two point clouds. Dowker complexes have been used to capture relations in molecular biology (Liu et al 2022), networks (Chowdhury and Mémoli 2018), PDF parsers (Ewing and Robinson 2021), and persistence diagrams (Yoon et al 2023). We propose using Dowker PH (Chowdhury and Mémoli 2018), a natural extension of Dowker complexes, for multispecies data. Dowker PH of the pair (Pi,Pj) creates a filtered Dowker complex on points Pi based on proximity to points in Pj. Dowker PH then examines the topological features of the Dowker complex that evolve as one varies the distances between Pi and Pj. The resulting Dowker persistence diagram is agnostic to the choice of Pi or Pj as the vertex set and can informally be interpreted as capturing shared topological features, i.e., connected components and loops, between Pi and Pj.

While Dowker PH encodes pairwise relations, it does not capture how one species, say P0, relates to all other species in P. To capture differences between all relations among every pair (P0,Pi), we present a second approach called multispecies witness PH, which is inspired by the lazy witness filtration (de Silva and Carlsson 2004). The multispecies witness filtration first creates a Delaunay triangulation (Delaunay 1934) on P0 and creates a filtration based on the number of points in Pi close to simplices in P0. We chose the Delaunay triangulation because of its simplicity and close relationship to the lazy witness filtration (see Theorem 3 in de Silva and Carlsson (2004))1. To encode P0’s different spatial relations to all other subpopulations, we construct m separate filtrations, measure the distance between their topological features, and combine these distances into a topological distance vector, which can then be used as input into classification or machine learning tools.

We apply Dowker PH and multispecies witness PH to two different, yet related biological problems arising from the data-generating model to showcase variations of relational PH. In both application pipelines, we combine the PH methods with vectorizations that emphasize relevant aspects of the encoded relations for the studied problem. In the first problem, we investigate the prediction of the dominant macrophage phenotype, and we use persistence images to vectorize the spatial interactions captured by the Dowker features. Because the dominant macrophage phenotype changes over time, we analyze data simulated at different time points. In the second problem, we classify qualitative regimes of the model which are determined by differences in relative spatial locations using multispecies witness PH. Here, we use distance vectors to detect differences between the cell type-dependent filtrations that we create. The qualitative regimes are classified at the final time point of the simulation (Bull and Byrne 2023), so we analyze the point clouds at the final time point of the simulation.

The paper is organized as follows. In Sect. 2, we describe the synthetic multispecies data and introduce the two questions arising in the study of data from the tumor microenvironment. In Sect. 3, we briefly review the mathematical preliminaries of PH. In Sect. 4, we present the relational PH approaches designed for capturing relations among multiple species: Dowker PH and multispecies witness PH. In Sect. 5, we showcase these methods on a simulated tumor microenvironment and address the biologically motivated questions introduced in Sect. 2. The paper concludes in Sect. 6 where we discuss our results and outline directions for future research.

Multispecies Spatial Data

We introduce the data set we later analyze, which is synthetic point clouds of multiple species in a simulated tumor microenvironment. Next, we state the two associated domain-specific questions that motivate this mathematical study.

Point Clouds Simulated Via Agent-Based Modeling

We study point clouds representing a dynamic and spatially-resolved tumor microenvironment generated by an agent-based model (ABM). ABMs simulate the emergent behavior of a system through the enactment of rules that determine the outcome of interactions between their constituent ‘agents’, here typically individual cells (Bonabeau 2002). They are ideally suited to create multispecies data. We use the ABM presented by Bull and Byrne (2023). See SI Section 1 and Bull and Byrne (2023) for details.

Each simulation produces a point cloud P consisting of five species P=PT∪PS∪PN∪PM∪PV. Each labeled point cloud represents the locations of tumor cells (PT), stromal cells (PS), necrotic cells (PN), macrophages (PM), and blood vessels (PV). The spatial locations of the blood vessels are randomized at the start of each simulation and then held fixed. By contrast, all other cell types are assumed to be motile. Their movement is determined by interactions among the cells and five different diffusible species (oxygen, CSF-1, TGF-β, CXCL12, and EGF). We focus on simulations that arise by varying two key parameters of the model that affect the behavior of macrophages: χcm, the chemotactic sensitivity of macrophages to spatial gradients of one of the chemical species (CSF-1), and c1/2, a parameter regulating the rate at which macrophage extravasate from the blood vessels (Bull and Byrne 2023). We consider 9 different values for each parameter. For each of the 81 possible parameter pairs (χcm,c1/2), we generate up to 20 realizations of the ABM in which the positions of the blood vessels are varied.2 For each simulation, the point clouds are generated at 6 time points (t=250,300,350,400,450,500 hours). We focus on the behaviors of macrophages and tumor cells. Each macrophage has an associated phenotype, Ω∈[0,1], which determines how it interacts with tumor cells. Macrophages with low Ω have high tumor-killing capacity. Those with high Ω assist the migration of tumor cells towards the vasculature, thereby promoting metastasis. We refer to macrophages with phenotype 0≤Ω<0.5 as M1 or anti-tumor macrophages; we refer to those with phenotype 0.5≤Ω≤1 as M2 or pro-tumor macrophages. The cutoff value of 0.5 is motivated by the model described in Bull and Byrne (2023) in which Ω=0.5 describes the tipping point of the macrophage behavior. For Ω<0.5, the model exhibits anti-tumor behavior, whereas for Ω>0.5, the model is skewed in a pro-tumor direction.

Simulations are initially seeded with a small cluster of tumor cells at the center of the domain, with blood vessels clustered around the edge. Blood vessels act as sources of oxygen, which is consumed by both stromal cells and tumor cells. Tumor cells are sources of CSF-1, which diffuses through the domain and acts as a stimulus for the recruitment of macrophages and as a chemoattractant for them. During each simulation, macrophages with phenotype Ω=0 enter the domain at a rate determined by CSF-1 levels at the blood vessels, with higher CSF-1 increasing the rate of macrophage extravasation. As a macrophage migrates through the domain, its phenotype changes in response to local levels of the different chemical species, including TGF-β. (For details, see SI Section 1).

For a given parameter set, at the end of each simulation (t=500 hours), we observe one of three distinct qualitative behaviors:tumor elimination, in which M1 macrophages dominate the simulation and the tumor cells have been eliminated;

tumor equilibrium, in which macrophages are unable to eliminate the tumor cells which form a compact mass, surrounded by macrophages that are predominantly of an M1 phenotype;

tumor escape, in which M2 macrophages enhance tumor cell migration to the vasculature. These simulations are characterized by the formation of perivascular niches in which M2 macrophages, tumor cells, and blood vessels are found in close proximity. Such behavior is associated with metastasis of tumor cells (Arwert et al 2018).

We consider two subsets of data generated by the ABM. The first data subset is generated from 2 realizations of 9×9 parameter combinations of c1/2 and Xcm. The point clouds are generated at 6 time points (t=250,300,350,400,450,500 hours) of the simulation, resulting in 972=6×2×9×9 point clouds. We use the first data subset to predict the dominant macrophage phenotype. All simulations at various time points are analyzed because the dominant macrophage phenotype changes over time.

For the second data subset, we consider up to 20 realizations of 9×9 parameter combinations of c1/2 and Xcm, i.e., a maximum of 1620 point clouds. As noted by Bull and Byrne (2023), limitations on HPC time meant that for some parameter combinations, fewer than 20 realizations were available, giving a total of 1485 point clouds generated at a single ‘endpoint’ time (t=500 hours). For each point cloud, we use the positions of tumor cells, blood vessels, and macrophages (with and without knowledge of macrophage phenotype) as input. For comparison, we also construct simple, i.e., non-topological, descriptor vectors with entries corresponding to the number of tumor cells, the number of macrophages, the number of necrotic cells, the average distance of tumor cells to the nearest blood vessel, the average distance of necrotic cells to the nearest blood vessel, and the average distance of macrophages to the nearest blood vessel. We use the second data subset to classify the parameter regimes that lead to different qualitative behaviors. The qualitative behaviors are labeled at the final time point of the simulation Bull and Byrne (2023). We therefore analyze the point clouds generated at the final time point.

Statement of Biologically Motivated Problems

Fig. 1 Relational PH pipeline and analysis. We use point clouds generated by an ABM as input to two different topological methods for encoding relations: Dowker PH (top row) and multispecies witness PH (bottom row). We vectorize Dowker topological descriptors using persistence images and vectorize witness topological descriptors via distances between them. Finally, we perform supervised binary classification to predict the dominant macrophage phenotype using Dowker features and perform unsupervised clustering to infer the parameter regimes of elimination, equilibrium, and escape using multispecies witness features

We address the following two biologically motivated questions regarding macrophage and tumor behavior: Can relational PH predict the dominant macrophage phenotype from the cell locations without knowledge of the phenotypes of individual macrophages?

Can relational PH identify the parameter regimes of the ABM that lead to different qualitative behaviors: tumor elimination, escape, and equilibrium with macrophages?

These two questions motivated the two pipelines shown in Fig. 1, with the first problem corresponding to the pipeline in the top row and the second to the pipeline introduced in the bottom row. While we present pipelines that were effective in this study, the pipelines may be adjusted by making different choices for the topological tools and vectorizations (see also Sect. 6).

Problem 1: Prediction of Dominant Macrophage Phenotype

We examine whether relational features can predict the dominance of M1 and M2 macrophages (see Fig. 2), which is an important predictor of a cancer patient’s overall survival time (Jayasingam et al 2020). Macrophage phenotype prediction problems may arise in experimental and clinical settings when analyzing imaging data that contains a single macrophage marker or when conventional time- and resource-intensive methods of characterizing macrophage phenotype are not viable (Rostam et al 2017; Huang et al 2019; Jayasingam et al 2020; Misharin et al 2013). We use Dowker PH for this task due to its pairwise encoding of relations. Dowker’s shared topological features allow biological interpretation of which relative cell locations directly influence macrophage phenotype. In particular, since macrophage phenotype is influenced by its spatial interactions with blood vessels and tumor cells, Dowker PH is ideally suited to quantify these interactions. In subsequent analysis, we use persistence images as a vectorization to retain the pairwise relations captured by the Dowker persistence diagrams. We demonstrate that relational PH can identify the dominant macrophage phenotype based on the spatial relations among the constituents.Fig. 2 Problem 1: Prediction of dominant macrophage phenotype. Given a simulated tumor microenvironment, can we predict the dominant macrophage phenotype?

Problem 2: Classification of Parameter Regimes Leading to Different Qualitative Behaviors of the ABM

Secondly, we explore the use of relational PH in understanding the parameter regimes used to generate different simulations, specifically to classify different parameter regimes from the spatial distribution of the different cell types (see Fig. 3). The ABM parameters influence the spatial distributions of different cell types in the tumor microenvironment, leading to different tumor compositions and morphology. The qualitative behaviors that arise from the different parameter combinations of the ABM are shown in Fig. 3 a). The qualitative behaviors were subjectively assigned in the paper (Bull and Byrne 2023). Capturing these differences objectively from the spatial patterns of cells could pave the way for the automated identification of disease stages in microscopy images. Since we are interested in classifying long-term tumor outcomes (escape, elimination, and equilibrium), we consider the ABM output at a single late ‘endpoint’ time (t=500 hours) for varying combinations of parameters c1/2 and Xcm. Multispecies witness PH is ideally suited to this task since it simultaneously takes into account all species in the data set and focuses on their relative spatial locations. When combined with distance computations, it can be used to highlight differences in relative spatial distributions.Fig. 3 Problem 2: Classification of parameter regimes leading to different qualitative behaviors of the ABM. a Parameter values of c1/2 and Xcm varied in the ABM. Depending on the parameter combination, a simulation of the tumor microenvironment results in one of three qualitative behaviors: elimination of the tumor (blue), equilibrium of tumor cells and macrophages (yellow), and escape of the tumor cells towards blood vessels (red). The parameter combinations are colored according to the subjective classification of the qualitative behavior observed in one simulation of the model. b Can we systematically determine the different qualitative behaviors of the ABM from the locations of the different cell types? (Color figure online)

Mathematical Preliminaries

We briefly introduce the standard PH, which can be used to analyze the spatial patterns of point cloud data. For details of PH, see Ghrist (2008); Edelsbrunner and Harer (2008); Carlsson (2009); Edelsbrunner et al (2000).

Persistent Homology

Let P denote a point cloud of data in Rn. Here, P is a point cloud of data in R2 describing the spatial location of biological cells such as cancer cells. The spatial patterns and structure of P can be studied by constructing filtered simplicial complexes, i.e., collections of vertices, edges, triangles, and their higher-order counterparts that can be glued together to approximate topological spaces. We refer to each building block as a simplex. A 0-simplex is a single point in P, a 1-simplex is an edge between two points in P, a 2-simplex is a triangle among three points, and so on. We denote an n-simplex by the collection of n+1 vertices (p0,⋯,pn) that are involved. The standard choice of a filtered simplicial complex is the Vietoris-Rips filtration VRP (Vietoris 1927):

Definition 1

(Vietoris-Rips filtration) Let P be a point cloud and let d be a distance on P. The Vietoris-Rips complex at parameter ε, denoted VRPε, is an abstract simplicial complex that has P as the vertex set and has the n-simplex σ=(p0,⋯,pn) if d(pi,pj)≤ε for all pi,pj∈σ. A Vietoris-Rips filtration VRP∙ is a nested sequence of simplicial complexes VRPε for varying ε.

Fig. 4 An example Vietoris-Rips filtration. a Example Vietoris-Rips complexes VRPε at various ε parameters. The top row shows the point cloud (in black) and ε/2-ball neighborhoods around each point (in green) for varying ε values. The bottom row shows the Vietoris-Rips filtration. For a fixed ε, we draw a 0-simplex for each point in the point cloud. Whenever the green balls intersect, we place a 1-simplex between the two corresponding points. We then fill in any higher-dimensional simplices that arise. b A persistence diagram provides a visual summary of the evolution of connected components (dimension 0, denoted pd0) and loops (dimension 1, denoted pd1). We show an overlay of the dimension-0 persistence diagram pd0(VRP∙) (in circle) and dimension-1 persistence diagram pd1(VRP∙)(in cross). A point on the persistence diagram represents a topological feature. The x-coordinate is the parameter ε at which the feature is born, and the y-coordinate is the parameter at which the feature dies. In pd0, all connected components share the same birth parameter, and the death of a component occurs when two components merge. The red line indicates an infinite death value. There is one point with an infinite death parameter, indicating that the Vietoris-Rips filtration has a single connected component that never vanishes as we increase ε. In pd1, there is a single point far from the diagonal, indicating that there is one significant loop with a small birth parameter and large death parameter. The remaining points can be considered as noise (Color figure online)

The Vietoris-Rips complex VRPε at parameter ε represents the connectivity of P up to proximity ε (see Fig. 4a). The Vietoris-Rips filtration VRP∙ encodes the connectivity of the point cloud at various proximity parameters. PH provides the means to study topological features such as connected components (H0) and cycles (H1) across nested simplicial complexes. Throughout this paper, we fix the field F=Z/2Z.

Definition 2

(Persistent homology) Given a nested sequence of simplicial complexes

the dimension-k persistent homology of X∙ is a collection of F-vector spacesPHk(X∙)=Hk(Xε1;F)→ϕε1Hk(Xε2;F)→ϕε2⋯→ϕεN-1Hk(XεN;F),

with ϕε being the maps induced by ιε.

The evolution of structural features across a filtration is obtained via the structure theorem.

Theorem 1

(Structure Theorem for persistent homology (Carlsson and Zomorodian 2005)) Any dimension-k persistent homology PHk(X∙) obtained from a finite filtered simplicial complex X∙ decomposes uniquely asPHk(X∙)≅⨁iIbi,di,

where each Ibi,di, called an interval module, is a sequence of F-vector spacesIbi,di=0→ϕ0⋯→ϕbi-1F→ϕbi⋯→ϕdi-1F→ϕdi0→ϕdi+1⋯→ϕN-10

with ϕε as identity maps for ε∈[bi,di) and zero otherwise.

Given an interval module Ibi,di, the parameters bi and di are referred to as the birth and death times of Ibi,di. The length (death - birth) is referred to as persistence. The decomposition of PHk(X∙) is often represented using the collection of birth and death times, and they are visualized using a persistence diagram (see Fig. 4b). We denote the dimension-k persistence diagram by pdk(X∙).

Persistence diagrams are stable (Chazal et al 2009). That is, there exist distances on persistence diagrams such that small perturbations of the input P result in small changes in the persistence diagram. Two commonly used distances on persistence diagrams are the Wasserstein distance (Cohen-Steiner et al 2010) and the bottleneck distance (Cohen-Steiner et al 2005), which are described as follows.

Definition 3

Given two points x=(xb,xd) and y=(yb,yd) in a persistence diagram let ‖x-y‖∞=max{|yb-xb|,|yd-xd|}. Given two persistence diagrams pdk(X∙) and pdk(Y∙), the q-Wasserstein distance isdW(pdk(X∙),pdk(Y∙))=infγ:pdk(X∙)→pdk(Y∙)∑x∈pdk(X∙)‖x-γ(x)‖2q1/q,

where ‖·‖2 is the L2 norm,3 and the bottleneck distance isdB(pdk(X∙),pdk(Y∙))=infγ:pdk(X∙)→pdk(Y∙)supx∈pdk(X∙)‖x-γ(x)‖∞.

where γ denotes a bijection between pdk(X∙) and pdk(Y∙). In the case of cardinality mismatches between the persistence diagrams, points on the diagonal are included to generate this bijection.

In Sect. 5.2, we use the 1-Wasserstein distance and the bottleneck distance to construct entries of distance vectors between pairs of persistence diagrams. Both distances capture distinct differences between the two persistence diagrams that are compared. While the 1-Wasserstein distance takes into account how well all points between the two persistence diagrams agree and is additionally influenced by the number of points in a persistence diagram, the bottleneck distance focuses on the worst such agreement.

Vectorization for Machine Learning

Given a persistence diagram pdk(X∙), various techniques can be used to convert it into a vector that is compatible with standard statistics and machine learning (Ali et al 2022; Bubenik 2015). Here, we use persistence images (Adams et al 2017), which summarize the distribution of points on the persistence diagram using a weighted sum of Gaussian distributions centered at each point of the persistence diagram (see Fig. 5). A persistence diagram is first transformed by mapping each point (birth,death) to (birth,death - birth) (Fig. 5a,b). We then place a Gaussian distribution centered at each transformed point and assign a non-negative weighting function (Fig. 5b,c). The function places zero weight for points along the horizontal axis of Fig. 5b. The weighted sum of Gaussians is then discretized to produce an array called a persistence image (Fig. 5d). The persistence image is often flattened into a vector. The resulting vector is influenced by several parameters, including the width, σ of the Gaussian, and discretization size. In this study, we use σ=1 and discretize images to size 20×20, resulting in flattened vectors of dimension 400.Fig. 5 Vectorization of persistence diagrams via persistence image. a An example persistence diagram. b The result of mapping each point (birth, death) in a persistence diagram to (birth, death-birth). c A weighted sum of Gaussians centered at each point of b. d A discretized array of image c. The resulting persistence image is often flattened into a vector. Figure adapted from Adams et al (2017)

Introducing Filtrations for Multispecies Data

While standard PH detects structure in a point cloud, it fails to encode how multiple point clouds are related. We present two extensions of the standard PH pipeline to capture multi-system interactions: Dowker PH (Chowdhury and Mémoli 2018) and multispecies witness PH, a new construction motivated by witness complexes (de Silva and Carlsson 2004).

Dowker Persistent Homology

Let U and V denote two distinct point clouds. In our study, U and V represent different biological cell types, such as tumor cells and macrophages. The structure of U from the viewpoint of V can be studied using a Dowker filtration:Fig. 6 Example Dowker complexes. We present two Dowker complexes built on point clouds U and V for some proximity parameter ε. (Top) Dowker complex with U as the potential vertex set. (Bottom) Dowker complex with V as the potential vertex set. Given a potential vertex set, the ε-neighborhoods of the vertices are shown in green if the neighborhood contains an element of the other point cloud. Otherwise, the neighborhood is shown in red. A vertex with a green neighborhood becomes a 0-simplex in the Dowker complex. We add a 1-simplex between two vertices if their ε-neighborhood intersection contains a vertex from the other point cloud. We add a 2-simplex among three vertices if their ε-neighborhood intersection contains a vertex from the other point cloud (Color figure online)

Definition 4

(Dowker filtration (Dowker 1952; Chowdhury and Mémoli 2018)) Let U and V be point clouds, and let dU,V be the distance function between elements of U and V. A Dowker complex at parameter ε, denoted DU,Vε, is a simplicial complex that has U as the potential vertex set and includes the n-simplex σ=(u0,…,un) if there exists a v∈V such that dU,V(ui,v)≤ε for all ui∈σ. A Dowker filtration DU,V∙ is a nested sequence of Dowker complexes DU,Vε for varying ε.

The Dowker complex DU,Vε at parameter ε captures relations between U and V, where the relations are restricted to points (u, v) whose distance is at most ε. The Dowker complexes DU,Vε (Fig. 6, top) and DV,Uε (Fig. 6, bottom) each have U and V as the potential vertex set. Note that the two Dowker complexes resemble one another even though their vertex sets are distinct. For example, both Dowker complexes have two connected components and three 1-dimensional cycles, i.e., loops. Dowker’s Theorem states that the two Dowker complexes have the same homology groups, i.e., connected components and loops4 (Dowker 1952). In fact, the geometric realizations of the two Dowker complexes are homotopy equivalent (Björner 1996). Dowker complexes can capture shared topological features between two point clouds, as illustrated in Fig. 6. Note that there are instances in which the Dowker complex captures a feature present in U that isn’t present in V, for example, if V is a dense sample of a region containing U (see Sect. 5.1.2 for details).

To study the features of Dowker complexes across a range of parameters ε, we compute the PH of the Dowker filtration DU,V∙. We call the resulting persistence diagram pdk(DU,V∙) the Dowker persistence diagram. The functorial Dowker’s Theorem states that the persistence diagrams of the two filtered Dowker complexes are the same.

Theorem 2

(Functorial Dowker’s Theorem (Chowdhury and Mémoli 2018)) pdk(DU,V∙)=pdk(DV,U∙) for all k.

The Dowker persistence diagram is a collection of birth and death parameters of k-dimensional topological features, i.e., connected components and loops for k=0 and k=1 respectively, in the Dowker filtration. In this study, we utilize dimension-0 and dimension-1 Dowker persistence diagrams. The Dowker persistence diagram can be vectorized via persistence images as described in Sect. 3 and then be used in various statistical and machine learning methods. Given points clouds of size n and m and a distance matrix of size n×m, one can create the filtered Dowker complex as the following. For each column j, sort the column. The collection of rows whose jth entry is at most p represent a simplex that is present at filtration values p or higher. One can then go through the rows of the sorted jth column to create a collection of simplices that are present at various filtration values. Creating the filtered Dowker complex thus has a computational complexity of O(mnlogn).

Multispecies Witness Persistent Homology

Our second approach is motivated by the construction of (lazy) witness filtrations. The (lazy) witness filtration was first introduced by de Silva and Carlsson (de Silva and Carlsson 2004) and has been used to study noisy artificial datasets (Kovacev-Nikolic 2012), primary visual cortex cell populations  (Singh et al 2008), and cancer gene expression data  (Lockwood and Krishnamoorthy 2015). Roughly, the lazy witness filtration is constructed via the following steps: Select a subset of landmark points L from the point cloud P.

Construct a lazy witness filtration where the landmarks L are the vertex set and the full point cloud P serve as witnesses for higher order simplices. Broadly speaking, points in P are witnesses to the simplices on L to which they are closest. De Silva and Carlsson (de Silva and Carlsson 2004) demonstrate that the resulting simplicial complex can be interpreted as an intrinsic Delaunay triangulation (Delaunay 1934) of the point cloud. A filtration of the resulting simplicial complex is typically created by measuring the spatial scale of the simplices, similar to the Dowker filtration as described above.5

For a multispecies point cloud P=∪i=0mPi, Pi∩Pj=∅ for i≠j, we use a similar construction to capture the spatial patterns of different Pi. However, rather than choosing a subset of landmarks L from P, we use one of the point species as landmarks, i.e., L=P0. Motivated by the close relationship of the witness complex and the Delaunay triangulation (de Silva and Carlsson 2004), we create the Delaunay triangulation (Delaunay 1934) D0 on the landmark set, i.e., for 2D point cloud data we create the triangulation of the 2D convex hull of P0. We include all simplices from the Delaunay triangulation and their faces in our simplicial complex, i.e., for 2D data we include all triangles, their edges, and their vertices as the 2-, 1- and 0-simplices of the simplicial complex. The remaining point species Pi for i=1,...,m in P are then used as witnesses for the simplices in the Delaunay triangulation:

Definition 5

(Pi-witness point) Let p∈Pi, L⊂P be a landmark set, and d a distance function on P. We say that p is a Pi-witness for the n-simplex σ=(l0,⋯,ln) if d(p,li)≤d(p,l^) for all l^∈L\{l0,⋯,ln} and i=0,...,n.

We now create species-dependent filtrations W0,i∙ on the landmark set P0 using witness points from Pi:

Definition 6

(Multispecies witness filtration W0,i∙) Let P=∪i=0mPi denote a collection of different point clouds, and let D0 be the Delaunay triangulation of P0. The multispecies witness filtration is a sequence of nested simplicial complexes W0,i∙ on P0 with respect to witness points in Pi where W0,iμ has P0 as its potential vertex set and includes the n-simplex σ=(p0,⋯,pn)∈D0 and all its faces, if μ~σ≤μ with μ~σ=μmax-μσμmax, where μσ is the number of Pi-witnesses of σ and μmax is the maximal number of Pi-witnesses for a simplex in D0.

We illustrate the multispecies witness filtration in an example point cloud in Fig. 7. Note that for the filtration to be well defined, it is necessary to assign filtration values such that all faces of newly added simplices in a particular filtration step are either already present in the filtration or are added in the same filtration step. It is not sufficient to compute only the filtration values of the top-dimensional simplices as their faces can have a higher number of witnesses (see Fig. 7 top row for an example).Fig. 7 Example multispecies witness filtration. Given a point cloud P=P0∪P1∪P2, we illustrate two multispecies witness filtrations on the Delaunay triangulation on P0 using witness points from P1 (depicted as yellow hexagons) and witness points from P2 (depicted as blue stars). Different witness points give rise to different filtrations of the same simplicial complex (Color figure online)

To compare the effect of the different types of witnesses on the filtration, we first compute the dimension-0 and dimension-1 persistence diagrams of the multispecies witness filtrations, denoted pd0(W0,i∙) and pd1(W0,i∙), for i=1,⋯,m, and we compute pairwise distance vectors among the different persistence diagrams. We focus on distances between persistence diagrams. The entries of our distance vectors are given by the pairwise bottleneck distances dB among pd0(W0,i∙), the pairwise bottleneck distances dB among pd1(W0,i∙), the pairwise 1-Wasserstein distances dW among pd0(W0,i∙), and the pairwise 1-Wasserstein distances dW among pd1(W0,i∙) for i=1,⋯,m. Given a point cloud P=⋃i=0mPi with m+1 species, this results in distance vectors with 2×2×m2=2m(m-1) entries. We note that this choice of distance vector sidesteps the additional steps (and parameter choices) of constructing other vectorizations such as persistence images or persistence landscapes. Note that using different witness points leads to differences in the filtrations of the Delaunay triangulation of P0 (see Fig. 7). Taking distances directly between the filtrations allows us to capture relative changes in spatial relations that are more relevant to our interpretation than the topological features of the filtrations themselves.

Results

We demonstrate the utility of relational PH in predicting the macrophage phenotype (Problem 1) and in classifying the qualitative behavior of different parameter regimes of the ABM (Problem 2). For the first task, we find that using Dowker PH features improves the performance of a classifier in comparison to using both non-relational topological and non-topological features. In particular, we find that Dowker PH between tumor cells and blood vessels is the best predictor for the dominant macrophage phenotype. For the second task, we perform classification using the multispecies witness filtration features and recover the previous subjective classification of Fig. 3.

Dowker Persistent Homology Predicts Dominant Macrophage Phenotype

Prediction Pipeline

Fig. 8 Pipeline for macrophage phenotype prediction using Dowker PH. a A point cloud representing a synthetic tumor microenvironment generated by an ABM. b Dowker complexes built on different pairs of cells at fixed proximity parameters. c Dowker persistence diagrams pdk(DU,V∙) for k=0,1. d Vectorization of (dimension-0) Dowker persistence diagrams via persistence images. e An SVM classifier takes a concatenation of flattened persistence image as input and predicts the dominant macrophage phenotype of the synthetic tumor microenvironment

We classify a synthetic tumor microenvironment as either anti-tumor (M1) macrophage dominant or pro-tumor (M2) macrophage dominant based on the spatial distributions of blood vessels, tumor cells, and macrophages. Since the M1 and M2 macrophages exhibit significantly different dynamics in the tumor microenvironment (see Sect. 2.1), we hypothesize that the relations of spatial distributions among the three cell types are good predictors of the dominant macrophage phenotype. Our input data is a point cloud P=PV∪PT∪PM that represents the locations of the three cell types. Note that the input data is blind to the phenotype of individual macrophages.

Given a point cloud P, if 50% or more macrophages are M1 macrophages, then we label the point cloud as M1 dominant. Otherwise, we label the point cloud as M2 dominant. A total of 731 images are labeled 0 (M1 dominant), and 241 images are labeled 1 (M2 dominant).

For each P, we use Dowker PH to capture relations between pairs of constituents of the tumor microenvironment (see Fig. 8a,b). Note that only the spatial information of macrophages, and not the macrophage phenotype, is used to create the topological descriptors. We consider the following three pairs of cell types: macrophages and tumor cells, tumor cells and blood vessels, and macrophages and blood vessels (see Fig. 8b). For each pair, we compute the dimension-0 and dimension-1 Dowker persistence diagrams. Since the Dowker persistence diagram is agnostic to the choice of the vertex set (Theorem 2), the vertex set was chosen to be the cell type with a smaller number of points for faster computation (see Fig. 8c). Each point cloud thus results in six Dowker persistence diagrams: pd0(DM,V∙), pd1(DM,V∙), pd0(DT,V∙), pd1(DT,V∙), pd0(DM,T∙), pd1(DM,T∙).

Each Dowker persistence diagram is vectorized via persistence images to an array of size 20 × 206 (see Fig. 8d) and flattened into vectors of size 400. We concatenate the resulting vectors and train a Support Vector Machine (SVM) for the image classification task. (see Fig. 8e).

We also train SVMs on non-relational topological features obtained from four Vietoris-Rips persistence diagrams: pd0(VRT∙), pd1(VRT∙), pd0(VRM∙), pd1(VRM∙). We further train an SVM on non-topological features such as the count of each cell type and the average distance of each cell type to the nearest blood vessels (see data description in Sect. 2.1).

For each SVM classifier, we optimize the hyperparameters via stratified 5-fold cross-validation, employing the synthetic minority oversampling technique (SMOTE) (Chawla et al 2002) in each fold to address the class imbalance. We train an SVM on 10 different random splits of train and test data and report the 10 classification accuracies on the test data.

Dowker Persistence Diagrams Capture Shared Topological Features

Fig. 9 Dowker persistence diagrams capture spatial relations between cell types. a A synthetic tumor microenvironment in which macrophages and blood vessels surround a compact tumor. The six Dowker persistence diagrams are shown in ai, aii, aiii. b A synthetic tumor microenvironment where the cancer cells and macrophages occupy different spaces from the blood vessels. The cancer cells and macrophages are in close proximity to blood vessels in two regions – the top left and bottom right corners of the tumor mass. The six Dowker persistence diagrams are shown in bi, bii, biii. ai The large birth parameters of points in pd0(DM,T∙) indicate that macrophages and tumor cells are far from one another. aii The small birth parameters of points in pd0(DM,V∙) indicate that macrophages and blood vessels are colocalized. The single cross far from the diagonal in pd1(DM,V∙) indicates that macrophages and blood vessels share a common loop. aiii Both Dowker persistence diagrams are similar to the diagrams in panel ai because the relationship between tumor cells and blood vessels is similar to the relationship between macrophages and tumor cells. bi The small birth parameters of pd0(DM,T∙) indicate that macrophages and tumor cells occupy similar regions. bii The spread of birth parameters for points in pd0(DM,V∙) indicates the variance in the extent to which macrophages and vessels occupy similar spaces. biii The two points in pd0(DT,V∙) far from the diagonal indicate that there are two regions (the top left and bottom right corners of the tumor mass) where the tumor cells and the blood vessels are close to each other

Before we discuss classification accuracy, we present example point clouds and interpretation of Dowker persistence diagrams (see Fig. 9).

Recall that pd0(DU,V∙) summarizes the birth and death of connected components of Dowker complexes as one varies the distances between PU and PV. One can thus consider a dimension-0 Dowker persistence diagram as summarizing shared connected components between two point clouds. There are multiple ways in which a shared connected component arises - PU and PV might occupy a similar region, or PU and PV may occupy different regions but have close contact. In such cases, the shared features will be represented by points in pd0(DU,V∙) with small birth parameters.

For example, consider the relationship between macrophages and tumor cells in Fig. 9a and Fig. 9b. In Fig. 9a, the macrophages are distant from the tumor cells, so the points in pd0(DM,T∙) have large birth times (see Fig. 9ai). On the other hand, in Fig. 9b, the macrophages and tumor cells occupy similar spaces, so the points in pd0(DM,T∙) have small birth times (see Fig. 9bi).

Consider the relationship between tumor cells and blood vessels in Fig. 9b. The tumor cells and blood vessels mostly occupy different spaces. However, the tumor cells and blood vessels are in close proximity in two regions, one on the top left corner and another on the bottom right corner of the tumor mass. The fact that there are two “contact points” between the tumor and blood vessels is reflected by two points in pd0(DT,V∙) that are far from the diagonal (Fig. 9biii). In Fig. 9a, the macrophages and blood vessels occupy very similar regions. Such colocalization between macrophages and blood vessels is reflected by the abundance of points in pd0(DM,V∙) with small birth times (see Fig. 9aii).

A dimension-1 Dowker persistence diagram summarizes the evolution of cycles of Dowker complexes as one varies the distances between U and V. We interpret points in pd1(DU,V∙) that are far from the diagonal line as representing shared loops between two point clouds. For example, the macrophages and blood vessels in Fig. 9a share a loop structure, and such shared loop is reflected by a point in pd1(DM,V∙) that is far from the diagonal (Fig. 9aii). We caution the reader that pd1(DU,V∙) may contain points far from the diagonal line even if PU and PV do not have shared cycles. Such a situation arises, for example, when PU is sampled from a circle while PV is a dense, uniform sample of the background.

SVM on Dowker Features Predicts Dominant Macrophage Phenotype

We first visually inspected whether Dowker persistence diagrams can distinguish M1 and M2 dominant tumor microenvironments. Recall that we computed six Dowker persistence diagrams, which resulted in six 400-dimensional vectors. We refer to the six vectors as Dowker features. We concatenated the six vectors into a 2400-dimensional vector, and we refer to the resulting vector as a concatenated Dowker feature vector. A two-dimensional visualization via Multidimensional Scaling (MDS) (Kruskal 1964) shows decent separation of classes (see Fig. 10b). For comparison, we computed four Vietoris-Rips persistence diagrams from tumor cells and macrophages and vectorized them. We refer to the four vectors as Vietoris-Rips features. We refer to the concatenated vectors as a concatenated Vietoris-Rips feature vector. Visualization of MDS (Fig. 10a, b) indicates that the concatenated Dowker feature vectors may be better predictors of the dominant macrophage phenotype.

We train two SVM classifiers, one that takes the concatenated Dowker feature vector as input and another that takes the concatenated Vietoris-Rips feature vector as input. The SVM trained on concatenated Dowker features has higher accuracy (median accuracy 86.6%) than the SVM trained on the concatenated Vietoris-Rips features (median accuracy 84.2%). Furthermore, the lower quartile of accuracy from the concatenated Dowker features is roughly equal to the upper quartile of accuracy from the concatenated Vietoris-Rips features (∼86%) (see Fig. 10c). Both models outperform an SVM trained on non-topological features such as the number of cells per cell type and average distances of cell types to the nearest blood vessels (see Fig. 10c).

Next, we investigate which cell types were most informative in predicting the dominant macrophage of the synthetic tumor microenvironment. To this end, we train ten additional SVM classifiers on the Dowker and Vietoris-Rips features without any concatenation. We train four classifiers on the four Vietoris-Rips features and six classifiers on the six Dowker features. Among the classifiers trained on Vietoris-Rips features, the model trained on pd1(VRT∙) has the highest median accuracy (83.7%). One possible explanation is that M2 macrophages assist metastasis of tumor cells by guiding them away from the tumor mass towards the blood vessels. During this process, the tumor cells may create many small loops as they navigate away from the tumor mass, creating many non-trivial points in pd1(VRT∙). The persistence diagram pd1(VRT∙) may then reflect the extent to which M2 macrophages assist the spread of cancer cells.

Among the classifiers trained on Dowker features, the model trained on pd0(DT,V∙) has the highest accuracy (median accuracy 88.9%), followed by the model trained on pd0(DT,M∙) (86.0%). It is perhaps surprising that the best predictor of the dominant macrophage phenotype uses the relations between tumor cells and blood vessels and not macrophages. One possible explanation for the improved performance of models using pd0(DT,V∙) in our application is that pd0(DT,V∙) captures colocalization between tumor cells and blood vessels, which can represent the extent to which M2 macrophages have assisted the tumor cells to navigate towards blood vessels for metastasis.

Note that dimension-1 Dowker features involving blood vessels are not particularly good predictors of the dominant macrophage phenotype (see Fig. 10c). The poor performance may be due to the lack of common loops between blood vessels and tumor cells and between the blood vessels and the macrophages.

The analysis at individual time points shows that the performance of the classifiers depends on the time point. In particular, there are time points at which the relational features are more valuable than the Vietoris-Rips features in predicting the dominant macrophage subtype (see SI Fig.4).Fig. 10 Dowker persistent homology features improve the prediction of dominant macrophage subtype. a MDS projection of Vietoris-Rips features. b MDS projection of Dowker features. The two classes have better separation when using Dowker features than the Vietoris-Rips features. c Classification accuracies of SVMs trained on Vietoris-Rips features (green), Dowker features (navy), and non-topological features (red). The box plot summarizes the accuracies from 10 different splits of train and test data. The red box shows the minimum (lower bounding line), median (middle line), and maximum (upper bounding line) accuracy values for SVM trained on non-topological feature vectors. The first two box plots show the accuracies of two SVMs, one trained on concatenated Vietoris-Rips features and another trained on concatenated Dowker features. SVM trained on Dowker features has higher accuracy than SVM trained on Vietoris-Rips features. The remaining box plots show the accuracies of SVMs trained on individual Vietoris-Rips or Dowker features. SVM trained on pd0(DT,V∙) has the highest accuracy among all SVM trained on the Dowker features (Color figure online)

Multispecies Witness Features Identify Qualitative Model Behaviors

To study the different qualitative behaviors of the ABM, we focused on differences between the spatial distributions of the different cell types and applied the multispecies witness PH. We illustrate how we applied multispecies witness PH to the output of our ABM in Fig. 11.Fig. 11 Multispecies witness PH on synthetic data from ABM. The point cloud given by the synthetic data P=PV∪PT∪PN∪PM1∪PM2 consists of blood vessels PV, tumor cells PT, necrotic cells PN, anti-tumor macrophages PM1 and pro-tumor macrophages PM2. We construct a Delaunay triangulation on the blood vessels PV and build cell type dependent filtrations WV,i∙ of the Delaunay triangulation where i∈{T,N,M1,M2}. We obtain one persistence diagram for each cell type-specific filtration

Our point cloud data P=PV∪PT∪PN∪PM1∪PM2 consists of blood vessels PV, tumor cells PT, necrotic cells PN, anti-tumor macrophages PM1 and pro-tumor macrophages PM2. We chose to fix P0=PV and considered two different versions for the witness filtrations: first, we did not distinguish macrophage phenotype, i.e., all macrophages are assumed to be identical and PM=PM1∪PM2. We obtained three different witness filtrations using tumor cells, necrotic cells, and macrophages as witness points. In the second case, we distinguished M1 and M2 macrophage subtypes and constructed four witness filtrations using tumor cells, necrotic cells, M1 macrophages, and M2 and macrophages as witness points. From the persistence diagrams, we computed multispecies PH distance vectors (see Sect. 4.2) to compare the effect of the different types of witnesses on the filtration. The entries of our distance vectors are listed in SI Table 1. We choose a combination of entries corresponding to the 1-Wasserstein distance and the bottleneck distance. This allows us to indirectly include the spatial relations of the witness cell types and the fixed blood vessel point cloud (via the sensitivity of the 1-Wasserstein distance to the number of features in a persistence diagram) while also capturing particularly large changes in individual topological features across the different witness cell type filtrations (via the bottleneck distance). The pairwise distances each contributed 3 entries when all macrophages are considered to be the same cell type and 6 entries when distinguishing between M1 and M2 macrophages for each topological dimension considered. In this way, we converted each point cloud P into a 12- (version 1) and a 24-dimensional (version 2) distance vector, respectively (for a summary, see SI Table 1). We used these distance vectors as input into k-means clustering (specifically, we apply the k-means version implemented in sklearn.cluster which uses Euclidean distance in the clustering algorithm). We summarize the full multispecies witness PH pipeline in Fig. 12. We compared our results to clustering performed on simple (non-topological) descriptor vectors (see data description in Sect. 2.1 for description of simple vectors and see SI Fig. 1 for results).Fig. 12 Multispecies witness PH pipeline. We use the point cloud generated by an ABM as input into our multispecies witness filtrations. We compute persistence diagrams for the multispecies witness filtrations, thereby obtaining topological descriptors of the spatial heterogeneity in the input images. We use the persistence diagrams to compute multispecies PH distance vectors. The entries of these vectors correspond to the pairwise bottleneck and 1-Wasserstein distances between the dimension-0 and dimension-1 persistence diagrams of the cell type-specific filtrations. We use the multispecies PH distance vectors as input into unsupervised classification to identify different qualitative behaviors of the ABM

Multispecies Witness Persistence Classification Disregarding Macrophage Subtype

We recovered the three qualitatively different behaviors of the ABM using the unsupervised multispecies witness PH pipeline without including knowledge about macrophage subtypes. We applied k-means classification for k=3. Figure 13 shows which of the three clusters is dominant amongst the 20 simulations for each parameter combination of Xcm and c1/2 that we consider. The results are consistent with the subjective classification of the qualitative behaviors of the model shown in Fig. 3, i.e., we recovered parameter regimes dominated by tumor elimination, tumor macrophage equilibrium, and escape of the tumor, with the exception of simulations in regimes at the boundaries between the three behaviors. We investigated the consistency of the cluster assignment, which we refer to as cluster purity by dividing the number of simulations attributed to the majority cluster by the total number of simulations for the parameter combination. We found that cluster assignment is less consistent in simulations of the ABM that lie in boundary regions between different qualitative behaviors than in parameter regimes far away from boundaries (see Fig. 13). Our results clearly surpass clustering obtained using simple descriptor vectors of the data (see SI Fig. 1), including information such as the number of cells per cell type and average distances of cell types to the nearest blood vessels with respect to cluster consistency with the subjective clusters shown in Fig. 3.Fig. 13 Classification of multispecies PH distance without distinguishing between macrophage subtypes. a Classification results. b Cluster purity scores. For each parameter combination Xcm and c1/2 of the ABM, we include 20 independent simulations in our analysis. The colors red, blue, and yellow represent the cluster to which the majority of simulations are attributed by the k-means algorithm for k=3. The purity score is computed by taking the ratio between the number of simulations attributed to the majority clusters and the total number of 20 simulations (Color figure online)

Multispecies Witness Persistence Classification Including Macrophage Subtypes

We also recovered the three qualitatively different behaviors of the ABM when information about macrophage subtypes M1 and M2 is included in the construction of our multispecies PH distance vectors. We show our results in Fig. 14. Comparison of the results in Fig. 13 and Fig. 14 shows that the inclusion of the additional information about macrophage subtype alters the prediction of the qualitative behaviors for only one parameter combination (see combination highlighted by square symbol in Fig. 14), Xcm=1 and c1/2=0.1, which is located at the phase transition between elimination and escape. We also computed the purity of clusters for each parameter combination by dividing the number of simulations attributed to the majority cluster by the total number of simulations for the parameter combination. We find that clusters assigned to parameter combinations located at the phase transitions between different parameter regimes are less consistent than those far away from boundaries. Again, our results surpass clustering obtained using simple descriptor vectors of the data, including information such as the number of cells per cell type and average distances of cell types to the nearest blood vessels (see SI Fig. 1 for results).Fig. 14 Classification of multispecies PH distance vectors distinguishing between macrophage subtypes M1 and M2. a Classification results. We highlight parameter combinations for which the classification changed due to the inclusion of macrophage subtype by using a square symbol as opposed to a circle. b Cluster purity scores. For each parameter combination Xcm and c1/2 of the ABM, we include 20 independent simulations in our analysis. The colors red, blue, and yellow represent the cluster to which the majority of simulations are attributed by the k-means algorithm for k=3. The purity score is computed by taking the ratio between the number of simulations attributed to the majority clusters and the total number of 20 simulations (Color figure online)

Multispecies Witness Persistence Classification Determines Phase Transitions as Separate Cluster

Multispecies PH distance vectors further stratified the parameter space of the ABM not only into the three qualitatively different behaviors but also into the regions of phase transitions. When applying k-means classification for k=4, the phase transitions between qualitative behaviors were identified as a separate cluster when including macrophage subtypes M1 and M2 in the analysis (see Fig. 15 b). Interestingly, when ignoring macrophage subtypes (see Fig. 15 a), this effect was less prominent. These results could not be obtained when using k-means classification for k=4 on simple descriptor vectors of the data including information such as the number of cells per cell type and average distances of cell types to the nearest blood vessels (see SI Fig. 1).Fig. 15 Classification of multispecies PH distance vectors for k=4 in k-means clustering. a Classification results without knowledge of macrophage subtypes. b Classification results when distinguishing between macrophage subtypes M1 and M2. The colors red, blue, yellow, and black represent the cluster to which the majority of simulations are attributed by the k-means algorithm for k=4 (Color figure online)

Multispecies Witness Persistence Classification is Robust to Mislabeling of Cell Types

The multispecies witness PH pipeline is robust to noise introduced through relabeling. For each point cloud generated by the ABM, we relabeled up to 50% of the necrotic cells, M1, and M2 macrophages. Relabeled cells were randomly attributed the label of one of the other two cell types. For example, a necrotic cell had a 50% chance of being relabelled as a M1 or M2 macrophage. We focused on these three cell types because their numbers are of comparable magnitude in the ABM output, e.g., relabeling tumor cells or vessels would lead to the addition of a disproportionately high or low number of the other three cell types to the simulation output.Fig. 16 Classification of multispecies PH distance vectors after relabeling 50% of the data. a Classification results. We highlight parameter combinations for which the relabelling changes the classification (compared to multispecies PH distance vectors distinguishing M1 versus M2 macrophages) using a square symbol as opposed to a circle. b Cluster purity scores. For each parameter combination Xcm and c1/2 of the ABM, we include 20 independent simulations in our analysis. The colors red, blue, and yellow represent the cluster to which the majority of simulations are attributed by the k-means algorithm for k=3. The purity score is computed by taking the ratio between the number of simulations attributed to the majority clusters and the total number of 20 simulations

Discussion

With the advancement of data collection techniques, there is a growing need for analysis tools that extract relational information from spatial multispecies data. We presented two novel topological approaches to study structural relations: Dowker PH and multispecies witness PH. Dowker PH produces interpretable persistence images, but its application is limited to pairwise relations. Multispecies witness PH, on the other hand, produces features that are more difficult to interpret, but it captures relations among three or more species. The two topological methods were incorporated into two example pipelines that can be adjusted with different choices of vectorizations. We tested the utility of relational topological features in understanding macrophage and tumor behavior in point cloud simulations of the tumor microenvironment. Our results show that topological relations provide biological insight beyond that contributed by non-relational topological features and non-topological features. Furthermore, our study demonstrates that Dowker PH and multispecies witness PH effectively encode topological relations.

This study contributes novel tools for capturing topological spatial relations that are missed in standard methods. A comparison of topological quantifications of relations to various spatial statistics (Wilson et al 2021), including the recently introduced weighted pair-correlation function (Bull and Byrne 2023), is postponed for future research. We believe that the topological methods, when combined with the computation of cycle representatives, may provide extra insight by identifying the local regions at which relational topological features occur in a point cloud.

Other viable topological methods include multiparameter persistence (Vipond et al 2021) and the chromatic alpha complex (di Montesano et al 2022). Multiparameter persistence creates multifiltrations of a simplicial complex using properties such as distances and density; one could potentially use within-species distance and cross-species distance to create such multifiltrations. However, there are many practical limitations to utilizing multiparameter persistence for multispecies data, such as computability and interpretability. The chromatic alpha complex (di Montesano et al 2022) creates a filtration on a multispecies version of the Delaunay triangulation. While the chromatic alpha complex is a viable and interesting approach for studying multispecies data, it will be computationally more expensive than the multispecies witness filtration since the chromatic alpha complex builds a Delaunay complex on the full point cloud, whereas the multispecies witness filtration builds a Delaunay triangulation on one of the point species. Furthermore, the chromatic alpha complex involves an additional convex optimization step for computing the radius function, whereas the filtrations in the multispecies witness filtration can be read directly from the distance matrix. We believe that the Dowker PH, the multispecies witness filtration, and the chromatic alpha complex characterize different types of spatial patterns, and a comparison of these approaches will be a topic of future study.

One of the limitations of the current work is that Dowker PH can be sensitive to outliers. For example, if P1 and P2 are point clouds that are excluded from one another, a single outlier point of P1 that lives in the neighborhood of P2 will create a shared feature that is encoded by the dimension-0 Dowker persistence diagram. An enhancement of Dowker PH for robustness against outliers, possibly through subsampling (Stolz 2023; Chazal et al 2015) and multiparameter persistence, is postponed for future work.

A further study could investigate the impact of choices in the pipelines, including the choice of the topological methods and the vectorizations. Here, Dowker PH, with its focus on shared topological features, captured the spatial relations underlying changes in the macrophage subtype, while multispecies witness PH successfully highlighted relative differences in spatial distributions across the different cell types. Motivated by the underlying biological problems, we made different choices of vectorizations for the two biological problems. The first choice focused on shared features (persistence images with Dowker PH) and the second method highlighted the differences in spatial patterns (distance vectors with multispecies witness PH). It is important to adapt these combinations to the context of the application. For example, using Dowker PH to address the second biological question and using the multispecies witness PH to address the first biological question led to a decrease in performance (see SI Fig.2 and 3). Using multispecies witness PH for the second question decreased the median accuracy in predicting the dominant macrophage subtype from a median accuracy of 86% to an accuracy of 71% (compare Fig. 10 and SI Fig.2). As the macrophage subtypes are influenced by cellular interactions rather than the coarse differences in spatial distribution among different cell types, multispecies witness features were not as well-suited for this task as Dowker features. Conversely, Dowker features did not lead to a clustering of the three qualitative regimes that resemble the expert annotations (compare Fig. 13 and SI Fig.3. See Fig. 3 for expert annotations). We hypothesize that this difference in performance reflects that the three qualitative regimes in the ABM are more dependent on global differences in spatial distributions among different cell types, rather than shared topological features as captured by the Dowker complex.

A study of the choices in the construction of multispecies witness PH is left for future work. While the multispecies witness PH was based on the lazy witness complex, one could extend the construction to “non-lazy” witness complexes. When applying the multispecies witness PH to simulated tumor microenvironments, we chose the blood vessels as landmarks. Investigation into the influence of the landmark cell type, along with the possibility of using randomly selected points in the domain as landmarks, are subjects of future work.

Recent developments in imaging techniques (Goltsev et al 2018; Giesen et al 2014) and cell identification techniques (Pratapa et al 2021; Yao et al 2019) produce multispecies immunohistochemistry images with detailed information about the locations of various constituents of a tissue microenvironment. In tumor tissue, these constituents may include tumor cells, T-cells, B-cells, stroma, blood vessels, and more. Relational PH can potentially be applied to such multiplex images to automatically extract interpretable quantifications of relations among tumor constituents. Furthermore, the relational topological features can more broadly be applied to many other data sets that carry information on spatial locations of multiple systems. In the future, we envisage the integration of our methods with machine learning tools such as graph neural networks, deep learning, and random forests to achieve increased performance on such relational data and achieve novel insights.

Supplementary Information

Below is the link to the electronic supplementary material.Supplementary file 1 (pdf 1853 KB)

Acknowledgements

BJS, HAH, HMB, and IHRY are members of the Centre for Topological Data Analysis and this research was funded in whole or in part by EPSRC EP/R018472/1. For the purpose of Open Access, the authors have applied a CC BY public copyright licence to any Author Accepted Manuscript (AAM) version arising from this submission. BJS is further supported by the L’Oréal-UNESCO UK and Ireland For Women in Science Rising Talent Programme. HAH gratefully acknowledges funding from EPSRC EP/K041096/1, EP/R005125/1 and EP/T001968/1, the Royal Society RGF\EA\201074 and UF150238, Leverhulme Trust and Emerson Collective. JAB was supported by Cancer Research UK grant number CTRQQR-2021/100002, through the Cancer Research UK Oxford Centre. IHRY gratefully acknowledges funding through the Mark Foundation for Cancer Research.

Data code availability

The data is available in the accompanying materials of Bull and Byrne (2023). All code is available at https://github.com/irishryoon/multiplex_relations. The Dowker PH was computed in Julia using https://github.com/irishryoon/Dowker_persistence. We implemented the multispecies witness PH in Python using the gudhi library (Maria et al 2014) to compute persistence diagrams, as well as bottleneck and Wasserstein distances.

Declarations

Conflict of interest

The authors have no Conflict of interest to declare that are relevant to the content of this article.

1 Note that our construction differs from the lazy witness filtration where the simplices in P0 and their filtration values are determined by their proximity to witnesses. Here, we use the number of witnesses to determine the filtration values on a Delaunay triangulation of P0.

2 These come from 2 sets of 10 realizations in which the threshold value of TGF-β required to change macrophage phenotype was varied (either 0.05 or 0.5). Varying this parameter had no qualitative effect on the simulations, and hence the parameter regimes have here been combined.

3 Note that the Wasserstein distance can be defined for any type of norm, a typical choice is Lp for p∈[1,∞].

4 Note that while the homology groups of the above constructions are isomorphic, their connectivity, as measured for example by Q-analysis (Atkin 1972), may differ.

5 The Dowker filtration can be viewed as a special case of the lazy witness filtration. In the general formulation of the lazy witness filtration (de Silva and Carlsson 2004), the distance to the ν-th closest witness is added to the proximity filtration scale ϵ. Given point clouds P and Q, a modified Dowker filtration in which all vertices have birth time 0 is a witness filtration with P as landmarks, Q as witnesses, and ν=0.

6 The classification accuracy is fairly robust to the size of the persistence image. Such robustness is known in the literature Adams et al (2017).

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Bernadette J. Stolz and Jagdeep Dhesi have contributed equally to this work.
==== Refs
References

Adams H Emerson T Kirby M Persistence images: a stable vector representation of persistent homology J Mach Learn Res 2017 18 8 1 35
Adams H, Emerson T, Kirby M et al (2017) Persistence images: a stable vector representation of persistent homology. J Mach Learn Res 18(8):1–35
Ali D, Asaad A, Jimenez MJ, et al (2022) A survey of vectorization methods in topological data analysis. arXiv preprint arXiv:2212.09703
Arwert EN Harney AS Entenberg D A unidirectional transition from migratory to perivascular macrophage is required for tumor cell intravasation Cell Rep 2018 23 1239 1248 29719241
Arwert EN, Harney AS, Entenberg D et al (2018) A unidirectional transition from migratory to perivascular macrophage is required for tumor cell intravasation. Cell Rep 23:1239–124829719241
Atkin RH From cohomology in physics to q-connectivity in social science Int J Man Mach Stud 1972 4 2 139 167
Atkin RH (1972) From cohomology in physics to q-connectivity in social science. Int J Man Mach Stud 4(2):139–167
Aukerman A, Carrière M, Chen C, et al (2020) Persistent homology based characterization of the breast cancer immune microenvironment: a feasibility study. In: International Symposium on Computational Geometry
Benjamin K Mukta L Moryoussef G Homology of homologous knotted proteins J R Soc Interface 2023 20 635
Benjamin K, Mukta L, Moryoussef G et al (2023) Homology of homologous knotted proteins. J R Soc Interface 20:635
Bhaskar D Zhang W Wong I Topological data analysis of collective and individual epithelial cells using persistent homology of loops Soft Matter 2021 17 210
Bhaskar D, Zhang W, Wong I (2021) Topological data analysis of collective and individual epithelial cells using persistent homology of loops. Soft Matter 17:210. 10.1039/D1SM00072A
Bhaskar D Zhang WY Volkening A Topological data analysis of spatial patterning in heterogeneous cell populations: I Clust Sort Vary Cell Cell Adhesion 2022 6 93
Bhaskar D, Zhang WY, Volkening A et al (2022) Topological data analysis of spatial patterning in heterogeneous cell populations: I. Clust Sort Vary Cell Cell Adhesion 6:93
Björner A Topological methods 1996 Cambridge MIT Press 1819 1872
Björner A (1996) Topological methods. MIT Press, Cambridge, pp 1819–1872
Bonabeau E Agent-based modeling: methods and techniques for simulating human systems Proc Natl Acad Sci USA 2002 99 Suppl 3 7280 7 12011407
Bonabeau E (2002) Agent-based modeling: methods and techniques for simulating human systems. Proc Natl Acad Sci USA 99(Suppl 3):7280–7. 10.1073/pnas.08208089912011407
Botnan MB, Lesnick M (2023) An introduction to multiparameter persistence. 10.48550/arXiv.2203.14289
Bubenik P Statistical topological data analysis using persistence landscapes J Mach Learn Res 2015 16 77 102
Bubenik P (2015) Statistical topological data analysis using persistence landscapes. J Mach Learn Res 16:77–102
Bull JA Byrne HM Quantification of spatial and phenotypic heterogeneity in an agent-based model of tumour-macrophage interactions PLoS Comput Biol 2023 3 19 e1010994
Bull JA, Byrne HM (2023) Quantification of spatial and phenotypic heterogeneity in an agent-based model of tumour-macrophage interactions. PLoS Comput Biol 3(19):e1010994
Cámara PG Topological methods for genomics: present and future directions Curr Opini Syst Biol 2017 1 95 101
Cámara PG (2017) Topological methods for genomics: present and future directions. Curr Opini Syst Biol 1:95–101
Carlsson GE Topology and data Bull Am Math Soc 2009 46 255 308
Carlsson GE (2009) Topology and data. Bull Am Math Soc 46:255–308
Carlsson G Zomorodian A Computing persistent homology Dis Comput Geom 2005 33 2 249 274
Carlsson G, Zomorodian A (2005) Computing persistent homology. Dis Comput Geom 33(2):249–274
Carlsson G Zomorodian A The theory of multidimensional persistence Dis Comput Geom 2007 42 71 93
Carlsson G, Zomorodian A (2007) The theory of multidimensional persistence. Dis Comput Geom 42:71–93. 10.1007/s00454-009-9176-0
Chan J, Carlsson G, Rabadan R (2013) Topology of viral evolution. In: Proceedings of the National Academy of Sciences of the United States of America 110. 10.1073/pnas.1313480110
Chawla NV Bowyer KW Hall LO Smote: synthetic minority over-sampling technique J Artif Intell Res 2002 16 321 357
Chawla NV, Bowyer KW, Hall LO et al (2002) Smote: synthetic minority over-sampling technique. J Artif Intell Res 16:321–357
Chazal F Guibas LJ Oudot SY Persistence-based clustering in riemannian manifolds J ACM 2013 60 6 63
Chazal F, Guibas LJ, Oudot SY et al (2013) Persistence-based clustering in riemannian manifolds. J ACM 60(6):63. 10.1145/2535927
Chazal F, Cohen-Steiner D, Glisse M, et al (2009) Proximity of persistence modules and their diagrams. In: Proceedings of the Twenty-Fifth Annual Symposium on Computational Geometry. Association for Computing Machinery, New York, NY, USA, SCG ’09, pp 237-246, 10.1145/1542362.1542407
Chazal F, Fasy B, Lecci F, et al (2015) Subsampling methods for persistent homology. In: Proceedings of the 32nd International Conference on Machine Learning, vol 37
Chittajallu D, Siekierski N, Lee S, et al (2018) Vectorized persistent homology representations for characterizing glandular architecture in histology images. pp 232–235, 10.1109/ISBI.2018.8363562
Chowdhury S Mémoli F A functorial dowker theorem and persistent homology of asymmetric networks J Appl Comput Topol 2018 2 1 115 175
Chowdhury S, Mémoli F (2018) A functorial dowker theorem and persistent homology of asymmetric networks. J Appl Comput Topol 2(1):115–175. 10.1007/s41468-018-0020-6
Cohen-Steiner D Edelsbrunner H Harer J Lipschitz functions have lp-stable persistence Found Comput Math 2010 10 127 139
Cohen-Steiner D, Edelsbrunner H, Harer J et al (2010) Lipschitz functions have lp-stable persistence. Found Comput Math 10:127–139. 10.1007/s10208-010-9060-6
Cohen-Steiner D, Edelsbrunner H, Harer J (2005) Stability of persistence diagrams. pp 263–271, 10.1007/s00454-006-1276-5
Curto C Itskov V Cell groups reveal structure of stimulus space PLoS Comput Biol 2008 4 e1000205 18974826
Curto C, Itskov V (2008) Cell groups reveal structure of stimulus space. PLoS Comput Biol 4:e1000205. 10.1371/journal.pcbi.100020518974826
Dabaghian Y Mémoli F Frank L A topological paradigm for hippocampal spatial map formation using persistent homology PLoS Comput Biol 2012 8 8 1 14 22629235
Dabaghian Y, Mémoli F, Frank L et al (2012) A topological paradigm for hippocampal spatial map formation using persistent homology. PLoS Comput Biol 8(8):1–14. 10.1371/journal.pcbi.100258122629235
de Silva V, Carlsson G (2004) Topological estimation using witness complexes. In: Gross M, Pfister H, Alexa M, et al (eds) SPBG’04 Symposium on Point-Based Graphics 2004. In: The Eurographics Association, pp 157–166
Delaunay B Sur la sphère vide. A la mémoire de Georges Voronoï Bulletin de l’Académie des Sciences de l’URSS Classe des sciences mathématiques et naturelles 1934 6 793 800
Delaunay B (1934) Sur la sphère vide. A la mémoire de Georges Voronoï. Bulletin de l’Académie des Sciences de l’URSS Classe des sciences mathématiques et naturelles 6:793–800
di Montesano SC, Draganov O, Edelsbrunner H, et al (2022) Persistent homology of chromatic alpha complexes. 10.48550/ARXIV.2212.03128
Dowker CH Homology groups of relations Ann Math 1952 56 84 95
Dowker CH (1952) Homology groups of relations. Ann Math 56:84–95
Edelsbrunner H (1993) The union of balls and its dual shape. In: Proceedings of the ninth annual symposium on Computational geometry, pp 218–231
Edelsbrunner H Harer J Persistent homology-a survey Dis Comput Geom 2008 5 453
Edelsbrunner H, Harer J (2008) Persistent homology-a survey. Dis Comput Geom 5:453. 10.1090/conm/453/08802
Edelsbrunner H, Letscher D, Zomorodian A (2000) Topological persistence and simplification. In: Proceedings 41st Annual Symposium on Foundations of Computer Science, pp 454–463, 10.1109/SFCS.2000.892133
Emmett K, Schweinhart B, Rabadan R (2015) Multiscale topology of chromatin folding. arXiv preprint arXiv:1511.01426
Ewing KP, Robinson M (2021) Metric comparisons of relations. 10.48550/ARXIV.2105.01690
Gardner RJ Hermansen E Pachitariu M Toroidal topology of population activity in grid cells Nature 2022 602 123 128 35022611
Gardner RJ, Hermansen E, Pachitariu M et al (2022) Toroidal topology of population activity in grid cells. Nature 602:123–12835022611
Ghrist R Barcodes: the persistent topology of data Bull Am Math Soc 2008 45 560
Ghrist R (2008) Barcodes: the persistent topology of data. Bull Am Math Soc 45:560. 10.1090/S0273-0979-07-01191-3
Giesen C Wang H Schapiro D Highly multiplexed imaging of tumor tissues with subcellular resolution by mass cytometry Nat Methods 2014
Giesen C, Wang H, Schapiro D et al (2014) Highly multiplexed imaging of tumor tissues with subcellular resolution by mass cytometry. Nat Methods. 10.1038/nmeth.2869
Giusti C, Pastalkova E, Curto C et al (2015) Clique topology reveals intrinsic geometric structure in neural correlations. In: Proceedings of the National Academy of Sciences of the United States of America, vol 112. 10.1073/pnas.1506407112
Goltsev Y Samusik N Kennedy-Darling J Deep profiling of mouse splenic architecture with codex multiplexed imaging Cell 2018
Goltsev Y, Samusik N, Kennedy-Darling J et al (2018) Deep profiling of mouse splenic architecture with codex multiplexed imaging. Cell. 10.1101/203166
Günther D Reininghaus J Wagner H Efficient computation of 3d morse-smale complexes and persistent homology using discrete morse theory Vis Comput 2012 28 959 969
Günther D, Reininghaus J, Wagner H et al (2012) Efficient computation of 3d morse-smale complexes and persistent homology using discrete morse theory. Vis Comput 28:959–969
Huang YK Wang M Sun Y Macrophage spatial heterogeneity in gastric cancer defined by multiplex immunohistochemistry Nat Commun 2019 10 1 3928 31477692
Huang YK, Wang M, Sun Y et al (2019) Macrophage spatial heterogeneity in gastric cancer defined by multiplex immunohistochemistry. Nat Commun 10(1):392831477692
Jayasingam SD Citartan M Thang TH Evaluating the polarization of tumor-associated macrophages into m1 and m2 phenotypes in human cancer tissue: Technicalities and challenges in routine clinical practice Front Oncol 2020 9 560
Jayasingam SD, Citartan M, Thang TH et al (2020) Evaluating the polarization of tumor-associated macrophages into m1 and m2 phenotypes in human cancer tissue: Technicalities and challenges in routine clinical practice. Front Oncol 9:560. 10.3389/fonc.2019.01512
Kovacev-Nikolic V (2012) Persistent homology in analysis of point-cloud data. Master’s thesis, University of Alberta, https://era.library.ualberta.ca/files/cv43nx33b/Kovacev-Nikolic_Violeta_Fall2012.pdf
Kruskal JB Multidimensional scaling by optimizing goodness of fit to a nonmetric hypothesis Psychometrika 1964 29 1 27
Kruskal JB (1964) Multidimensional scaling by optimizing goodness of fit to a nonmetric hypothesis. Psychometrika 29:1–27
Lawson P Sholl A Brown J Persistent homology for the quantitative evaluation of architectural features in prostate cancer histology Sci Rep 2019 9 1139 30718811
Lawson P, Sholl A, Brown J et al (2019) Persistent homology for the quantitative evaluation of architectural features in prostate cancer histology. Sci Rep 9:1139. 10.1038/s41598-018-36798-y30718811
Liu X Feng H Wu J Dowker complex based machine learning (dcml) models for protein-ligand binding affinity prediction PLoS Comput Biol 2022 18 4 960
Liu X, Feng H, Wu J et al (2022) Dowker complex based machine learning (dcml) models for protein-ligand binding affinity prediction. PLoS Comput Biol 18(4):960
Lockwood S, Krishnamoorthy B (2015) Topological features in cancer gene expression data. In: Pacific Symposium on Biocomputing, pp 108–119
Maria C, Boissonnat JD, Glisse M, et al (2014) The gudhi library: Simplicial complexes and persistent homology. In: International congress on mathematical software, Springer, pp 167–174, software available at https://gudhi.inria.fr (software retrieved in 2020)
Masoomy H Askari B Tajik S Topological analysis of interaction patterns in cancer-specific gene regulatory network: persistent homology approach Sci Rep 2021 11 630 33436651
Masoomy H, Askari B, Tajik S et al (2021) Topological analysis of interaction patterns in cancer-specific gene regulatory network: persistent homology approach. Sci Rep 11:63033436651
Misharin AV Morales-Nebreda L Mutlu GM Flow cytometric analysis of macrophages and dendritic cell subsets in the mouse lung Am J Respir Cell Mol Biol 2013 49 4 503 510 23672262
Misharin AV, Morales-Nebreda L, Mutlu GM et al (2013) Flow cytometric analysis of macrophages and dendritic cell subsets in the mouse lung. Am J Respir Cell Mol Biol 49(4):503–51023672262
Nardini JT Stolz BJ Flores KB Topological data analysis distinguishes parameter regimes in the Anderson-Chaplain model of angiogenesis PLoS Comput Biol 2021 17 6 e1009094 34181657
Nardini JT, Stolz BJ, Flores KB et al (2021) Topological data analysis distinguishes parameter regimes in the Anderson-Chaplain model of angiogenesis. PLoS Comput Biol 17(6):e100909434181657
Nicolau M Levine AJ Carlsson GE Topology based data analysis identifies a subgroup of breast cancers with a unique mutational profile and excellent survival Proc Natl Acad Sci 2011 108 7265 7270 21482760
Nicolau M, Levine AJ, Carlsson GE (2011) Topology based data analysis identifies a subgroup of breast cancers with a unique mutational profile and excellent survival. Proc Natl Acad Sci 108:7265–727021482760
Pratapa A Doron M Caicedo J Image-based cell phenotyping with deep learning Curr Opin Chem Biol 2021 65 9 17 34023800
Pratapa A, Doron M, Caicedo J (2021) Image-based cell phenotyping with deep learning. Curr Opin Chem Biol 65:9–17. 10.1016/j.cbpa.2021.04.00134023800
Rostam H Reynolds P Alexander M Image based machine learning for identification of macrophage subsets Sci Rep 2017 7 21 28154422
Rostam H, Reynolds P, Alexander M et al (2017) Image based machine learning for identification of macrophage subsets. Sci Rep 7:21. 10.1038/s41598-017-03780-z28154422
Singh G Mémoli F Ishkhanov T Topological analysis of population activity in visual cortex J Vis 2008 8 11 1 18 19350710
Singh G, Mémoli F, Ishkhanov T et al (2008) Topological analysis of population activity in visual cortex. J Vis 8(11):1–1819350710
Singh N, Couture H, Marron J, et al (2014) Topological descriptors of histology images. pp 231–239, 10.1007/978-3-319-10581-9_29
Stolz BJ Outlier-robust subsampling techniques for persistent homology J Mach Learn Res 2023 24 1 35
Stolz BJ (2023) Outlier-robust subsampling techniques for persistent homology. J Mach Learn Res 24:1–35
Stolz BJ Kaeppler J Markelc B Multiscale topology characterises dynamic tumour vascular networks Sci Adv 2020 8 23 eabms2456
Stolz BJ, Kaeppler J, Markelc B et al (2020) Multiscale topology characterises dynamic tumour vascular networks. Sci Adv 8(23):eabms2456
Vietoris L Über den höheren Zusammenhang kompakter Räume und eine Klasse von zusammenhangstreuen Abbildungen Math Ann 1927 97 454 472
Vietoris L (1927) Über den höheren Zusammenhang kompakter Räume und eine Klasse von zusammenhangstreuen Abbildungen. Math Ann 97:454–472
Vipond O Bull JA Macklin PS Multiparameter persistent homology landscapes identify immune cell spatial patterns in tumors Proc Natl Acad Sci 2021 118 41 e2102166118 34625491
Vipond O, Bull JA, Macklin PS et al (2021) Multiparameter persistent homology landscapes identify immune cell spatial patterns in tumors. Proc Natl Acad Sci 118(41):e210216611834625491
Wilson CM Ospina OE Townsend MK Challenges and opportunities in the statistical analysis of multiplex immunofluorescence data Cancers 2021 13 12 63
Wilson CM, Ospina OE, Townsend MK et al (2021) Challenges and opportunities in the statistical analysis of multiplex immunofluorescence data. Cancers 13(12):63. 10.3390/cancers13123031
Yang J, Fang H, Dhesi J, et al (2023) Topological classification of tumour-immune interactions and dynamics. 10.48550/arXiv.2308.05294
Yao K Rochman N Sun S Cell type classification and unsupervised morphological phenotyping from low-resolution images using deep learning Sci Rep 2019 9 1 13 30626917
Yao K, Rochman N, Sun S (2019) Cell type classification and unsupervised morphological phenotyping from low-resolution images using deep learning. Sci Rep 9:1–13. 10.1038/s41598-019-50010-930626917
Yoon HR Ghrist R Giusti C Persistent extension and analogous bars: data-induced relations between persistence barcodes J Appl Comput Topol 2023
Yoon HR, Ghrist R, Giusti C (2023) Persistent extension and analogous bars: data-induced relations between persistence barcodes. J Appl Comput Topol. 10.1007/s41468-023-00115-y
