稻草人新闻RSS 聚合阅读

← 返回 🔬 科学 & 医学

A cluster analysis of neonates using clinical signs of possible serious bacterial infection at hospital admission in Kenya: A retrospective multicentre cohort study

PLOS One Timothy Tuti 1 天前 journals.plos.org

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

Neonatal sepsis remains a major cause of mortality in Sub-Saharan Africa (SSA). Despite presenting with considerable clinical heterogeneity, suspected cases are managed uniformly with broad-spectrum antibiotics. Typical data-driven approaches developed in high-resource settings to identify clinically meaningful phenotypes and support management of neonatal sepsis have limited generalisability to many SSA public hospital settings, due to inclusion of variables that are largely unavailable at admission. This study’s objective was to identify sepsis clusters using signs of possible Serious Bacterial Infection (pSBI) readily available at the time of admission, and to assess the clusters performance in predicting mortality.

We conducted unsupervised model-based cluster analysis using Latent Class Analysis based on pSBI data collected at admission. All neonates <28 days old admitted to 21 Kenyan hospitals between January 2022 and December 2024 with ≥1 pSBI sign/symptom at admission were eligible for inclusion. We further explored the external validity of this clustering approach on new patient populations, and assessed the ability of the identified clusters to accurately predict in-hospital mortality compared to the World Health Organization neonatal sepsis severity classification guidelines.

Five clusters of minimal, low, moderate, substantial and critical mortality risk were identified from development dataset with 33094 patients from eight hospitals. The models had an accuracy, positive predictive value and specificity of at least 83.16% (82.72% to 83.62%), 81.02% (80.58% to 81.45%) and 86.91% (86.61% to 87.23%) respectively in predicting cluster membership of 23704 patients in the external validation dataset admitted to thirteen different hospitals. From an internal-external cross-validation approach of the in-hospital mortality risk, the model-based clustering approach had discrimination (AUROC) of 0.867 (0.863 to 0.871) and calibration intercept and slope of −0.004 (−0.031 to 0.023) and 0.996 (0.979 to 1.014) respectively, outperforming the WHO sepsis severity classification whose discrimination was 0.721 (0.715 to 0.727) and calibration intercept and slope being 0.018 (−0.005 to 0.041) and 1.015 (0.986 to 1.043) respectively.

The identified clusters can complement clinicians’ judgement in assessing risk among neonates with sepsis at admission. Future work evaluating the utility of these clusters and potential differences in treatment response across clusters are therefore recommended to help strengthen the case for more targeted, risk-based neonatal sepsis management.

Citation: Tuti T, Muema T, English M, on behalf of The Clinical Information Network Group, Aluvaala J (2026) A cluster analysis of neonates using clinical signs of possible serious bacterial infection at hospital admission in Kenya: A retrospective multicentre cohort study. PLoS One 21(9): e0359212. https://doi.org/10.1371/journal.pone.0359212

Editor: Khin Thet Wai, Freelance Consultant, Myanmar, MYANMAR

Received: January 19, 2026; Accepted: September 10, 2026; Published: September 24, 2026

Copyright: © 2026 Tuti et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Data Availability: The datasets generated and/or analysed during the current study are not publicly available because of ethical restrictions on sharing the de-identified data set. The dataset comprises of patient-level clinical data collected from neonates included in this multicentre study. The primary data are owned by third-party institutions (i.e. the participating hospitals and county governments), with restrictions imposed by the KEMRI scientific and ethics review unit (SERU) and the Office of Data Protection Kenya (ODPK). The investigators obtained the necessary permissions to access and use these data for the purposes of this research. However, access to the underlying data by other researchers is subject to the relevant Kenyan data governance and KEMRI institutional requirements. The investigators did not receive any special privileges to access the data that would not be available to other researchers. Requests for access to primary quantitative research data by researchers other than the investigators should be submitted as a first step to the KEMRI-Wellcome Trust Research Programme’s Data Governance Committee by emailing dgc@kemri-wellcome.org. Based on the nature of the data requested, the Data Governance Committee will advise on the requirements for access and whether additional approval is required from SERU committee and the relevant facility or county government.

Funding: Wellcome Trust Early Career Research Fellowship (#227562/Z/23/Z) awarded to TT and a Wellcome Trust Senior Fellowship (#097170) awarded to ME. Additional support was provided by a Wellcome Trust core grant awarded to the KEMRI-Wellcome Trust Research Programme (#092654). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Competing interests: The authors have declared that no competing interests exist.

Abbreviations: CIN, clinical information network; EHRs, electronic health records; HICs, high income countries; HCWs, health care workers; KPA, Kenya paediatric association; LCA, latent class analysis; LMICs, low and middle-income countries; MoH, ministry of health; NBU, newborn unit; pSBI, possible serious bacterial infection (pSBI); SERU, KEMRI’s scientific and ethics review unit; SSA, sub-Saharan Africa; WHO, World Health Organization

The most vulnerable time for a child’s survival are the first 28 days of life (the neonatal period), with 2.3 million children dying during this period globally, with a child from Sub-Saharan Africa (SSA) being 10 times more likely to die than a child from a high-income country [1]. Sepsis, which is a dysregulated immune response to infection that leads to acute organ dysfunction, is a leading contributor to burden of disease in neonates in SSA, both as primary cause of death and as a frequent contributor [2,3], and carries a high risk of death even when care is provided promptly [4].

In newborns, sepsis occurs in two forms: early-onset neonatal sepsis, which occurs within the first 72 hours of life, often linked to complications during birth, and late-onset neonatal sepsis, which develops between after 72 hours of life [5]. These infections are more common in premature or low birth weight babies, and their rates are higher in regions with limited or poor quality healthcare [6]. One-sixth of neonates treated for sepsis face significant challenges including functional limitation, cognitive impairment, and mental health disorders after recovery with 40% of sepsis patients requiring rehospitalisation within 90 days of discharge [3].

Many of these neonatal sepsis deaths in SSA are preventable with early detection and proper treatment. To achieve this healthcare providers rely on a high index of suspicion of sepsis (in the absence of laboratory confirmation) based on signs of possible Serious Bacterial Infection (pSBI) like poor feeding, lethargy, temperature instability, or respiratory distress [7,8]. While laboratory tests such as blood cultures to identify pathogens, complete blood counts to assess the immune response, and urine cultures for accurate sepsis diagnosis are crucial for effective treatment [9], in many health facilities in SSA, these diagnostic tests are unavailable at the time of admission [10].

There is no consensus on the clinical definition of neonatal sepsis where microbiological confirmations are unavailable [11,12]. It is a heterogeneous syndrome representing a diverse group of patients, ranging from neonates with minor infections that will resolve quickly to those with severe life-threatening sepsis. Using broad-spectrum, empiric antibiotic treatment for all suspected cases of neonatal sepsis may not be ideal, as it can lead to unnecessary prolonged antibiotic use in non-infected infants, potentially causing harm [4]. Tailoring treatment based on specific diagnoses and risk of mortality could improve outcomes and reduce overuse of antibiotics [13].

The challenge lies in developing alternative effective strategies for early diagnosis and treatment of neonatal sepsis in environments such as SSA where microbiological investigation or laboratory testing is either unavailable or inaccessible [10]. Different combinations of pSBI signs may naturally cluster into previously undescribed subsets or phenotypes that may have different risks for the outcome and may respond differently to treatments. Identification of distinct clinical phenotypes may allow more precise therapy and improve care [4]. However, many guidelines and hospital protocols continue to recommend a one-size-fits-all approach; recommending the same initial inpatient treatment and follow up [4].

Efforts to address the variability in neonatal sepsis have included the WHO’s triaging guidelines for pSBI [4,14] as illustrated in S1 Table. Other approaches have included use of (1) cluster analysis of neonates based on clinical and laboratory data from electronic health records (EHRs) collected within the first six hours of the patients’ hospital stay, combined with serum biomarkers [15], (2) latent class analysis (LCA) to identify molecular phenotypes in sepsis patients using both clinical and protein biomarker data [16], and even used self-organizing maps to identify sepsis groups susceptible to multiple organ dysfunction syndrome based on their age and sequential organ failure assessment scores [17].

However, these approaches generally rely on laboratory, biomarker, or other variables that are not routinely available at admission in many SSA health facilities, limiting their applicability in these settings [4,10,18]. Moreover, there is limited evidence on whether clinically derived phenotypes based on routinely collected admission data are reproducible across hospitals or can be translated into simple classifiers for assigning new patients to predefined risk groups [19]. Addressing these gaps could provide a more context-appropriate approach to neonatal risk stratification that is both externally valid and feasible for routine use in resource-constrained settings [18,19].

Using routinely collected data that is available in typical SSA hospital settings, the aim of this study is to identify clusters of neonates at time of admission that can be candidates for appropriate interventions to maximise patient outcomes given the resource constraints. More specifically, the objectives of this study were:

The reporting of this study follows the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guidelines, which provide a set of recommendations for transparent and comprehensive reporting of observational studies using cohort, case-control, or cross-sectional designs [20]. The reporting also follows the Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD) guidelines, which is a set of recommendations for the reporting of studies developing, validating, or updating prediction models for prognostic purposes [21].

The Scientific and Ethics Review Unit of the Kenya Medical Research Institute (KEMRI) approved the collection of a copy of the de-identified anonymised patient-level data that provides the basis for this study as part of the Clinical Information Network (CIN) which runs in partnership with the Ministry of Health and the participating hospitals (SERU #3459). We elaborate on the CIN in the study design and settings section. Individual consent for retrospective access to the de-identified anonymised patient records was not required and was waived by KEMRI with the authors having no access to the information that could identify the patients.

This is a retrospective cohort analysis utilizing data from the Clinical Information Network (CIN). The CIN collects standardised routine admission and discharge data from newborn units (NBU) across 21 public county hospitals in 14 out of the 47 counties in Kenya with a detailed description of the methods of data collection and management is provided elsewhere [22]. In brief, neonatal admission data are recorded using paper forms such as Neonatal Admission Records (NAR), treatment sheets, continuous monitoring charts, supplementary forms etc. These documents provide a structured checklist for typically junior clinicians [23], covering nine essential domains: demographics, admission details, maternal history, presenting complaints, cardinal signs, physical examinations, nursing monitoring, discharge status and supportive care. The hospital forms are designed to record the clinical data outlined in the “Predictors” subsection of the Methods in accordance with WHO and national clinical guidelines, thereby supporting evidence-based clinical decision-making [14,24]. Consequently, the recorded variables correspond directly to the relevant WHO criteria without requiring additional mapping. This approach to clinical form design and data collection is widely used across East and Southern Africa [25]. Each hospital has a clerk who then extracts data from these typical hospital forms into a Research Electronic Data Capture (REDCap) database immediately after patient death or discharge [26].

The study participants included inborn neonates aged 0–28 days admitted into newborn units between 01st January 2022 and 31st December 2024 with at least one pSBI sign/symptom at admission, regardless of whether they had a neonatal sepsis admission diagnosis, or were prescribed first-line intravenous antibiotics (e.g., Crystalline Penicillin and Gentamicin) at admission [27]. Data from 8 hospitals (n = 33094) was used as model derivation dataset, with the external geographical validation dataset coming from 13 hospitals (n = 23704); Hospitals in the derivation dataset represent large hospitals with ≥1000 NBU admissions per year, with the validation dataset having hospitals with <1000 NBU admissions per year. Although hospital volume may capture differences in patient populations and care settings, it was used as a pragmatic criterion to define the development and validation datasets to ensure that the development dataset contained a sufficiently large number of neonates to support stable estimation of the model-based clusters while minimising the influence of very small hospital-specific samples.

The primary objective of this study is to identify clinically meaningful sub-populations of neonates with pSBI signs at admission based on key clinical features, derived using unsupervised clustering techniques fitted on the development dataset (n = 33094 admissions from 8 hospitals). The selected pSBI clinical variables included signs of local infection, fever, hypothermia, difficulty feeding, difficulty breathing, convulsions, apnoea, floppy, indrawing, grunting, crackles, jaundice, slow capillary refill, central cyanosis, irritable, hypoxic, and bulging fontanelle [14]. The second objective’s outcome was the predicted cluster membership in the external dataset (n = 23704 admissions from 13 hospitals) from a parsimonious classification model compared to the derived cluster membership from using the original model-based clustering approach from objective 1. For the third objective, the outcome of interest was mortality at discharge.

This study analysed two sets of predictors corresponding to the study objectives. For the first and second objective, clinical variables listed in the section above which are commonly associated with sepsis/pSBI [14] were used to identify and characterise distinct clusters of neonates. Additionally, objective two adopted feature selection based on feature importance to identify a parsimonious set of pSBI predictors for cluster membership classification [28].

The identified clusters from objective one are the primary predictor variables for objective three, when assessing their association with mortality outcome; To ensure robust estimation, adjustments were made for potential confounders, neonate’s age and birth weight, given their well-established impact on length of stay and mortality [29–31].

Previous studies using CIN data showed that missing at random (MAR) is a reasonable assumption for our dataset [32]; For objective one, the model-based clustering approach adopts a Full Information Maximum Likelihood (FIML) approach where the cluster memberships are computed using all available information while taking the missingness into account under the MAR assumption, therefore no imputation was required [33].

For objective three, with the MAR assumption, multiple imputation using the chained equation approach was separately applied for both the development and validation datasets [34] in predicting in-hospital mortality outcome.

Latent Class Analysis (LCA) which is a Finite Mixture Modelling approach was used for objective one since it offered a model-based (i.e., probabilistic) approach to derive clusters [33,35]; LCA adopts Full Information Maximum Likelihood (FIML) approach making use of observations with missing data in the cluster estimation process [33,35,36]. LCA uses the distribution of the dataset to assess probabilities that certain patients are members of certain clusters. Best model-fit was determined by evaluating different metrics from LCA approach summarised in S2 Table; In summary the most probable number of clusters is determined by clear delineation (i.e., entropy thresholds >0.8), the smallest cluster size having considerable number of admissions, with the cluster model able to accurately predicting class membership for individual patients (i.e., Average Latent Class Posterior Probability (ALCPP) thresholds >0.85) (S2 Table). The final LCA model used to generate the clusters in the development dataset was also used to also predict the cluster membership of the validation dataset.

For each LCA-derived cluster, the label was assigned by the research team after examining the mortality outcome and the cluster’s characteristic pattern of conditional response probabilities across the observed clinical variables; and assigning a clinically meaningful descriptor that reflected its dominant features [37].

As part of sensitivity analysis, we compared the model-based clustering approach to a distance-based unsupervised machine learning technique of agglomerative hierarchical cluster analysis (HCA) that groups patients with similar pSBI into clusters without imposing a specific sequential order on them [38,39]. Agglomerative HCA was selected for sensitivity analysis due to its efficiency in handling mixed data types and its ability to generate clinically meaningful clusters that are stable and easy to interpret [40]. Given our dataset consists of both continuous and categorical variables, standard Euclidean distance measures suited for numerical data, were unsuitable [41]; Gower’s distance which is well suited for mixed data types was used instead [42,43], together with Ward’s minimum variance method as the linkage criterion to group neonates into meaningful clusters by minimising the total within-cluster variance [44]. The optimal number of clusters from HCA approach was determined using the silhouette score [45]. The Silhouette Score evaluated how well individual data points fit within their assigned clusters relative to the nearest alternative cluster, and ranges from −1–1 with higher scores indicating better clustering [45].

For model external validation in objective 2(a), the predicted class-membership from the original fitted LCA model in objective 1 was compared to class-membership from a new LCA model fitted directly to the external dataset.

For objective 2(b), we used eXtreme Gradient Boosting (xgboost), a machine learning technique with implicit features selection that is robust to missing data patterns, to identify a subset of pSBI features based on feature importance to predict cluster membership [46,47]. Feature selection was based on SHAP (SHapley Additive exPlanations) values where the SHAP values derived from xgboost model which explain how the patient-level variables contribute to the cluster assignment [48]; The process is explained in detail in S1 Text. Variables for inclusion in the final model were selected by visual inspection (“eyeballing”) of the SHAP summary plots, prioritising predictors that demonstrated relatively large and consistent absolute SHAP contributions across observations. The parsimonious multivariate logistic regression model was developed from the development dataset (n = 33094 admissions from 8 hospitals).

For evaluation of model transportability in objective 2(b) using the validation dataset (n = 23704 admissions from 13 hospitals), the predicted class-membership from the parsimonious model (which was determined based on which cluster the patient had the highest probability of being a member of), was compared to class-membership of the validation dataset predicted from the original fitted LCA model in objective one.

Typical classification metrics of accuracy, sensitivity, specificity, positive predictive value (PPV) etc. were used to report performance of the parsimonious classification model to correctly cluster patients into the model-based clusters.

For objective three, logistic regression without variable selection was used together with the imputed datasets with parameter estimates being combined using Rubin’s rule. The number of imputation datasets to use was determined from the integer value of the percentage of patients in the derivation dataset that had one or more missing values, rounded upwards. To examine heterogeneity in model performance while incorporating geographical external validation, we compared the logistic regression models internal-external cross-validation performance where we omitted one hospital at a time using it as the validation dataset, built the model on the remaining hospitals, and evaluated model’s discrimination and calibration performance on the hospital left out. We repeated this process with each iteration using a different hospital as the validation data source [49]. The predictive performance of in-hospital mortality by the model-based and hierarchical clustering approaches was then compared to the WHO’s pSBI guideline-based clustering approach (S1 Table). Model performance was assessed using calibration and discrimination performance metrics detailed in S3 Table, with the confidence intervals for both c-statistic and calibration slope and intercept, calculated through 1000 iterative bootstrap samples with replacement, by resampling 10% of the held-out hospital dataset (i.e., the validation dataset) in each iteration.

For the first and second objective, given clustering is an unsupervised statistical learning technique that does not rely on predefined sample sizes but rather extracted patterns from the available data [50,51], coupled with a large development dataset (n = 33094 admissions) and validation dataset (n = 23704 admissions), a formal sample size calculation was deemed unnecessary. All eligible neonatal records were thus included to maximize the robustness of the clustering approach.

For the third objective, we based our sample size calculation on previously reported mortality outcome prevalence of between 9% and 14% with R-squared values of 0.453 derived from previously-developed mortality risk prediction models in similar neonatal patient populations [52]. Using the pmsampsize library in R [53] and assuming a maximum of six clinically-meaningfully clusters identified [54,55] after including age (in days) and birth weight, the required sample size for the mortality risk model development and validation would be 260 patients with 24 deaths; making both our development dataset with 33094 admissions (with 9768 deaths) and validation dataset with 23704 admissions (with 4180 deaths) satisfactory for in-hospital mortality clinical prediction model development and external validation (S4 Table).

Between 01st January 2022 and 31st December 2024, 73140 neonates aged 0–28 days were admitted to newborn units across the 21 CIN hospitals. For analysis, 59302/73140 (81.08%) of the admissions had at least one pSBI sign/symptom and were eligible for inclusion in the study (Fig 1) with patterns of pSBI signs/symptoms illustrated in Table 1 and S1 Fig; Of the 59302 admissions with ≥1 pSBI sign/symptom, 2504/59302 (4.22%) had signs of systematic missingness (i.e., ≥ 7/19 (≥35%) pSBI signs not documented; Within the CIN, this level of missingness is indicative of missing documentation due to patients with either severe symptoms, in emergency situations or paper documents removed from patient files (S2 Fig); The patients with systematic missingness were excluded due to high likelihood of introducing bias if included in subsequent analysis due to violating the missing at random (MAR) assumption. The neonatal admissions in CIN during this period were divided geographically by hospitals into a development dataset from 8 CIN hospitals (n = 33094/56798, 58.27%), and a validation set from the remaining 13 CIN hospitals (n = 23704/56798, 41.73%) (Table 1). Excluding sepsis admission diagnoses, respiratory distress syndrome, birth asphyxia, and meconium aspiration were the most common admission diagnoses in the included patients (S3 Fig).

https://doi.org/10.1371/journal.pone.0359212.g001

https://doi.org/10.1371/journal.pone.0359212.t001

Table 2 illustrates the model-based clustering fit statistics in the development dataset. The identified clusters are of patients with pSBI managed for sepsis grouped to maximise similarity in pSBI characteristics. From the model fit statistics described in S2 Table, the maximum number of clusters with clear delineation (i.e., entropy thresholds > 0.8), whose smallest cluster size had arguably considerable number of admissions, and whose cluster model was able to accurately predicting class membership for individual patients (i.e., Average Latent Class Posterior Probability (ALCPP) thresholds > 0.85), is five (Table 2).

https://doi.org/10.1371/journal.pone.0359212.t002

Fig 2 illustrates the difference in the probability of pSBI signs and symptoms given cluster membership. The pSBI signs of floppy, tachypnoea, grunting, hypothermia, indrawing, hypoxia, difficulty in feeding and difficulty breathing appear to be useful in distinguishing different clusters (Fig 2). The labels “minimal,” “low,” “moderate,” “substantial,” and “critical” were intended as clinically meaningful descriptors of the relative mortality risk across the clusters; no predefined mortality thresholds were used to assign these labels [37].

https://doi.org/10.1371/journal.pone.0359212.g002

There is substantive variation with in-hospital mortality based on cluster membership, with the minimal risk cluster having the least mortality cases (0.98%) and critical risk having the highest mortality cases (67.8%). The same pattern is present in the validation dataset, with 3.06% and 61.76% case fatality rate between the minimal risk cluster and critical risk cluster (Table 3). Patients in the critical risk cluster had low gestational age in weeks and low birth weight while patients in the substantial risk cluster had normal gestational age and birth weight but both clusters had at least 4 pSBI signs/symptoms and substantially high mortality rates of >40% in both the development and validation datasets (Table 3). Admissions in the minimal and low mortality risk clusters had 2 (IQR: 1–3) pSBI signs/symptoms with antibiotics prescribed in ≥45% of the cases within each of these clusters in both the development and validation datasets.

https://doi.org/10.1371/journal.pone.0359212.t003

External validation demonstrated high concordance between cluster membership predicted by the original LCA model and cluster membership derived from an independently fitted LCA model in the external dataset (Table 4). Balanced accuracy ranged from 93.13% to 98.59%, with F1 scores ranging from 91.10% to 97.49% across the five risk classes. Sensitivity was highest for critical (97.53%), moderate (96.96%), and low risk (96.76%) classes, while specificity remained high across all classes (94.90% to 99.66%). Overall, the findings support the reproducibility and generalisability of the five-cluster risk structure, with strong classification performance maintained across all risk classes (Table 4).

https://doi.org/10.1371/journal.pone.0359212.t004

The pSBI signs and symptoms used in model-based clustering were 19, excluding birth weight and gestational age and the mortality outcome. From feature selection using xgboost, we subjectively judged 11 (52%) of pSBI signs and symptoms to be sufficiently capable of assigning the patients into the mortality risk clusters using a multivariate logistic model with the outcome as the clusters based on values (Table 5, S4 Fig). From table 5, the cluster membership of any neonate with a pSBI sign/symptom can be determined based on which cluster the patient had the highest probability of being a member of, with the cluster-specific probabilities computed from the sum of the logits based on the presence/absence of the sign or symptom. The transportability assessment demonstrated good overall performance of the simplified classifier in assigning patients to the predefined LCA cluster membership, with balanced accuracy ranging from 83.16% to 95.36% and F1 scores from 78.01% to 93.14% across the five classes. Performance was particularly strong for the low, moderate, and critical Risk classes, while the high specificity across all classes (86.91% to 99.39%) supports the classifier’s ability to reliably distinguish patients belonging to the predefined risk groups (Table 6).

https://doi.org/10.1371/journal.pone.0359212.t005

https://doi.org/10.1371/journal.pone.0359212.t006

Fig 3 compares the performance of the three clustering approaches in predicting mortality risk. Overall, 33004/33094 (99%) admissions of the development dataset and 23601/23704 (99%) of the validation dataset were eligible for this comparison; The omitted admissions were excluded since their clinical signs and symptoms did not align with any of the predefined WHO clusters (S1 Table), due to variable missingness, posing a challenge in assigning them to a WHO cluster.

https://doi.org/10.1371/journal.pone.0359212.g003

The c-statistic (discrimination) for the WHO expert-based clustering approach was 0.721 (95% CI: 0.715 to 0.727) with a calibration intercept of 0.018 (95% CI: −0.005 to 0.041) and calibration slope of 1.015 (95% CI: 0.986 to 1.043). The c-statistic (discrimination) for the distance-based clustering approach was 0.721 (95% CI: 0.714 to 0.727) while its calibration intercept was −0.002 (95% CI: −0.025 to 0.021) and calibration slope was 0.987 (95% CI: 0.96 to 1.014). The WHO and distance-based clustering approach calibration curves both showed substantive variances across the mortality risk spectrum with over-estimation of mortality at the lowest predicted risks and under-estimation at moderate to high risks of mortality (Fig 3).

The model-based clustering approach had the highest c-statistic (discrimination) of 0.867 (95% CI: 0.863 to 0.871), with a calibration intercept of −0.004 (95% CI: −0.031 to 0.023) and a calibration slope of 0.996 (95% CI: 0.979 to 1.043). The model-based clustering approach’s calibration curve showed close and relatively better alignment between the predicted risks and observed outcomes across the probability range compared to the alternative clustering approaches (Fig 3).

This study used a pragmatic approach with a limited set of signs easy to use by clinicians (usually busy junior clinicians) to identify residual heterogeneity in the neonates admitted with signs indicative of sepsis, that can be used to inform antibiotic use treatment decisions and strategies. From applying model-based clustering approaches on routinely collected neonatal data at hospital admission, we identified five clinically distinct clusters of neonates with pSBI managed for sepsis, labelled based on the mortality risk in each cluster: minimal risk, low risk, moderate risk, substantial risk and critical risk (Table 3, Fig 2). From external validation, we further demonstrate that the model-based clusters outperform distance-based and WHO expert-based clustering approaches in predicting in-hospital mortality based on discrimination and calibration statistics (Fig 3). Among the 19 pSBI signs/symptoms collected at admission and used in the initial model-based clustering, a subset of 11 pSBI signs and symptoms together with birth weight, gestational age, early-onset age clinical variables were sufficient to determine cluster membership of a patient and subsequently the associated mortality risk from pSBI at admission (Table 5, Table 6).

Hypothermia, indrawing, floppy, tachypnoea, grunting, hypoxia, difficulty in feeding and difficulty breathing were key in distinguishing the clusters. These signs were mostly prevalent in the substantial and critical risk groups with tachypnoea being more common in the substantial risk cluster while hypothermia and indrawing were more common in the critical risk cluster. The critical risk cluster also comprised of neonates with low gestational age and birthweight while those in substantial risk cluster had normal gestation age and birthweight yet still faced substantially high mortality rate. The moderate risk cluster on the other hand, included neonates with low birth weight but demonstrated a relatively lower rate of mortality risk, highlighting the heterogeneity in mortality risk beyond birth-related factors.

The findings of this study contribute to the growing evidence of how unsupervised machine learning methods might identify meaningful phenotypes in neonatal sepsis populations. For instance, Seymour et al. and Jang et al. [4,15] employed k-means clustering and identified up to four clusters. Sinha et al. used model-based clustering and derived two clusters, unlike our findings. These studies used patients with different characteristics and clinical variables whose availability at the time of admission varied considerably. Unlike this current study which used routinely collected clinical data from low-resource neonatal units in an SSA country and relied only on clinical signs and symptoms, some of the other prior studies were conducted in high-income settings and included a set of variables combining both clinical and laboratory parameters that are typically unavailable at neonatal admission such as white blood cell count, platelet count, C-reactive protein, premature neutrophil count and serum levels.

Furthermore, these studies neither directly assessed the external validity of the clustering approach to classify new patients into the identified clusters nor did they provide the probability of patients having a combination of signs and symptoms given their cluster membership. These methodological differences highlight how selection of variables and validation approaches influence findings across studies. The observed differences in mortality risk among the clusters in this study and how the risk differs based on the probability of pSBI in each cluster agree with findings from Puri et al [56] although our findings support five distinct neonatal sepsis clusters. Hypothermia, among the signs previously identified as strongly predictive of death, was particularly common in the critical risk cluster of our study which had the highest mortality risk.

The clusters identified through model-based unsupervised learning approach reliably stratify neonatal sepsis into patient sub-populations predictive of in-hospital mortality at admission, outperforming the WHO expert-based and distance-based clustering methods as exemplified by the reported discrimination and calibration statistics (Fig 3). This work shows that there is residual heterogeneity in the neonatal sepsis in-hospital admissions linked to mortality as an outcome. However, mortality may also have been influenced by differences in treatment and other aspects of care received during hospitalisation, which were not accounted for in the clustering model. Therefore, the observed differences in mortality between clusters should not be interpreted as reflecting differences in underlying disease severity alone, and prospective studies are needed to disentangle the contribution of baseline phenotype from subsequent treatment and care. The improved mortality risk stratification exemplified by our model-based clustering approach provides a foundation for more tailored clinical decision making and potential development of phenotype-specific treatment guidelines that does not neglect the typical multimorbid nature of neonatal sepsis admissions in routine SSA hospitals. Additionally, the identification of substantial and critical mortality risk phenotype with distinct clinical profiles could guide prioritisation of care or more aggressive interventions, potentially improving outcomes in these groups. In the context of the growing global threat of antimicrobial resistance (AMR), the finding that more than 45% of neonates admitted to CIN facilities in the minimal and low mortality-risk clusters received antibiotics warrants further evaluation of whether antibiotic use could be safely reduced in some of these patients. The clusters identified in this study provide a more granular stratification than the WHO categories and could provide a basis for prospectively evaluating whether different antibiotic treatment strategies, such as shorter intravenous courses or, where clinically appropriate, oral therapy, can be used safely and effectively.

These findings can either (a) be used to augment the clinical judgement of clinicians in risk stratification of neonatal sepsis patients at admission [57] and/or (b) inform the design and implementation of emulated trials [58] to evaluate different antibiotic treatment strategies for neonatal sepsis at admission informed by the identified clusters [59,60].

By relying on clinical signs and symptoms readily available at admission, the model-based clusters have better chances of being incorporated into routine clinical workflows at the point of care in similar SSA hospital settings [18,25], with the potential to enhance early tailored interventions and treatment to ensure critical care is efficiently administered where it is most needed.

This study has several notable strengths. First, it is based on a large multi-centre cohort of neonatal admission (n = 56798) across 21 hospitals over three years, ensuring geographical and time coverage thus increasing confidence in the applicability of the modelling approaches across diverse neonatal care contexts. The inclusion criteria ensured the study sample was reflective of the typical multi-morbidity of neonatal admissions, thereby strengthening the real-world relevance of the findings. From a methodological standpoint, this study advances neonatal sepsis phenotyping by moving beyond conventional distance-based clustering techniques to more superior probabilistic, multivariate model-based approaches, which unlike distance-based methods, do not use a forced-choice mechanism to assign patients to clusters.

The absence of microbiological validation to confirm the presence of bacteria indicative of sepsis is a limitation to this study. Since the aim was to classify neonates based on clinical pSBI signs in recognition of limited laboratory capacity in similar SSA hospital settings [18,25], the clusters should hence be interpreted as clinical phenotypes rather than pathogen confirmed syndromes. Future work would involve validating and extending this study’s findings by integrating microbiological testing with clinical clustering to ascertain whether the data-driven phenotypic clusters truly reflect bacterial infections. Prospective studies could also evaluate whether these groups respond differently to interventions or management strategies, providing insights into their clinical utility and potential to inform targeted risk-based neonatal care.

In this retrospective cohort study, we used pSBI signs that are readily available at neonatal admission across many SSA hospital settings for model-based clustering approach and identified five clinically distinct clusters differentiated by mortality risk: minimal, moderate, substantial and critical risk. From discrimination and calibration statistics, the model-based clustering approach outperformed clustering patients using WHO expert-based guidelines and typical distance-based clustering approaches. These clusters demonstrate the potential to complement clinical judgement in mortality risk stratification at admission and could inform future randomised clinical trials to test whether triage guided by these clusters could improve clinical outcomes such as mortality and length of hospital stay. Future work could also integrate microbiological validation to confirm whether clusters map onto pathogen-specific syndromes and investigate possible variation in treatment responses across the clusters. These efforts could help strengthen the case for more targeted and risk-based neonatal sepsis management, informed by data-driven clustering.

This study is published with the permission of the Director of Kenya Medical Research Institute (KEMRI).

WHO guidance spans newborns aged 0–59 days but this analysis is restricted to those aged 0–28 days (i.e., neonatal period).

https://doi.org/10.1371/journal.pone.0359212.s001

https://doi.org/10.1371/journal.pone.0359212.s002

https://doi.org/10.1371/journal.pone.0359212.s003

Data from the included patients. Numerator: total number of deaths; Denominator: total number of admissions to the NBU.

https://doi.org/10.1371/journal.pone.0359212.s004

pSBI signs and symptoms grouped by the different datasets.

https://doi.org/10.1371/journal.pone.0359212.s005

Patterns of pSBI signs and symptoms missingness grouped by the different datasets.

https://doi.org/10.1371/journal.pone.0359212.s006

Common admission diagnoses for neonates managed for sepsis (n = 8012), excluding sepsis diagnosis.

https://doi.org/10.1371/journal.pone.0359212.s007

https://doi.org/10.1371/journal.pone.0359212.s008

https://doi.org/10.1371/journal.pone.0359212.s009

The Clinical Information Network (CIN) Group: The members of the CIN group who collaborated in the network’s development, data collection, data management, implementation and who participated in this study are:

a. Paediatricians: Juma Vitalis (Department of Health, County Government of Vihiga), Nyumbile Bonface ((Department of Health, County Government of Kakamega), Roselyne Malangachi (Department of Health, County Government of Kakamega), Christine Manyasi (Department of Health, County Government of Nairobi), Catherine Mutinda (Department of Health, County Government of Nairobi), David Kibiwott Kimutai (Department of Health, County Government of Nairobi), Rukia Aden (Department of Health, County Government of Nairobi), Caren Emadau (Department of Health, County Government of Nairobi), Elizabeth Atieno Jowi (Department of Health, County Government of Nairobi), Cecilia Muithya (Department of Health, County Government of Nairobi), Charles Nzioki (Department of Health, County Government of Machakos), Supa Tunje (Department of Health, County Government of Machakos), Penina Musyoka (Department of Health, County Government of Machakos), Wagura Mwangi (Department of Health, County Government of Nyeri), Agnes Mithamo (Department of Health, County Government of Nyeri), Magdalene Kuria (Department of Health, County Government of Kisumu), Esther Njiru (Department of Health, County Government of Embu), Mwangi Ngina (Department of Health, County Government of Embu), Penina Mwangi (Department of Health, County Government of Kirinyaga), Rachel Inginia (Department of Health, County Government of Trans-Nzoia), Melab Musabi (Department of Health, County Government of Trans-Nzoia), Emma Namulala (Department of Health, County Government of Busia), Grace Ochieng (Department of Health, County Government of Kiambu), Lydia Thuranira (Department of Health, County Government of Kiambu), Felicitas Makokha (Department of Health, County Government of Bungoma), Josephine Ojigo (Jaramogi Oginga Odinga Teaching and Referral Hospital), Beth Maina (Department of Health, County Government of Nairobi), Catherine Mutinda (Department of Health, County Government of Nairobi), Mary Waiyego (Kenyatta National Hospital), Bernadette Lusweti (Department of Health, County Government of Kiambu), Angeline Ithondeka (Department of Health, County Government of Nakuru), Julie Barasa (Department of Health, County Government of Nakuru), Meshack Liru (Department of Health, County Government of Homabay), Elizabeth Kibaru (Department of Health, County Government of Nakuru), Alice Nkirote Nyaribari (Department of Health, County Government of Nakuru), Joyce Akuka (Department of Health, County Government of Migori), Joyce Wangari (Department of Health, County Government of Migori);

b. Nurses: Amilia Ngoda (Department of Health, County Government of Vihiga), Aggrey Nzavaye Emenwa (Department of Health, County Government of Vihiga), Patricia Nafula Wesakania (Department of Health, County Government of Kakamega), George Lipesa (Department of Health, County Government of Kakamega), Jane Mbungu (Department of Health, County Government of Nairobi), Marystella Mutenyo (Department of Health, County Government of Nairobi), Joyce Mbogho (Department of Health, County Government of Nairobi), Joan Baswetty (Kenyatta National Hospital), Ann Jambi (Department of Health, County Government of Nairobi), Josephine Aritho (Department of Health, County Government of Nairobi), Beatrice Njambi (Department of Health, County Government of Nairobi), Felisters Mucheke (Department of Health, County Government of Nairobi), Zainab Kioni (Department of Health, County Government of Machakos), Jeniffer, Lucy Kinyua (Department of Health, County Government of Nyeri), Margaret Kethi (Department of Health, County Government of Nyeri), Alice Oguda (Department of Health, County Government of Kisumu), Salome Nashimiyu Situma (Department of Health, County Government of Busia), Nancy Gachaja (Department of Health, County Government of Embu), Loise N. Mwangi (Department of Health, County Government of Embu), Ruth Mwai (Department of Health, County Government of Embu), Virginia Wangari Muruga (Department of Health, County Government of Kirinyaga), Nancy Mburu (Department of Health, County Government of Kirinyaga), Celestine Muteshi (Department of Health, County Government of Trans-Nzoia), Abigael Bwire (Department of Health, County Government of Trans-Nzoia), Salome Okisa Muyale (Department of Health, County Government of Busia), Naomi Situma (Department of Health, County Government of Kisumu), Faith Mueni (Department of Health, County Government of Kiambu), Hellen Mwaura (Department of Health, County Government of Kiambu), Rosemary Mututa (Department of Health, County Government of Bungoma), Caroline Lavu (Department of Health, County Government of Bungoma), Joyce Oketch (Jaramogi Oginga Odinga Teaching and Referral Hospital), Jane Hore Olum (Jaramogi Oginga Odinga Teaching and Referral Hospital), Orina Nyakina (Department of Health, County Government of Nairobi), Faith Njeru (Department of Health, County Government of Nairobi), Rebecca Chelimo (Department of Health, County Government of Nairobi), Margaret Wanjiku Mwaura (Department of Health, County Government of Kiambu), Ann Wambugu (Department of Health, County Government of Nairobi), Epharus Njeri Mburu (Department of Health, County Government of Nakuru), Linda Awino Tindi (Department of Health, County Government of Homabay), Jane Akumu (Department of Health, County Government of Homabay), Ruth Otieno (Department of Health, County Government of Migori), Slessor Osok (Department of Health, County Government of Migori);

c. Health Record Information Officers (HRIOs): Seline Kulubi† (Department of Health, County Government of Bungoma), Susan Wanjala (Department of Health, County Government of Busia), Pauline Njeru (Department of Health, County Government of Embu), Rebbecca Mukami Mbogo (Department of Health, County Government of Embu), John Ollongo (Jaramogi Oginga Odinga Teaching and Referral Hospital), Samuel Soita (Department of Health, County Government of Kakamega), Judith Mirenja (Department of Health, County Government of Kakamega), Mary Nguri (Department of Health, County Government of Kirinyaga), Margaret Waweru (Department of Health, County Government of Kiambu), Mary Akoth Oruko (Department of Health, County Government of Kisumu), Jeska Kuya (Department of Health, County Government of Trans-Nzoia), Caroline Muthuri (Department of Health, County Government of Machakos), Esther Muthiani (Department of Health, County Government of Machakos), Esther Mwangi (Department of Health, County Government of Nairobi), Joseph Nganga (Department of Health, County Government of Nairobi), Benjamin Tanui (Department of Health, County Government of Nakuru), Alfred Wanjau (Department of Health, County Government of Nyeri), Judith Onsongo (Department of Health, County Government of Nairobi), Peter Muigai (Department of Health, County Government of Kiambu), Arnest Namayi (Department of Health, County Government of Vihiga), Elizabeth Kosiom (Department of Health, County Government of Trans-Nzoia), Dorcas Cherop (Department of Health, County Government of Trans-Nzoia), Faith Marete (Department of Health, County Government of Nakuru), Johanness Simiyu (Jaramogi Oginga Odinga Teaching and Referral Hospital), Collince Danga (Department of Health, County Government of Homabay), Arthur Otieno Oyugi (Department of Health, County Government of Migori), Fredrick Keya Okoth (Department of Health, County Government of Migori).

The Clinical Information Network (CIN) Group’s monitored email address is CIN@kemri-wellcome.org and the list can change when new paediatrician(s), nurse(s) or HRIO leave or come into the hospital. The lead author for this group is Prof. Mike English (email: mike.english@ndm.ox.ac.uk).

This is an open access article distributed in accordance with the Creative Commons Attribution 4.0 Unported (CC BY 4.0) license, which permits others to copy, redistribute, remix, transform and build upon this work for any purpose, provided the original work is properly cited, a link to the licence is given, and indication of whether changes were made.

在原文站打开 ↗

Cloudflare Workers 每 3 分钟抓一批,9 批轮完最快约 27 分钟 · 点右上 ↻ 立刻全量抓一次