Research & evidence

A practical library on the scientific foundation of polygenic scores: from publications, model specifications, and reference data to validation, calibration, technical quality, and responsible implementation.

Reading evidence responsibly

This page shows how Starling reads evidence for polygenic-score applications. Not every publication, model, or analytical capability is automatically ready for routine use; value depends on intended use, data quality, external validation, calibration, and local context.

The central theme is caution with purpose. Polygenic scores are probabilistic instruments, and conclusions must never exceed what the data, model performance, and current evidence support.

Evidence across the PRS lifecycle

Responsible PRS implementation connects discovery evidence, model specification, genotype and imputation quality, external validation, calibration, ancestry context, and comparison with established clinical predictors.

Starling therefore makes clear which sources support each score, what validation is needed, which conclusion can be reported, and where population, pathway, or territory-specific limitations apply.

  • Model performance
  • Model provenance
  • Ancestry and transferability
  • Clinical utility
  • Reporting and governance
  • Cohorts and implementation

Deep dive: polygenic scores at a glance

Across most clinical areas, PRS is best supported as a risk modifier that complements standard clinical assessment. Evidence is strongest where PRS is integrated with existing pathways (screening, monitoring, classification), and weakest where PRS is proposed as a stand-alone decision tool. [2]

What PRS does well

Stratifies inherited risk at population scale, often improving discrimination when combined with age, family history, biomarkers, and clinical factors. [7]

What PRS does not do

It does not diagnose disease, and a low score does not rule out disease. Clinical context remains essential. [2]

Main implementation barrier

Ancestry portability and calibration: performance can drop when a score is transferred across populations without local validation. [13]

PRS deep dive: meaningful evidence

Meaningful evidence is rarely a single headline result. It usually combines discovery studies, population biobanks, ancestry-reference resources, and external validation cohorts that show whether the score remains interpretable beyond its original training environment.

For real implementation questions, the strongest studies are the ones that compare polygenic models with existing predictors, test calibration carefully, and examine how reporting could work inside routine pathways rather than in isolation.

Reading PRS evidence by domain

We use a simple three-part framework for each domain: clinical validity, incremental value, and implementation readiness. [3]

1. Clinical validity

Can the PRS distinguish relatively higher-risk versus lower-risk individuals in relevant populations?

2. Incremental value

Does PRS materially improve prediction beyond established predictors already used in care?

3. Implementation readiness

Is the score reproducible, calibrated, ancestry-appropriate, and interpretable for safe reporting?

Cross-domain reality

Different PRS for the same endpoint can perform similarly at population level but classify some individuals differently, which matters for hard percentile cutoffs. [5]

PRS source cohorts

The currently available indications sit within a broader PGS literature that is rarely built from a single dataset. Most usable models combine discovery GWAS, ancestry reference resources, and external validation cohorts for development, calibration, and transfer testing.

Population biobanks

Large resources such as UK Biobank and All of Us are commonly used to train, benchmark, or validate broad disease and quantitative trait scores at scale.

Ancestry reference panels

Datasets such as 1000 Genomes are commonly used to support ancestry assignment, allele-frequency checks, and more careful interpretation across diverse populations.

Clinical validation cohorts

Prospective or deeply phenotyped cohorts such as MESA and ARIC are especially valuable because they test whether performance holds up outside the original discovery environment.

Disease-specific sources

Some indications also rely on specialist consortia or registries, for example CARDIoGRAMplusC4D, Biobank Japan, CIMBA, and WHI, to strengthen domain-specific risk modeling.

PRS literature across multiple domains

The wider literature across oncology, ophthalmology, immune disease, bone health, neurology, psychiatry, renal medicine, and quantitative traits remains highly relevant. It shows where polygenic methods are gaining traction and where translational potential is beginning to appear.

At the same time, those domains differ substantially in readiness. Some have strong population-level signal but weak pathway integration; others have promising screening or surveillance relevance but still depend on better calibration, ancestry transfer, or prospective outcome data.

PRS evidence matrix

Quick comparison of where evidence is strongest today and what still limits routine use. Ratings are synthesis-level judgments based on the cited studies and consensus statements on this page.

Domain Evidence Strength Near-Term Clinical Readiness Main Caveat
Oncology High High Outcome impact and equity across ancestry groups still need stronger prospective data. [9] [10]
Cardiometabolic & Vascular High Moderate Incremental value over strong clinical models varies by endpoint and care setting. [6] [7]
Renal Moderate Moderate Integration with APOL1/clinical data helps, but prospective decision-impact data is limited. [8]
Bone Health Moderate Moderate Thresholds and pathway standardization are still developing across populations. [14]
Gastroenterology & Immune Moderate Moderate Actionable clinical pathways are uneven despite strong genetic signal. [15]
Ophthalmology Moderate Moderate Calibration and validated treatment or surveillance thresholds remain key constraints. [11]
Neurology & Psychiatry Moderate Low-Moderate Effect sizes, portability, and ethical complexity limit routine susceptibility use. [16]
Quantitative Traits High Moderate Strong statistical signal does not automatically support direct treatment decisions. [12]

Quantitative traits and conditional use

Trait PRS can be statistically strong because traits are directly measured at scale. Clinical value depends on whether trait information connects to validated intervention pathways. [12]

  • Best-supported use today: integrate trait PRS with disease PRS and clinical factors in endpoint-specific models.
  • More caution is needed when acting directly on trait PRS without prospective evidence of improved outcomes.
  • Implementation need: population-specific reference distributions and consistent reporting standards.

Requirements for clinical use

Clear use case
Define whether the score is for risk stratification, classification support, or pathway triage.
Ancestry-aware validation
Validate locally and report clearly where performance is less transferable.
Calibration and thresholding
Avoid unvalidated hard cutoffs; monitor drift over time and across cohorts.
Interpretation safeguards
Communicate probabilistic risk clearly for both clinicians and patients.
Operational reproducibility
Use transparent pipelines, versioned models, and quality controls suitable for clinical settings.
Outcome evidence
Prioritize studies showing whether PRS-guided decisions improve real-world outcomes.

Technical discussion?

We discuss evidence, model selection, cohort context, and interpretation questions with clinics, healthcare providers, laboratories, and research teams implementing polygenic scores.