All articles Methodology

Multi-Source Evidence Integration: The Foundation of Reliable Target Identification

7 min read Naomi Stein
Multi-Source Evidence Integration: The Foundation of Reliable Target Identification

Every target that makes it to a nomination decision is supported by some evidence. The question is never whether evidence exists; it is what kind of evidence exists, from how many independent sources, and what the convergence or divergence of those sources implies about target reliability. A target supported by a single dense evidence stream is a different kind of bet than one supported by multiple independent streams pointing in the same direction. The difference shows up in clinical success rates, though it is difficult to trace precisely because target selection records are rarely published with the same rigor as development outcomes.

Multi-source evidence integration is the practice of assembling and weighting these streams systematically before making a target nomination decision. This post describes why it matters and how we approached it in building the Avenzo evidence framework.

What Single-Source Evidence Gets Wrong

The most common single-source case in contemporary oncology target identification is heavy reliance on large-scale genomic datasets: cancer genomics cohort data identifying recurrent somatic alterations, CRISPR functional genomics screens identifying essential genes, or expression data linking gene activity to survival outcomes. These datasets are genuinely informative and represent a real advance over the hypothesis-driven target identification of earlier decades. The problem is using them as the primary or sole evidence source rather than as one stream in a broader synthesis.

Genomic alteration frequency tells you that a gene is altered at meaningful rates in a tumor type. It does not tell you whether those alterations are functionally driving tumor growth or are passenger events occurring in a genomically unstable context. Genes that are recurrently altered in solid tumors include many that have no functional consequence for tumor survival. The alteration frequency is a biological observation, not a target validation.

CRISPR dependency screens in cancer cell lines provide direct functional evidence that a gene is necessary for cell viability in that model. They are a significant step beyond pure genomic correlations. But dependency in a panel of cancer cell lines does not predict dependency in the in-vivo tumor context, where the microenvironment, metabolic landscape, immune pressure, and stromal survival signals are all different from standard culture conditions. Essential gene findings in screens are better starting points for target investigation, not endpoints of target validation.

The pattern across single-source approaches is the same: each source type captures one dimension of the biological question and is silent on others. Treating a strong signal in one dimension as sufficient validation is a category error, substituting depth in one source for breadth across multiple orthogonal sources.

The Case for Orthogonal Evidence Streams

The principle of requiring convergent evidence from orthogonal sources is well-established in experimental biology and is the reason replication across independent experimental systems is the standard of scientific validity. For target identification, orthogonal evidence streams are those that test different biological aspects of the target hypothesis using different experimental approaches and, ideally, different research groups.

Consider a target candidate with the following evidence profile: genomic alteration data showing recurrent somatic mutations in a specific tumor type (genomic evidence); in-vitro CRISPR screen data showing dependency in cell lines carrying the alteration (functional genetic evidence); in-vivo genetic perturbation experiments in relevant tumor mouse models showing reduced tumor growth (in-vivo functional evidence); and clinical biomarker data from trials where patient subgroups carrying the alteration showed differential outcomes (clinical context evidence). Each of these streams tests a different aspect of the target hypothesis. When all four point in the same direction, the convergence is meaningful. It is not proof of clinical success, but it represents a fundamentally different risk profile than strong signal in a single dimension.

The integration challenge is that these evidence streams vary in availability, quality, and what they actually say. Not all targets have strong in-vivo functional evidence. Not all targets have clinical context data. Some of the most important targets for a given tumor type may have excellent functional genetic evidence and pathway context support but thin clinical biomarker data simply because no relevant trials have been run yet. Evidence integration has to handle heterogeneous evidence availability without defaulting to requiring all streams to be present as a condition for target consideration.

How Evidence Integration Differs from Evidence Counting

A common oversimplification in computational target prioritization is treating evidence integration as a sum of citations or publications. More papers about a target equals better-supported target. This is not evidence integration; it is evidence counting, and it produces systematically biased rankings.

Evidence counting favors targets that have attracted research attention, which is partly correlated with biological importance but also correlated with funding patterns, scientific fashion, and the work of influential research groups who generated initial publications that attracted follow-on attention. Targets in well-funded cancer types with established research communities accumulate publications faster than biologically equivalent targets in less-studied contexts. Citation counts are a social signal as much as a scientific signal.

True evidence integration requires assigning weights to evidence types based on what they say about target validity, not just how often they appear. An in-vivo genetic perturbation study replicated across two independent mouse models is more informative for target validation than five cell-line papers from the same laboratory reporting incremental confirmations of the same finding. A weight-by-quality integration will score the former combination higher. A citation count will score the latter higher. These different scoring approaches produce different rankings, and the differences are largest precisely for the targets where the decision is most consequential.

The Integration Model in Practice

In the Avenzo evidence framework, integration happens across three primary dimensions: preclinical functional evidence, clinical context evidence, and computational signal convergence. Each dimension is scored independently before being combined. Keeping the dimensions separate preserves the interpretive transparency that teams need to evaluate the output.

Preclinical functional evidence is scored by quality tier, with in-vivo replicated evidence at the top and single cell-line results at the bottom, weighted for model system diversity and independent replication. The key design choice is that multiple low-tier evidence items do not accumulate to the weight of a single high-tier item. A target with fifteen cell-line papers and no in-vivo evidence does not score equivalently to one with strong in-vivo replication, even though its citation count may be much higher.

Clinical context evidence captures the relationship between the target and clinical outcomes in human patient data: biomarker correlations, association with survival in subgroup analyses, or evidence from trials where the target's pathway was implicated in response or resistance. This dimension is often the weakest in early discovery, because the clinical data frequently does not yet exist for novel targets. A target scoring low in this dimension is not necessarily a poor target; it may simply be in a space where the clinical investigation has not happened yet.

Computational signal convergence integrates the network-level evidence: pathway centrality, co-alteration patterns with known drivers, genetic interaction predictions, and cross-database corroboration of the target's network positioning. This dimension captures what the biology of the surrounding network implies about the target's functional role, independent of direct experimental perturbation data.

What Multi-Source Integration Cannot Solve

Multi-source evidence integration is a better methodology than single-source assessment. It is not a solution to the fundamental biological uncertainty of early drug discovery. Convergent evidence from multiple high-quality sources reduces the probability of target failures driven by evidence gaps or evidence misinterpretation. It cannot reduce failures that stem from biology that is not yet in any data source, from model-to-human translation gaps that are not predictable from existing preclinical systems, or from clinical population heterogeneity that only manifests when an unselected patient cohort enters a trial.

We are explicit about this in the Avenzo output. The evidence quality scores describe what the current publicly available data says about a candidate. They do not predict clinical success. The claim we make is narrower: that targets with converging multi-source high-quality evidence have a different failure risk profile than targets with weak or single-source evidence, and that making the evidence profile visible before commitment is better than discovering it mid-investigation.

If you are working on building or refining a target selection process where the evidence integration step feels like it has not been formalized, and you want to see what a structured multi-source evidence assessment looks like for your specific candidates, reach out to discuss the Avenzo analysis for your program.