All articles Science

Evidence Quality Tiers: A Framework for Drug Discovery Target Validation

8 min read Naomi Stein
Evidence Quality Tiers: A Framework for Drug Discovery Target Validation

The evidence that bears on a drug target is not uniform. A finding from a replicated in-vivo study in a tumor-type-specific model is not the same kind of signal as a computational prediction derived from protein interaction network proximity. Treating them as equivalent inputs to target prioritization is a category error that quietly distorts the ranking.

This is the reasoning behind the evidence quality tier system we built into the Avenzo scoring model. The framework was developed through our own prior work in translational oncology programs, and it reflects the distinctions that practicing solid-tumor biologists actually apply when they read the literature and decide whether a finding is convincing. Here is how we think about it.

The Core Problem: Not All Sources Are Equal

Drug target validation literature spans an enormous range in terms of experimental rigor. At one end: replicated genetic loss-of-function studies in syngeneic or patient-derived xenograft tumor models, with consistent phenotypic outcomes and a mechanistic explanation for why the target matters. At the other end: a computational prediction based on sequence homology to a known oncogene, or a single cell-line sensitivity assay in a commercially available cell line that may or may not be representative of the tumor type in question.

Both types of evidence appear in target identification workflows. Both are cited in target review documents. And both are represented in the published literature in roughly comparable numbers. The difference in what they tell you about target confidence is not subtle. One is strong evidence for pursuing assay development. One is a starting hypothesis that requires substantial additional support before it justifies bench investment.

Without a systematic way to distinguish between them, scoring models that aggregate evidence by count are treating these as equivalent signals. Four Tier-4 computational predictions do not equal one Tier-1 in-vivo replication. But they can look similar in a citation-count or paper-volume summary.

The Four-Tier Framework

Our tier structure is built around four dimensions that together determine how much epistemic weight to assign a piece of preclinical evidence: experimental system type, independence of replication, tumor type specificity, and mechanistic clarity.

Tier 1 is the highest-confidence category. It requires at least one of the following: replicated in-vivo findings across independent research groups in a tumor-type-relevant model, or a single high-quality in-vivo study with a clear genetic mechanism and strong tumor-type specificity. Tier 1 evidence tells you that the target has a validated functional role in relevant tumor biology under conditions that are at least partially predictive of in-vivo human tumor behavior.

Tier 2 covers unreplicated in-vivo findings from a single group, or replicated in-vitro findings across multiple independent cell-line systems with consistent directional signal. These are meaningful but require the interpretive caveat that in-vitro replication does not solve the translation gap. Consistent behavior across six cell-line systems still leaves open whether the finding will hold in a tumor microenvironment with stromal interactions and immune components.

Tier 3 covers single-study in-vitro findings in one or two cell-line systems, or in-vivo findings limited to a single non-tumor-type-specific model. These are hypothesis-generating signals that may justify deeper investigation but do not yet constitute meaningful validation evidence on their own.

Tier 4 is computational evidence without independent experimental corroboration: pathway network proximity scores, protein interaction predictions, sequence similarity to known targets, or expression correlation in tumor datasets not independently validated for the target in question. Tier 4 evidence is not worthless. It is an important starting point for target hypothesis generation. But it is the beginning of a validation process, not the end of one.

How Tiers Apply Across Evidence Types

The tier system applies differently across the three evidence dimensions we score. For preclinical support depth, the tier weighting works as described above. For clinical context alignment, the structure is adapted: Tier 1 equivalent is direct clinical outcome evidence from a relevant patient population, even observational data that shows a clear link between target activity and clinical signal. Tier 4 equivalent is a trial enrollment context where the target was mentioned in the mechanism hypothesis but where no target-specific clinical outcome data is available.

Computational signal convergence has its own tier logic. Here, Tier 1 is convergent support from multiple independent computational approaches (pathway network, protein structure prediction, and expression analysis all pointing the same direction). Tier 4 is a single computational database hit with no corroboration from orthogonal computational methods. The independence of signals matters in the computational dimension just as it does in the experimental one.

The Boundary Between Tiers and Druggability

One clarification worth making: evidence quality tiers assess target validation, not druggability. A target can be Tier 1 validated by every preclinical and clinical context measure and still be undruggable by current medicinal chemistry approaches. Protein-protein interaction targets are a well-documented example: strong biological validation, difficult or impossible to address with small molecules using current methods.

The tier framework does not address druggability. It addresses the biological question of whether the target has a validated role in the relevant tumor biology. Druggability assessment is a separate filter that operates downstream of target selection. Teams using the Avenzo evidence ranking should apply druggability evaluation in addition to, not instead of, the evidence quality assessment.

Why This Matters for Scoring Model Calibration

The practical effect of the tier weighting is that it changes which targets rank highly. A target with ten Tier 3 and Tier 4 evidence items will rank below a target with three Tier 1 and Tier 2 items, even though the first target has more total evidence. This inversion happens frequently in our analyses and is often the most useful output for a discovery team: it surfaces targets whose evidence depth is more substantial than their volume suggests, and it flags targets whose volume is driven by lower-quality signals that would not withstand scrutiny in an expert evidence review.

We built the tier weights based on our understanding of how different evidence types predict downstream validation success, informed by the published literature on preclinical-to-clinical translation. The weights are not a fixed formula that applies identically across all tumor types or target classes. We have initial calibration, and we refine it based on feedback from early-access partners who bring expert knowledge of specific target families and tumor contexts. The system is designed to be transparent: every report shows tier breakdowns so a team can inspect the weighting and apply their own judgment.

The Limits of Any Tier System

Tier frameworks are useful precisely because they impose a discipline on evidence evaluation. They are not useful if they are applied mechanically without domain expertise. A finding that technically qualifies as Tier 2 by our classification criteria may be regarded by a specialist in that tumor biology as much weaker than the tier suggests, because the experimental system used has known limitations in that specific context. Conversely, a finding we would classify as Tier 3 may be regarded by an expert as particularly compelling because it came from a highly relevant model system that is not widely used but is considered gold-standard in the field.

This is why we designed the Avenzo output to be inspectable, not just a single number. The tier-weighted score is an input to expert judgment, not a replacement for it. When a scientist who has studied a target for years looks at our tier classification and thinks we got something wrong, that is valuable calibration data for the model, not a reason to override the model's output without examination.

If you are working through target selection decisions and want to understand how a structured evidence quality framework would reorder your current candidate list, the Avenzo platform is designed for exactly that evaluation. Reach out to learn what early access involves.