All articles Methodology

The Data Problem in Oncology Target Selection: A Structural Analysis

7 min read Avenzo Research Team
The Data Problem in Oncology Target Selection: A Structural Analysis

Oncology research programs have access to more target-relevant data than at any previous point in the history of the field. The published preclinical literature runs into the tens of millions of papers and preprints. Genomic alteration databases index somatic mutation frequencies across tumor cohorts of tens of thousands of samples. Trial registries record clinical context that was inaccessible to discovery teams a decade ago. Protein interaction and pathway databases have grown orders of magnitude in coverage and resolution.

None of this has solved the evidence problem in target selection. If anything, it has intensified it. More data does not automatically mean better access or better synthesis. The challenge is structural.

The Fragmentation Problem

The evidence relevant to a single target nomination lives across hundreds or thousands of individual documents. A portion of it is in high-visibility journals. A portion is in lower-visibility publications, conference proceedings, and preprint servers. Some is in trial registry records that are not indexed by the literature databases most research teams use as a default starting point. Some is in computational database entries that require a different retrieval workflow than text-based literature search. And pathway context data, which can significantly alter the interpretation of direct target evidence, requires its own source layer.

No single retrieval system covers all of these sources. A PubMed search captures a portion of the peer-reviewed literature but misses preprints, conference data, and genomic database context. A computational database query captures pathway and protein interaction signals but not the clinical context or the nuances of preclinical experimental design that determine evidence quality. Building a complete picture requires pulling from multiple source types and then doing the synthesis work across them.

The synthesis step is where the structural problem gets acute. Bringing together evidence across source types, classifying each piece by quality and relevance, and arriving at an integrated assessment of target confidence is a skilled task that takes significant time. It is also a task that tends to fall at the stage in the discovery cycle when time is shortest: between nomination and resource commitment.

How This Shapes Decision Quality

When evidence synthesis is time-constrained, the signals that surface first are the ones that are easiest to retrieve. High-citation papers in prominent journals. Targets with established biological narratives. Recent preprints that reinforce the nomination hypothesis. This creates a systematic bias toward the visible rather than the complete.

Counter-evidence is the most underweighted signal in a compressed synthesis. A finding from a smaller lab, published in a lower-visibility journal, showing that the same target was non-functional in an in-vivo model of the relevant tumor type, may be perfectly valid and highly relevant evidence. But it will rarely surface in a rapid synthesis pass. The target advances with a more favorable evidence picture than its actual evidence base supports.

This is not a failure of scientific judgment by the discovery team. It is a predictable consequence of the access and synthesis problem. Teams that would make different decisions if they had complete evidence access make the decisions they can make given the evidence they can practically retrieve.

The Scale Issue Is Not Improving on Its Own

The oncology literature grows faster than manual review capacity. A single busy area of solid tumor biology can generate several thousand new publications in a year. A discovery team running target identification cycles on a quarterly cadence is working with an evidence landscape that is materially different every cycle, with no practical way to refresh the synthesis manually at that frequency.

This means evidence assessments made at nomination have a shelf life. A target that was assessed as well-validated two years ago may now have significant counter-evidence in the literature that the original assessment did not reflect. Programs that carry a target forward based on evidence snapshots from early in the program's life are operating with increasingly stale information. This contributes to failures at later stages that look like validation surprises but are actually information lag.

The computational tools that are available for this problem have historically been built around retrieval and not synthesis. A good literature search tool can surface relevant papers more efficiently than a manual search. What it does not do is classify those papers by evidence quality, reconcile contradictory findings, or integrate them across source types into a composite target confidence score. Retrieval and synthesis are different problems. The synthesis piece is where the structural gap remains largest.

What a Structural Solution Looks Like

A meaningful response to the data problem in target selection requires addressing the synthesis bottleneck, not just the retrieval problem. That means pulling evidence from the full range of relevant source types, classifying each piece by experimental system, evidence quality, and tumor specificity, and integrating across source types using a scoring model that reflects the hierarchy of evidence strength.

It also means doing this at the pace of real decision cycles. An evidence synthesis that takes three weeks is still a structural fix if the team's decision window is four weeks. But it is not a fix if the team is running portfolio reviews monthly and needs to refresh assessments for fifteen active targets on that cadence.

At Avenzo, we built the platform around this specific constraint. The architecture is designed to process evidence across millions of source documents, classify and tier that evidence for each candidate target, and return a ranked, source-attributed output that the team can use directly in a target review meeting. The time from query to ranked report is measured in hours, not weeks.

We should be clear about what this does and does not solve. It addresses the access and synthesis problem for the evidence that exists in indexed sources. It does not create evidence that does not exist, and it does not substitute for internal experimental validation. A target can have a high evidence ranking and still fail in the assay phase for reasons that no literature synthesis could predict. The synthesis problem is one lever on decision quality, not all of them.

Where This Matters Most in the Discovery Workflow

The structural data problem has its highest impact at two stages: initial target screening and go/no-go decisions before first-assay commitment. These are the two points where evidence quality most directly determines which targets consume bench resources. They are also the two stages where the time pressure on evidence synthesis is most acute.

Teams that can run a structured evidence assessment at these two stages, rather than relying on compressed manual synthesis, end up with a different distribution of targets in their assay queues. The change is not dramatic in every case. But at the margin, getting the evidence right before commitment is where the difference between a productive and unproductive assay cycle is most often made.

The data exists to make better target decisions. The structural problem has been access and synthesis at decision speed. That is what we built Avenzo to address. If you are working through a target screening cycle and want to understand how the evidence picture changes with structured synthesis, reach out.