All articles Company

The Data Problem in Oncology: Why We Started Avenzo

8 min read Athena Countouriotis

Before starting Avenzo, I spent several years working in oncology drug development, initially in early discovery and later in clinical-stage programs. The work I found most frustrating, not in a discouraging way but in a this-should-not-be-this-hard way, was target selection. Specifically, the moment when a team sits down to decide which of a set of candidates to invest real investigation resources in, and realizes that the evidence picture for each candidate is assembled from a combination of memory, informal literature familiarity, and whatever the scientist most recently assigned to that target happened to read.

That is not a criticism of discovery scientists. It is an observation about the information problem they face. The evidence that informs target selection is distributed across an enormous and growing literature, across internal assay systems with their own unwritten history, across pathway databases that do not talk to each other, and across the personal knowledge of researchers who have changed employers, retired, or published a fraction of what they know. No one person or team can hold all of it. And the synthesis capacity that would let you assemble it rigorously and consistently, at nomination speed and across a full candidate list, does not exist as a standard workflow tool. It exists as the output of individual effort applied unevenly across candidates.

This is the specific problem Avenzo was built to address.

The Pattern I Kept Seeing

The failure pattern that crystallized the problem for me was not a single catastrophic case. It was a pattern that I saw variations of repeatedly: a target advances to an investigation cycle, accumulates significant assay and chemistry investment, and then fails at a mechanistic challenge that was visible in the literature from the beginning, but was buried in a paper that did not make it into the standard narrative around the target.

This happens because target selection conversations default to a different kind of evidence than what is most relevant for the nomination decision. Teams discuss genomic alteration frequency because that data is easily retrievable from public databases. They discuss the landmark papers in the space because those papers are the ones that shaped the field's understanding of the target. They discuss the internal assay data they have. What they discuss less is the full quality distribution of the functional evidence: whether the in-vivo work exists, whether it has been independently replicated, whether there are published results from non-standard models that were done precisely because someone wanted to stress-test the finding and did not see the result they expected.

The landmark papers got cited because they were landmark. The stress-test papers got cited less. Both are in the literature. Only one tends to make it into the standard evidence narrative that informs a nomination meeting.

Why This Is Harder Than It Looks

The obvious response to this problem is: do better literature review. Read more papers. Be more systematic. I understand why that response is intuitive, and I spent years trying to operationalize it at the team level. The problem is not motivation or rigor. It is structural.

A discovery team doing serious work in an oncology therapeutic area is tracking a moving literature that grows by hundreds of relevant papers a year. The scientists doing that tracking are also running experiments, interpreting results, attending program meetings, and doing the work of actual discovery. The bandwidth available for comprehensive evidence synthesis at the moment a target nomination decision is being made is limited by definition. Even a team that genuinely prioritizes evidence quality will be making nomination decisions with a literature synthesis that is current to whenever each scientist last had time to read, not current to today.

The second structural problem is consistency. When a scientist synthesizes evidence for a target they have been tracking closely, they naturally bring pattern recognition from everything they have read in the space. That pattern recognition is a genuine scientific asset. It is also unevenly distributed across a candidate list. Targets that have an internal champion who has been tracking them closely get a different quality of evidence synthesis than targets that are being evaluated for the first time by someone who is doing the literature review under time pressure. Nomination decisions made from this uneven synthesis basis are systematically biased toward targets with internal advocates, independent of the actual evidence quality distribution.

What We Decided to Build

When Ravi and Naomi and I started talking through what kind of company could address this problem, we had a long conversation about the difference between a tool that helps scientists search and read the literature faster versus a tool that produces a structured evidence quality assessment. The first category is useful for many things. It is not sufficient for the specific decision problem we had identified.

What makes target nomination hard is not that the relevant papers are hard to find with a good search. It is that once you have the papers, you have to classify what each one says in evidence quality terms, aggregate across hundreds of papers with different experimental systems and outcome measures, weight by quality not just volume, and produce an assessment that reflects what the evidence actually supports at a claims level that matches what you are trying to decide. That is not a search problem. It is a structured synthesis problem.

We built Avenzo around the evidence quality tier framework because it directly encodes the decision-relevant classification. The question a nomination decision is trying to answer is not how many papers exist about a target. It is whether the evidence supports the mechanistic hypothesis that the target is functionally relevant to tumor survival in the specific tumor type being targeted, and whether that support comes from the kinds of experimental systems that are meaningful predictors of clinical relevance. The tier framework structures the synthesis around exactly those questions.

What We Are Not Trying to Do

There are adjacent problems in drug discovery that we are not trying to solve with the current platform. We are not trying to predict druggability or selectability: whether a target can be addressed with a small molecule, biologic, or cell-based approach is a chemistry and structural biology problem that is separate from whether the target should be pursued at all. We are not trying to integrate internal wet-lab data streams in real time, though that is a direction we are thinking about for later phases of the platform. We are not trying to build a clinical prediction model.

The scope we are working in is the evidence quality picture for solid-tumor targets from the public biomedical literature and pathway databases, at the point of nomination and during active investigation cycles. That scope is narrow enough to do well and broad enough to address the specific structural problem that led us here. Staying in that scope is a deliberate choice.

Why We Are Starting With Solid Tumors

The solid tumor focus is not arbitrary. Solid tumors have the most complex heterogeneity problem at both the target biology and evidence synthesis level. They also represent the bulk of unmet clinical need in oncology. The challenges that make target identification difficult in solid tumors, intratumoral heterogeneity, microenvironment interactions, model system translation gaps, are exactly the challenges where evidence quality assessment at the nomination stage makes the most difference.

There are also a large number of solid tumor types where the biology is well enough characterized that the pathway databases and functional genetics literature are genuinely informative, but where systematic evidence synthesis capacity does not exist in most discovery organizations. That gap between available information and synthesis capacity is where we can add value now, in the current state of the science and current state of discovery team bandwidth.

What We Have Learned Since Starting

Building the platform has taught us things we did not fully anticipate. The evidence extraction problem is harder than the classification problem: getting the machine to correctly identify what a paper reports in terms of experimental system, directionality, and outcome specificity, in the variety of writing styles and experimental reporting conventions across the oncology literature, required more annotation work than our initial estimates. The pathway context dimension of the scoring model was more useful than we initially weighted it. And the teams we have talked with most extensively are, in many cases, less interested in having a ranked list and more interested in having a structured description of the evidence landscape that they can reason about alongside their own knowledge. The output has evolved in that direction.

None of this is a surprise in retrospect. Building in a domain with this much scientific complexity means learning from the people doing the work, not just from the problem framing you start with. We are still learning.

If you are working in solid tumor discovery and the evidence problem I have described here sounds familiar from your own experience, we would like to talk. The platform is in early access; reach out to discuss whether and how it fits your program context.