All articles Technology

Literature Mining at Scale: How Avenzo Processes Target Evidence Across Thousands of Papers

8 min read Ravi Nair
Literature Mining at Scale: How Avenzo Processes Target Evidence Across Thousands of Papers

The published oncology literature grows by roughly a million new papers per year across the major biomedical databases. For target identification specifically, the relevant portions of that literature span molecular biology, pharmacology, tumor biology, clinical context reporting, and an expanding body of computational analysis publications. No research team can manually track this volume in real time. The question is not whether to automate parts of the evidence retrieval process, but which parts to automate and how to do it well.

This post describes our approach to large-scale literature mining for oncology target evidence, including where it works well and where it does not.

The Core Challenge: Evidence Is Embedded in Prose

The fundamental difficulty in biomedical literature mining is that the evidence we need is not stored in structured format. A paper that describes a loss-of-function experiment in a tumor model, demonstrating that a specific target is necessary for tumor cell survival, does not arrive pre-labeled with fields like "evidence_type: in_vivo_functional," "model_system: syngeneic mouse model," and "outcome: reduced tumor growth." That information is embedded in the methods and results sections, written in natural language, and requires contextual understanding to extract correctly.

Early biomedical text mining systems addressed this with keyword matching and pattern extraction: look for combinations of gene names, experimental terms, and outcome language in the same sentence or paragraph. This produces reasonably good recall for well-defined evidence types but poor precision, because natural language is ambiguous and biomedical writing is convention-laden in ways that simple pattern matching does not handle well. A sentence that says "X was not required for tumor growth in this model" contains nearly the same keywords as "X was required for tumor growth in this model" but carries opposite meaning for evidence quality purposes.

The approaches that do better on precision require deeper linguistic processing: semantic role labeling to identify the experimental agent, the action, the biological object, and the outcome; negation detection to handle "X was not required" correctly; co-reference resolution to track that "the target" in paragraph five refers to the same gene mentioned by name in paragraph two; and named entity recognition fine-tuned for biomedical vocabulary to correctly identify the gene, tumor type, and model system references throughout the paper.

What We Built and Why

The Avenzo text mining architecture was designed from the start around the specific task of evidence quality classification, not just evidence retrieval. Retrieval solves the problem of finding papers that mention a target. Classification solves the problem of determining what those papers actually tell us about whether the target is worth pursuing.

For each paper, we extract: the experimental system type (cell-line, in-vivo mouse model, patient-derived model, organoid, or other), the directional outcome (positive for target relevance, negative, or ambiguous), the tumor type specificity of the finding (relevant to the queried tumor type, a different tumor type, or unspecified), and the replication context (first published finding, confirmed replication by an independent group, or internal replication within the same study).

These extracted fields are what feed the evidence quality tier classification. A paper that contributes a confirmed in-vivo replication finding for a relevant tumor type classifies as Tier 1 evidence. A paper that contributes a single cell-line finding with ambiguous outcome language classifies differently. The text mining extraction accuracy determines how well this tier classification reflects the actual content of the paper.

Getting extraction accuracy to a level that produces useful output required extensive annotation work: domain-expert-labeled examples of evidence extractions across hundreds of oncology papers, used to fine-tune the extraction models for oncology-specific language and experimental conventions. The model's behavior on novel papers is only as reliable as the annotation examples it learned from, which is why the annotation work is ongoing rather than a one-time build.

What We Get Right and What We Miss

The extraction pipeline performs well on full-text papers from major oncology journals where the experimental methods sections follow recognizable conventions and the results language is relatively standard. It performs less well on older papers where experimental reporting conventions were different, on papers from certain non-English-origin journals that have been machine-translated, and on preprints where the writing has not gone through editorial review and may use less-standard notation.

We also miss evidence that is presented in supplementary figures and tables that are not full-text indexed in the databases we query. A substantial fraction of functional validation data appears in supplementary materials, particularly for papers published in high-tier journals where the main figures are reserved for the central narrative and supporting data is in the supplement. Our current architecture does not fully process supplementary figure legends and tables, which means we undercount functional evidence in high-visibility papers relative to lower-visibility ones. We are working on this but it is not solved.

Ambiguous outcome language is a persistent source of error. Phrases like "our results suggest," "is consistent with a role for," and "may contribute to" are common in oncology writing and reflect legitimate scientific hedging. The extraction model needs to classify these as lower-confidence positive signals rather than as definitive positive or negative findings, and it does not always do this correctly. In practice, we apply a confidence score to extracted evidence statements that downstream affects the quality tier classification, with uncertain extractions contributing less weight.

Scale and Currency

The corpus we index spans over forty million papers and preprints, with daily ingestion of new publications from the major biomedical preprint servers and journal feeds. For a given target query, we pull and process evidence from the relevant subset of this corpus, which for a well-studied oncology target may be several thousand papers.

Currency matters for this application. A target's evidence base can change materially over six to twelve months as new publications arrive. The Avenzo platform reflects the evidence state at the time of query, which we note explicitly in the output. An evidence assessment run six months ago may not reflect current literature for a rapidly-moving target area. We recommend reassessing high-priority targets at regular intervals in an active discovery program, not treating a single assessment as a permanent characterization.

The Human-in-the-Loop Requirement

Automated text mining at this scale produces a lot of correctly classified evidence and some incorrectly classified evidence. The incorrect classifications are not always obvious from the output alone. A domain expert looking at the evidence report for a target they know well will sometimes identify classification errors: a paper they know describes a negative finding that appears in the positive evidence count, or a tumor-type-specific finding that has been incorrectly generalized.

This is why we designed the Avenzo output to show evidence source attribution and classification reasoning, not just aggregate scores. When a target scores unexpectedly high or low, the team can drill into the specific papers driving that score and verify whether the classification is correct. Catching errors is possible because the reasoning is visible.

We consider this transparency a design requirement rather than an optional feature. An opaque score with no audit path puts the team in the position of either trusting or rejecting the output wholesale. A score with visible source attribution and classification rationale lets the team act as the quality check on the automated analysis, which is the right role for domain expertise in this workflow.

If you are interested in seeing how large-scale literature processing translates into target evidence rankings for your specific candidates, reach out to discuss early access to the platform.