How to Find Relevant Papers for a Medical Literature Review with AI
Linda
Learn how to find relevant papers for a medical literature review with AI, from inclusion criteria and evidence screening to relevance ranking and a prioritized reading list using a real Noah AI workflow.
Finding papers is not the same as finding the right papers.
A search for a biomedical topic can return dozens or hundreds of papers that contain the correct keywords but do not actually answer the research question.
Some may study the wrong patient population. Others may use the right biomarker in the wrong treatment setting. Some are useful reviews but should not replace primary evidence. Others look highly relevant from the title but provide no outcome needed for the review.
For this guide, we tested Noah AI, an AI Agent built for medical and life-science research, on a real paper-selection task.
In adults with resectable non-small cell lung cancer receiving neoadjuvant or perioperative immunotherapy, how well do circulating tumor DNA dynamics predict pathological response, recurrence, event-free survival, or other clinically relevant outcomes?
The goal was not to generate a full literature review. We wanted to build a smaller, defensible set of papers that a researcher would actually want to read and cite.
Quick Answer
A useful paper-selection workflow has four stages:
define relevance → search broadly → screen systematically → prioritize the papers that directly answer the question.
Keyword overlap alone is not enough.
Step 1: Define What “Relevant” Means Before Searching
The first mistake in literature searching is deciding whether a paper is relevant only after seeing the search results.
Instead, define the boundaries first.
For our example, a directly relevant paper needed meaningful overlap across four domains:
| Domain | Requirement in Our Example |
|---|---|
| Population | Adults with resectable NSCLC |
| Treatment setting | Neoadjuvant or perioperative immune-checkpoint inhibitor treatment |
| Biomarker | ctDNA, clearance, persistence, molecular response, or MRD |
| Outcome | Pathological response, recurrence, EFS, DFS, PFS, OS, or molecular recurrence |
A paper that matched only one or two of these domains could still be useful as background, but it should not automatically enter the core evidence set.
Step 2: Write Inclusion and Exclusion Criteria
Before searching, convert the research question into explicit screening rules.
Example inclusion criteria
- Adults with resectable NSCLC.
- Neoadjuvant or perioperative immunotherapy exposure.
- ctDNA or MRD measured before, during, or after treatment.
- At least one relevant pathological, recurrence, survival, or molecular outcome.
- Sufficient publication information to identify and verify the study.
Example exclusion criteria
- Metastatic or unresectable NSCLC only.
- Chemotherapy-only cohorts without separately reported ICI data.
- General liquid-biopsy studies without ctDNA response or MRD outcomes.
- Studies of unrelated cancers.
- General perioperative trial reports without ctDNA results.
- Protocols or registry entries without reported biomarker outcomes.
This step prevents the selection criteria from changing every time an interesting paper appears.
Step 3: Give AI the Research Question — Not Just Keywords
We entered the full research objective into Noah rather than asking for something broad such as:
Find papers about ctDNA and lung cancer.
The actual task specified the population, treatment setting, biomarker, outcomes, evidence hierarchy, inclusion criteria, exclusion criteria, and the final type of output we wanted.
In Noah's current Agent interface, we used:
- Medical & Academia
- Deep Research
- 5.6 Terra · Balanced

A real medical literature-review task submitted to Noah AI with a clearly defined research question and paper-selection objective.
Step 4: Let the Research Process Produce Structured Evidence Files
One useful part of this workflow is that the research process does not have to end as one long chat response.
In this run, Noah generated separate research outputs for:
- database searches,
- professional searches,
- citation records,
- the paper-selection evidence set,
- an outline,
- and the research plan.
This makes it easier to inspect where the evidence came from rather than treating the final narrative as a black box.

Noah AI organized the research run into evidence files, search records, an outline, and a paper-selection output.
Step 5: Screen Papers by Four Domains
Once candidate papers are collected, screen them systematically.
For every paper, ask four separate questions:
Population: Is this really the population in the review?
Treatment: Is the patient receiving the intervention or clinical pathway we care about?
Biomarker: Does the paper actually measure the target biomarker?
Outcome: Does the study report an outcome capable of answering the question?
A paper should not become core evidence merely because the title contains the correct disease and biomarker.
Step 6: Build a Relevance Hierarchy
Not all included papers deserve equal weight.
In the Noah output, the evidence was separated into three practical priority levels:
| Relevance Level | Meaning |
|---|---|
| Essential | Direct, decision-relevant papers with comparatively strong study designs or especially important evidence. |
| High relevance | Directly useful studies with meaningful limitations, such as small samples or retrospective designs. |
| Supporting | Preliminary, methodological, or synthesis evidence that helps interpret the core literature. |

Noah AI separates the evidence landscape into Essential, High relevance, and Supporting papers rather than treating every search result equally.
Step 7: Separate Direct Evidence From Papers That Only Look Relevant
This is one of the most important steps in medical literature screening.
Consider a paper on:
perioperative ctDNA and recurrence in resectable NSCLC
It sounds almost perfect.
But if the cohort was not treated with neoadjuvant or perioperative immunotherapy, it does not directly answer whether ctDNA predicts outcomes in the ICI treatment setting.
It may still be valuable for understanding recurrence or molecular lead time — but it belongs in a different evidence category.

Noah AI separates contextual or methodological literature from the direct clinical evidence set and records the reason for each decision.
Free to use · Free credits included · No credit card required
Find Relevant Medical Papers with Noah AI →
Keyword Match Is Not Evidence Relevance
This distinction is easy to miss when using AI search tools.
Example 1: Right biomarker, wrong treatment setting
A study may evaluate postoperative ctDNA in resectable NSCLC but contain no neoadjuvant or perioperative immunotherapy.
That study can support recurrence biology, but it should not be presented as direct evidence of ctDNA performance during perioperative ICI.
Example 2: Right disease, wrong biomarker
A major perioperative NSCLC trial may report pCR, MPR, EFS, PD-L1, or tumor-mutational burden but no ctDNA result.
It should not suddenly become ctDNA evidence simply because the trial is important.
Example 3: Right topic, wrong evidence type
A review article may summarize the field very well but should not replace the primary prospective or randomized biomarker paper when numerical clinical findings are available.
Step 8: Prioritize Primary Evidence
For clinical questions, we generally want the reading order to start with the studies that most directly answer the question.
In this example, the evidence hierarchy prioritized:
randomized-trial biomarker analyses,
prospective biomarker studies,
other direct primary cohorts,
systematic reviews and meta-analyses,
methodological and guideline sources for interpretation.
Step 9: Create a “Read First” List
Researchers rarely need to read 50 papers in random order.
A more useful output is a short prioritized list that answers:
If I only have time to read five papers first, which five should they be?
For this case, Noah's recommended set contained 10 core and supporting sources, with five papers prioritized for first-pass reading.
| Priority | Paper | Why Read It First? |
|---|---|---|
| 1 | Biomarkers of nivolumab benefit in resectable non-small cell lung cancer | Randomized perioperative evidence linking ctDNA dynamics with pathological response, recurrence, and EFS. |
| 2 | Minimal Residual Disease Enhances Prognostic Stratification beyond Pathologic Response in Resectable NSCLC | Important for understanding how postoperative MRD adds information beyond pCR. |
| 3 | Overall Survival and Biomarker Analysis of Neoadjuvant Nivolumab Plus Chemotherapy in Operable Stage IIIA NSCLC | Prospective evidence linking ctDNA status with PFS and OS. |
| 4 | Circulating Tumor DNA Is Associated With Pathologic Response and Survival Outcomes in NSCLC Treated With Neoadjuvant Immunotherapy | Provides a broader systematic synthesis and highlights heterogeneity across the evidence base. |
| 5 | ESMO recommendations on the use of circulating tumour DNA assays for patients with cancer | Helps distinguish prognostic association from clinical utility and treatment-guiding evidence. |
Step 10: Read the Highest-Value Papers in the Right Order
Reading order matters.
A practical sequence is:
strongest direct study → major validation study → additional prospective evidence → systematic synthesis → guideline or methodological context.
This gives the researcher a working model of the evidence before moving into weaker or more peripheral studies.
Step 11: Verify Metadata Before Adding a Paper to the Review
AI-generated literature lists should not be copied directly into a review.
Verify:
- exact title,
- authors,
- journal,
- publication year,
- PMID or DOI,
- study design,
- trial identifier where relevant,
- and whether the publication is full text, abstract-only, or protocol-only.
This matters particularly for very recent evidence, conference abstracts, and ongoing trials.
Step 12: Distinguish Three Very Different Evidence Claims
Biomedical papers often use similar language for fundamentally different claims.
| Claim | What It Means |
|---|---|
| Prognostic | Patients with different biomarker states have different outcomes, regardless of which treatment caused the difference. |
| Predictive | Biomarker status modifies the relative benefit of one treatment versus another. |
| Clinical utility | Changing treatment because of the biomarker result actually improves patient outcomes. |
These should not be treated as interchangeable conclusions.
Step 13: Use Background Papers for the Right Purpose
Excluding a paper from the core evidence set does not mean the paper is useless.
Background and methodological papers can help with:
- assay sensitivity,
- ctDNA detection limits,
- sample timing,
- definitions of molecular residual disease,
- pathological-response assessment,
- and interpretation of clinical utility.
The key is to label their role correctly rather than mixing them with direct treatment-specific evidence.
A Repeatable 7-Step Workflow
The same approach can be reused for other medical literature reviews.
Define the research question and eligibility rules. Specify population, intervention or exposure, biomarker, outcomes, and evidence cutoff.
Search biomedical databases and authoritative sources. Start broad enough to avoid missing important terminology or named trials.
Deduplicate and verify records. Confirm titles, publication type, journal, year, PMID, DOI, and trial identifiers.
Screen each study across the core relevance domains. Do not rely on title or keyword overlap alone.
Extract the details needed to judge the evidence. Capture population, design, assay, timing, endpoints, results, and important limitations.
Classify and prioritize. Separate direct, partial, background, and excluded evidence, then rank the direct set.
Update and synthesize cautiously. Keep protocols, conference abstracts, ongoing trials, and mature publications clearly separated.
Where AI Helps Most
AI can reduce the manual work involved in:
- turning a research question into explicit screening criteria,
- organizing candidate studies,
- checking whether all required relevance domains are present,
- separating direct evidence from contextual evidence,
- ranking papers by relevance,
- building a reading order,
- and identifying evidence gaps.
But relevance decisions still need to be inspectable.
A useful AI workflow should show why a paper was included, downgraded, retained only as background, or excluded.
What AI Should Not Do for You
Researchers should not accept a paper simply because an AI system labels it relevant.
Always verify:
- that the publication exists,
- that the population matches the review question,
- that the study actually contains the biomarker or outcome claimed,
- that numerical findings come from the correct paper,
- and that conference or registry evidence is not presented as mature peer-reviewed evidence.
Frequently Asked Questions
How do I find relevant papers for a literature review?
Define your research question and inclusion criteria first, search broadly, screen each paper against the required population, intervention or exposure, outcomes, and study design, then prioritize the studies that most directly answer the question.
Can AI find papers for a medical literature review?
AI can help discover, organize, screen, and prioritize biomedical papers. Important publications and study details should still be verified in the original database and article.
How many papers should I include?
There is no fixed number. The goal is not to collect the largest possible bibliography but to capture the evidence necessary to answer the research question without filling the review with redundant or tangential papers.
Should review articles be included?
Reviews are useful for orientation, terminology, reference discovery, and synthesis. For numerical evidence and study-level claims, researchers should usually return to the underlying primary studies.
How do I know whether a paper is directly relevant?
Check whether it meaningfully matches the required population, clinical setting or intervention, biomarker or exposure, and relevant outcome. Keyword similarity alone is not sufficient.
What is the difference between an important paper and a relevant paper?
A landmark paper may be important to the broader field but still be only background evidence for a narrowly defined review question. Relevance depends on how directly the study helps answer that specific question.
Final Takeaway
The goal of a medical literature search is not to accumulate papers.
It is to build a defensible evidence set.
In practice, that means moving from:
“Does this paper contain my keywords?”
to:
“Does this study's population, treatment setting, biomarker, outcome, and design actually help answer my research question?”
In our real Noah AI workflow, papers were not simply retrieved. They were separated into direct, partial, background, and excluded evidence, ranked by relevance, and converted into a practical reading order.
That is a more useful endpoint than a long list of search results: a smaller set of papers with an explicit reason for why each one matters.
Free to use · Free credits included · No credit card required
Find Relevant Medical Papers with Noah AI →
Selected Papers From This Example
- Biomarkers of nivolumab benefit in resectable non-small cell lung cancer
- Minimal Residual Disease Enhances Prognostic Stratification beyond Pathologic Response in Resectable NSCLC
- Overall Survival and Biomarker Analysis of Neoadjuvant Nivolumab Plus Chemotherapy in Operable Stage IIIA NSCLC
- Circulating Tumor DNA Is Associated With Pathologic Response and Survival Outcomes in NSCLC Treated With Neoadjuvant Immunotherapy
- ESMO Recommendations on the Use of Circulating Tumour DNA Assays for Patients With Cancer