Home Blog [Interview] Beyond AI Protocol Writing: How Pharma Can Detect Trial Risks Before Protocol Lock

[Interview] Beyond AI Protocol Writing: How Pharma Can Detect Trial Risks Before Protocol Lock

Szymon Komorowski, Pharma and Life Sciences Advisor at deepsense.ai, discusses how AI can help Clinical Development teams balance scientific ambition with operational feasibility, identify trial risks earlier, and turn fragmented historical knowledge into a governed decision-support capability.

The question “Can large language models write clinical trial protocols?” attracts attention, but it frames the opportunity too narrowly.

A protocol is not simply a scientific and regulatory document, but an operational blueprint that influences who can participate in a study, how demanding participation will be for patients and healthcare professionals, where the study can be conducted, how quickly sites can recruit, and how convincingly the eventual results will support regulatory, reimbursement, and clinical decisions.

This means that the most valuable role for AI may not be generating protocol text faster. It may be helping expert teams detect contradictions, unrealistic assumptions, excessive operational burden, and downstream execution risks before the protocol is locked and the study begins.

Exclusively for us, Szymon Komorowski explains what AI can realistically support today, why private historical trial knowledge matters, and why successful implementation depends more on workflow design, data quality, validation, and change management than on access to another model.

In this interview, we use the term “protocol and feasibility intelligence” as an umbrella term for AI-assisted workflows that help teams review protocol complexity, feasibility assumptions, inclusion and exclusion criteria, enrollment plans, site-selection inputs, and potential amendment risks. It does not describe autonomous clinical decision-making. The objective is to give responsible experts better evidence, earlier warnings, and more systematic access to organizational knowledge.

Is “Can LLMs write clinical trial protocols?” the wrong question?

It is certainly too narrow.

LLMs can already support protocol drafting, summarization, information extraction, and evidence synthesis. But protocol generation is only one part of the problem, and it should remain a human-led process.

The more important question is whether AI can help the people responsible for the protocol make better-informed decisions. Can it reconstruct the key elements of a protocol, convert them into structured data, compare them with internal and external evidence, identify red flags, and show where further clinical or operational review is needed?

That is where protocol review and protocol intelligence come together. You start with a written protocol or a sufficiently advanced draft. The system extracts the relevant attributes, structures them for comparison, evaluates them against own historical data and external sources, and then surfaces potential problems or optimization opportunities. Human experts review those findings and decide what action to take.

Protocol generation is a related but separate workflow. In that case, AI supports the author while the protocol is being created. It can pre-populate sections, retrieve relevant evidence, incorporate prior trial experience, and flag possible issues as the document develops. The human author remains in control throughout.

What makes clinical trial protocol design so difficult?

Protocols are naturally optimized to generate the strongest possible evidence.

The study needs to produce safety and efficacy results that are convincing to regulators such as the FDA or EMA. Those results may subsequently need to support reimbursement discussions with payers and, ultimately, persuade physicians to use the treatment in clinical practice.

That creates pressure to select convincing endpoints and measure them frequently and precisely. Teams may want large, geographically diverse samples that increase confidence in the results. They may also want to divide participants into subpopulations to understand where the treatment works best and to prioritize future regulatory, reimbursement, or commercial efforts.

But many of the decisions that strengthen the scientific design can make the study harder to execute.

Frequent or complex endpoint measurements increase the burden on patients and healthcare professionals. Large target samples may be incompatible with restrictive eligibility criteria. Numerous subpopulations and study arms can fragment the sample, making recruitment more difficult.

A good protocol has to balance the strength of the conclusions with the speed and practicality of execution. Achieving that balance requires considerable attention, particularly as the volume of protocol work increases and organizations are expected to design protocols and begin studies faster.

Which protocol risks are most often discovered too late?

One of the most common problems is a conflict between the eligibility criteria and the required sample size.

The criteria may be too restrictive for the number of participants the study needs to recruit. Alternatively, the target sample may simply be too large if the scientific team wants to preserve those criteria.

Patient and healthcare-professional burden is another frequent issue. A protocol may require too many measurements, visits, or procedures. Some patients may drop out, while others may remain enrolled but fail to complete all required procedures. Their data may then be incomplete for some analyses.

Site selection can create a similar problem. A site that performed exceptionally well in a comparable study last year may be overloaded this year because it has started two other trials in the same therapeutic area.

The competitive context matters. If every sponsor approaches the same institutions because they have the right specialists and patient volumes, some of those studies will inevitably experience delays.

“Too late” usually means that the study has already begun. At that stage, material protocol changes can require additional submissions and approvals, as well as updated documentation and retraining for sites and healthcare professionals. The problem is no longer an observation in a draft. It has become operational rework.

Is pharma genuinely moving from AI experimentation to operational workflows?

Yes and no. Everyone is experimenting, but not everyone is succeeding.

The front-runners have already developed high levels of automation and can estimate how a study is likely to perform while the protocol is still being designed. Even these organizations continue to improve their capabilities.

Their initiatives may include standardizing protocol ontologies to enable more reliable benchmarking against historical data, AI-assisted authoring, and automated red-flag reviews of every protocol for which a CRO receives a request for proposal.

A significant part of the industry, however, is still in an experimental phase. AI is applied to individual tasks, requires manual orchestration, operates only at a limited scale, and must be substantially redesigned when the organization tries to expand it.

The technology itself is not the principal constraint, because the harder problems are implementation, product and architecture decisions, change management, data preparation, and integration with the way teams already work.

Which protocol decisions are most likely to generate downstream rework?

Three areas appear repeatedly.

The first is the relationship between inclusion and exclusion criteria and the required sample size.

The second is the selection of endpoints and the frequency or complexity of their measurement.

The third is the allocation of participants across countries and sites.

Each of these decisions looks reasonable in isolation. The difficulty emerges when they interact with actual patient behavior, healthcare systems, competing studies, local treatment access, and the operational capacity of individual sites.

Consider four examples.

Visit burden in elderly or seriously ill populations

Elderly patients or patients with a high disease burden may not be able to attend a research site frequently. Some may drop out. Others may miss individual assessments, leaving gaps that reduce the completeness of the analytical population.

Remote treatment or procedure administration

Asking a patient to administer a treatment or complete a procedure during a video consultation may appear convenient. But teams must first determine whether the patient is comfortable with remote care, can use the necessary technology, and can safely follow the procedure.

Site capacity changes over time

Historical site performance is important, but it cannot be treated as static. A site that recruited effectively in the past may now be running several competing studies and may no longer have the same capacity.

Reimbursement and treatment access

For an expensive treatment that is already reimbursed in a given country, patient interest in joining a trial may be lower than in a country where the treatment is not yet reimbursed and trial participation provides access to otherwise unavailable therapy.

That difference can materially change country-level enrollment assumptions.

What does a weak AI-based protocol or feasibility review look like?

One warning sign is a high-level LLM review lacking a predetermined workflow, a measurement framework, or a reliable connection to trusted sources.

A generic model may produce plausible observations, but plausibility is not sufficient. Teams need to know which questions were asked, which evidence was used, how outputs were assessed, and where human review is required.

A second problem is the lack of standardized ontologies and dictionaries. If different projects describe the same protocol attributes differently, benchmarking and quantitative analysis become unreliable. Superficial standardization can be just as problematic, particularly when it is performed without a human-in-the-loop process.

The third problem is inconsistency. If every protocol, project, or client follows a different ad hoc process, the organization fails to build a reusable knowledge base. It repeatedly solves similar problems from the beginning instead of compounding what it has already learned.

What can LLMs realistically support around protocols today?

They can support drafting, summarization, extraction, and evidence synthesis.

They can also help structure protocol attributes, compare assumptions, identify inconsistencies, retrieve relevant evidence, and prepare reviewable risk flags for experts.

The boundary for autonomous action remains high. In clinical trials, there is a strong tendency to include human review not only when modifying protocol content, but even when standardizing protocol data for analytical purposes. This helps maintain consistency and accountability.

Over time, selected retrospective analyses of study execution could become more autonomous. For example, an AI system might independently analyze the operational flow of completed studies, provided it does not interpret their clinical results or make clinical decisions.

However, the people who sign off on the protocol remain responsible for it. Authors, review boards, ethics committees, and other designated decision-makers need to understand and approve what is being submitted.

The objective should not be to remove that responsibility. AI should reduce the burden of manually connecting evidence, historical experience, operational inputs, and fragmented data sources, enabling responsible experts to review them more effectively.

What can AI actually do before protocol lock?

There are several realistic use cases:

  • protocol complexity review;
  • inclusion and exclusion criteria operationalization;
  • feasibility risk assessment;
  • enrollment assumption review;
  • historical trial comparison;
  • site selection support;
  • literature review vs future evidence synthesis;
  • identification of protocol elements that may increase the risk of amendment.

AI should not make the final clinical or operational decision. Its role is to structure information, identify patterns, expose assumptions, and direct expert attention to the areas where additional validation is needed.

Which of these use cases can work as an MVP?

When an organization has historical data on which protocols it executed and how those studies performed, practically any of these areas can serve as an MVP.

Historical data creates a point of reference. A new protocol can be compared with earlier trials by examining their enrollment performance, amendments, delays, site results, and other operational outcomes.

Without internal historical data, the best starting points are usually the areas where reliable external information is available. This may include population and epidemiological data, public trial registries, competing studies, and information about trial activity at individual sites.

The availability and timing of detailed trial documents differ by jurisdiction and study stage. The practical implication is that an MVP should be designed around evidence that is genuinely accessible in the chosen geography and therapeutic area, rather than assuming identical data will be available everywhere.

Logical protocol review can also provide value without a historical dataset. A system can still identify contradictions, operational burden, implausible assumptions, or criteria that may be difficult to apply in practice. However, probability estimates will be less precise without a relevant historical benchmark.

Which use cases require internal historical data?

Feasibility risk assessment benefits significantly from a standardized database of historical studies and their performance. Without that reference point, it is difficult to estimate how a new study compares with previous operational experience.

The same applies to enrollment assumption review. Public epidemiological or registry data can help, but accurate probability estimates usually require internal or purchased datasets that reflect actual recruitment performance.

Historical trial comparison requires such data by definition.

Other use cases can begin with logical analysis and publicly available information, although they become more valuable as internal knowledge is added.

Which use case has the greatest business value?

There is no universal answer.

The most valuable use case is the one that addresses the type of problem most likely to create a serious delay or loss in a particular organization’s studies.

In one portfolio, site selection may be relatively straightforward because trials have small samples and broad recruitment criteria. The major problem may instead be a burdensome or invasive endpoint. In that case, optimizing site selection contributes little, whereas improving endpoint strategy or identifying countries with more motivated patients may yield a much greater return.

Every sponsor or CRO should examine the types of trials it conducts most frequently and identify the problems that recur most often. That should determine which capability is built first.

Protocol authoring, protocol review, and site selection should therefore be viewed as parallel opportunities. Their relative value depends on the organization, therapeutic area, portfolio, and available data.

Should pharma talk about “amendment prediction” or “amendment risk flags”?

I do not see prediction itself as problematic, provided everyone understands that it is a prediction rather than a certainty.

Any estimate needs to be interpreted in context. It should show what evidence and assumptions contributed to it and where expert validation remains necessary.

For external communication, “amendment risk flags” is generally the more precise framing. It makes clear that the system identifies protocol characteristics associated with potential downstream rework and directs them to human reviewers. “Prediction” can still be used where an organization has sufficient historical data, a validated methodology, and clearly communicated performance measures.

Why are public benchmarks not enough?

Public protocols, scientific literature, ClinicalTrials.gov, CTIS, epidemiological sources, and external benchmarks are useful. But the most valuable organizational knowledge is often private.

The protocol itself may eventually become public. What normally remains private is whether the study was completed on time, whether it stayed within budget, where enrollment underperformed, which sites caused delays, why amendments were required, and which operational assumptions proved wrong.

A sponsor can use historical protocols, amendments, feasibility assessments, site performance, enrollment data, screen failure patterns, SOPs, and CRO collaboration history from its own trials.

A CRO may hold comparable knowledge across work performed for multiple clients. In that situation, data usage depends on contractual permissions, appropriate anonymization, and ensuring that one client cannot identify another client’s protocols or performance when they are a part of the shared benchmarking dataset.

This private knowledge is where the value compounds. It allows an organization to learn systematically from what it has already done rather than relying only on general market benchmarks.

What is the main obstacle to using that internal knowledge? Ownership, compliance, data, or formats?

Ownership and compliance are often manageable because organizations have processed clinical data for decades and already have relevant contractual and governance structures in place.

Data quality and formats are usually the more difficult problem.

Many existing datasets were not created for AI-supported analysis. Terminology may be inconsistent. Important information may be stored in documents rather than structured fields. Definitions can vary across business units, therapeutic areas, CROs, and time periods.

AI can assist with extraction and standardization, but this cannot simply be handed off to an automated process without supervision. Building the capability to generate reliable datasets for specific analytical objectives is often a separate project.

The organization needs to define what each field means, how it will be validated, how exceptions will be handled, and where human review is mandatory.

Does a client need to share confidential protocols to get started?

Not necessarily. A demonstration or early proof of concept can use a public protocol, a completed study, synthetic data, a fictional trial, or an appropriately anonymized protocol synopsis.

If the objective is to generate analytical data directly from confidential protocols, the documents will eventually have to be made available under the appropriate confidentiality and security arrangements.

However, a client that already has a structured data-generation capability may be able to provide anonymized protocol attributes without exposing identifying details. The required approach depends on the use case and the maturity of the client’s internal data systems.

Should pharma build this capability or buy an existing platform?

This is rarely a pure build-versus-buy decision.

Existing platforms can provide valuable external datasets, benchmarks, and standardized analytics. They may be entirely sufficient for organizations that need a broadly established workflow and do not require extensive adaptation.

A private AI layer becomes more relevant when an organization wants to use confidential internal knowledge, reflect its own protocol-review process, integrate multiple clinical systems, apply organization-specific governance, or avoid making a single model or vendor the permanent center of its architecture.

The most realistic direction is often hybrid: retain established platforms and databases where they already provide value, while building a governed intelligence layer that connects them with internal protocols, historical performance, SOPs, expert decisions, and enterprise systems.

What implementation principles matter most?

The five most important principles are broadly applicable:

  1. Start with the decision workflow, not the model.
  2. Treat private organizational knowledge as a strategic asset.
  3. Use AI for governed decision support rather than autonomous clinical decisions.
  4. Design for integration from the beginning.
  5. Build governance into the MVP.

I would place additional emphasis on validation and MLOps.

Validation is necessary to establish whether the workflow performs reliably for its intended purpose. MLOps is required to keep the system observable and maintainable as models, data, prompts, integrations, and organizational processes change.

Data residency is not necessarily the primary difficulty in this context, as we are primarily discussing protocols and study plans rather than individual patient results. GxP compliance is, of course, required where applicable, but it should be treated as a baseline requirement rather than the only implementation challenge.

The practical challenge is to create a validated system that works consistently inside the actual workflow and can continue to operate as its technical and organizational environment changes.

What mistakes do companies make when implementing AI in Clinical Development?

Many of the challenges are understandable. Clinical Development teams operate under significant time pressure, across complex systems, fragmented data, and strict regulatory requirements. In that environment, it is easy for an AI initiative to begin with a promising model or prototype before the underlying decision workflow has been fully defined.

Data preparation may receive less attention than the visible application. Early pilots may remain isolated from existing systems. Different projects may adopt their own terminology and evaluation methods, while validation, governance, monitoring, and operational ownership are deferred until the solution appears ready to scale.

These are not failures of ambition or expertise. They are signs that moving from an impressive prototype to a reliable clinical workflow requires a broader implementation approach—one built around the five principles we have just discussed.

What is the strongest business case for AI in protocol and feasibility work?

Speed.

The entire competitive game in clinical development is to be ready earlier than competitors without sacrificing quality or precision.

Quality and precision are threshold conditions. A study cannot be made “110% correct.” But the time required to design, review, operationalize, and begin the study can be improved significantly.

A drug can be first to market, or it can arrive 12 months later.

At the same time, speed cannot come from oversimplifying the study. A trial that finishes quickly because its procedures or evidence requirements were weakened may produce conclusions that are less convincing to regulators, payers, and physicians.

The real objective is therefore to increase speed while preserving the strength of the evidence.

Is this still a technically difficult problem?

Many individual components are now relatively straightforward. But they still have to be implemented correctly.

In the hands of an experienced team, a well-scoped project in this area can be closer to painting by numbers than open-ended research. The technical building blocks exist. The workflow can be defined. The system can be evaluated against measurable criteria.

In the hands of a team without the necessary experience, the same initiative can become a risky operation. The organization may produce plausible outputs without reliable grounding, create inconsistent datasets, omit human decision gates, or build a prototype that cannot be validated or scaled.

For a well-selected workflow with a measurable baseline, relevant data, and experienced implementation, the business case should be strong. The value should be evaluated through avoided delays, faster review cycles, improved feasibility decisions, and the ability to reuse historical knowledge—not through a generic promise of cost reduction.

What should pharma leaders take away from this?

AI should not replace clinical expertise, but it should give clinical experts a better operating layer: one that connects fragmented sources, structures historical knowledge, evaluates assumptions consistently, and identifies potential risks before they become delays, amendments, or recruitment problems.

The opportunity is not limited to faster document production. It is to help teams design and prepare trials faster without weakening the evidence they ultimately need to present to regulators, payers, and physicians.

For organizations exploring this direction, the starting point does not need to be a confidential protocol upload or a large platform implementation.

A focused Protocol and Feasibility AI Readiness Assessment can first examine how protocol and feasibility decisions are currently made, where knowledge is fragmented, what historical data is available, which governance mechanisms are required, and which MVP could produce measurable value. It is a first, light sweep, using a narrow set of data to illustrate our points.

From there, teams can evaluate a private protocol risk review, inclusion and exclusion criteria operationalization, historical trial comparison, feasibility risk workflow, or site-selection capability using public, synthetic, anonymized, or appropriately protected data.

The objective is a validated, auditable workflow that helps Clinical Development teams see trial risk earlier and move faster with confidence..