Drug Discovery and Development

  • Home Drug Discovery and Development
  • Drug Discovery
  • Women in Pharma and Biotech
  • Oncology
  • Neurological Disease
  • Infectious Disease
  • Resources
    • Video features
    • Podcast
    • Views
    • Webinars
    • PharmSci360
  • Pharma 50
    • 2026 Pharma 50
    • 2025 Pharma 50
    • 2024 Pharma 50
    • 2023 Pharma 50
    • 2022 Pharma 50
    • 2021 Pharma 50
  • Advertise
  • SUBSCRIBE

As the dream of the autonomous clinical trial dawns, agent supervision is today’s human job

By Brian Buntz | August 27, 2026

[Modified based on Adobe Stock image]

In 2020, the average phase 3 protocol collected approximately 3.56 million data points. By 2025, that figure had reached about 5.96 million, a five-year increase of 67% and 6.4 times the 2012 average of 929,203, according to collaborative research from the nonprofit industry group TransCelerate BioPharma and the Tufts Center for the Study of Drug Development. The study found potential room for greater discipline as nearly one-third of phase 3 procedures and associated data were classified as non-core or non-essential. The peer-reviewed paper also noted that AI and ML processing power “may be serving as a disincentive to reduce data volume.”

Nearly eight months later, that observation about AI looks prescient. “I don’t think [sponsors] are as worried about collecting too much data,” said Venu Mallarapu, chief transformation and AI officer at eClinical Solutions. He argued that AI has made expanding datasets easier to manage and analyze, while sensors and wearables could drive volumes higher still. “Of course, you still want to make sure that you’re collecting the right data,” Mallarapu said. “But crunching that data and generating the insights that would be required, AI has certainly made that a lot easier with the powerful capability of some of the models.”

Venu Mallarapu,

Venu Mallarapu

Where agents are landing first

The data processing layer is one layer; the other lies in the complexity of collecting the data, whether it is staff, say, collecting a patient’s blood, running the imaging or sitting with the patient through the questionnaire. As clinical trial complexity has grown, so has burnout among clinical research staff. “The sites are understaffed and overwhelmed,” said Janice Chang, CEO of TransCelerate BioPharma, in a January interview. She said site staff report being overburdened by administrative work including training, paperwork and contracts.

At eClinical Solutions, the initial agent deployments are concentrated in the data-processing side of that workload. Mallarapu said an average customer may have three to five defined AI use cases within data review and reporting. Most remain assistive, with an agent drafting or triaging material and a person taking the action. “The most advanced ones are running scoped agents inside a workflow which has been reengineered to take advantage of the agents,” he said.

That deployment pattern aligns with the broader industry picture. In an April poll of roughly 300 life sciences professionals at a Pistoia Alliance conference, 30% of organizations claimed enterprise-wide AI implementation, with reported value concentrated in regulatory and reporting work.

The latest clinical-trial numbers are broadly similar. In an Everest Group survey of 200 senior pharma, biotech and CRO decision-makers commissioned by clinical trial software company Medidata, one-third of organizations reported using AI in a handful or a majority of their trials, while 82% of those using it in trial operations had 18 months or less of experience. A total of 63% said they mandate human oversight. The survey measured AI use broadly rather than agents specifically, and respondents reported the strongest above-expectation results in bounded applications such as task and workflow automation (46.5%), data cleaning (40.5%) and query resolution (36.5%).

Dr. Pamela Tenaerts

Dr. Pamela Tenaerts

Another example comes from clinical trial technology company Medable, which built a clinical monitoring agent to help clinical research associates work across data scattered among more than a dozen systems. Its chief medical officer, Pamela Tenaerts, offered the limits of that approach. “I would argue protocol optimization still needs to happen,” Tenaerts said. An agent may be able to help a monitor navigate the millions of data points in typical phase 2 and 3 protocols, yet its processing capacity leaves the scientific justification for collecting them untouched. “We should figure out a way to decrease the numbers,” Tenaerts said of the tendency to focus on data volume for its own sake.

Agents are also changing the economics of data collection. “For a while, the data volume was sort of capacity limiting,” said Ken Getz, executive director of the Tufts Center for the Study of Drug Development and lead author of the data collection study. An agent can absorb routine data management work and free clinical trial professionals to concentrate on higher-priority decisions, he said, giving sponsors a more efficient way to handle rapid growth in volume. The paper he led had identified that same processing capacity as a possible disincentive to collecting less.

Clinical-trial quality guidelines already treat prioritization as a design requirement. ICH E8(R1), the international guideline titled “General Considerations for Clinical Studies,” says critical-to-quality factors should remain clear and “uncluttered” by extensive secondary objectives, processes and data. Yet deciding what to omit becomes more difficult when sponsors want to preserve information for questions they have yet to formulate. Mallarapu said sponsors collect data for primary and secondary endpoints along with possible future needs. “Or, we’ll collect that data just in case,” he said.

Ken Getz

Ken Getz

For now, contextualizing the data is largely a human affair. “The CRA [Clinical Research Associate] has to put all that data together in context and decide what to focus on,” Tenaerts said. But agents are already freeing up their attention. Agents can help with processing data, making connections between diverse data points, freeing up the human to have a global view of the trial, “so you no longer have to have 13 tabs open on your computer,” she said. Such workflows make it easier to focus on addressing a potential problem rather than manually hunting for the data to diagnose the problem in the first place. The agent “tells you things like, ‘We’ve noticed here in the EDC [electronic data capture] that there’s a new medication started, but in your safety system I don’t see an adverse event at the same time.’” Tenaerts said. “The agent is saying there’s a discrepancy, you may want to go check that out.”

Delegating the data work

In other words, more sponsors are using agents to deputize more data collection and analysis tasks. “You can watch the agent do it, but you don’t always watch what the agent does. You can go to some other activity and then it comes back to you with, ‘This is what I found,’” Tenaerts said. “Or at the beginning of the day you can ask, ‘What are the most pressing issues at my site, what do I need to pay attention to?,’ and it does all that work for them.”

Data architecture is another core consideration. What gives humans and agents a coherent view across all those diverse data sources? Spreadsheets remain popular as a working layer because they are familiar and flexible. There’s a reason for the enduring popularity of spreadsheet tools like Excel in clinical trials, Mallarapu said. Users can write macros, annotate decisions and “drop it in SharePoint, share it with people,” he added. But such workflows can sometimes be messy, and keeping track of the most recent version of the data can be complicated. “You don’t know how many copies of the same data exist in enterprises today,” Mallarapu noted.

Fragmented data create extra work for people and agents because each workflow must identify the authoritative source, connect related records and preserve traceability. Mallarapu said companies are increasingly using Snowflake, Databricks and similar enterprise platforms to provide a central data lake or lakehouse, govern access and trace how information is used. Such an architecture keeps the underlying data queryable while letting a person or agent pull a defined slice and send the result into Excel or another familiar interface for review. The same architecture supports the clinical monitoring agent’s consolidated view: connected data underneath, with a focused set of findings presented to the CRA. “If you are introducing a system or automation, you need to think about reengineering the process as well, so that the process is tweaked to make use of the new system or automation that you are bringing in,” Mallarapu said.

When processing scales faster than judgment

While the ability to query and process dispersed data streams has evolved considerably with cloud data warehouses, lakehouses and AI-assisted analysis, dealing with massive troves of data remains something of an unsolved problem, decades after “Big Data” emerged as a buzzword.

A recent reminder comes from METR’s investigation of an OpenAI cybersecurity evaluation in which roughly 1,200 agent instances used a shared, unsanctioned message board to coordinate, and approximately 700 participated in an attack on Hugging Face. In making sense of that incident, METR examined roughly 1,300 transcripts recording the actions and reasoning of individual agent runs, along with more than 70,000 messages and files posted by the agents. Faced with that volume, the researchers turned to GPT-5.6 Sol analysis agents, which METR described as “often managing large nested trees of sub-agents.” “We heavily delegated our analysis to often-unreliable AI agents,” they wrote. Those analysis agents generated well over 1,000 pages, yet often failed to highlight the most important findings and made errors that researchers caught later. METR said a human researcher given enough time would likely have produced more calibrated and useful analysis. Completing the investigation manually within the available time would have been “completely infeasible,” they said.

METR surveyed 349 technical researchers, engineers and managers about AI’s contribution to their work. Depending on the question, respondents estimated that AI made their early 2026 output 1.6 to 2.1 times more valuable and anticipated a 2.9-fold contribution by March 2027. The findings reflect self-reported perceptions. Image credit: METR, CC BY

METR surveyed 349 technical researchers, engineers and managers about AI’s contribution to their work. Depending on the question, respondents estimated that AI made their early 2026 output 1.6 to 2.1 times more valuable and anticipated a 2.9-fold contribution by March 2027. The findings reflect self-reported perceptions. Image credit: METR, CC BY

The incident agents, for their part, treated one another as the relevant authority. Once agents reached the message board, METR found, they received requests and assignments from other agents “that they may have taken to be instructions.”

Still, METR’s broader research places “often unreliable” agents within a steadily improving trend. Its May 2026 survey of 349 technical workers found median self-reported value gains from AI tools roughly doubling year over year, though METR cautioned that self-reports may overstate real productivity.

There are also core differences between METR’s analyzed workflows and regulated clinical ones. The agents METR studied had been set loose on an effectively impossible benchmark task inside a sandbox with cybersecurity classifiers switched off, running for days with a strong incentive to defeat an automated scorer they had wrongly inferred existed, and with no reviewer assigned to any individual action. Clinical agents by necessity are far more constrained. They operate under a protocol, a statistical analysis plan and a monitoring plan, with human oversight at defined decision points. The narrower point survives those differences: the assumption that a human can meaningfully review what an agent proposes is an assumption about human attention, and there is now a documented case of that attention degrading under volume.

Human oversight across layered workflows

The agentic trajectory, meanwhile, may not remain unique to frontier labs. Asked whether agents could eventually connect clinical data back to early-stage preclinical work, Getz described the increasingly autonomous capabilities of AI models. “It’s sort of agents that are managing agents, right?” Getz said. “It’s somewhat analogous to what’s happening on the internet today, where very few people are actually visiting websites. They’re relying on AI engine optimization to interact with AI optimization in the searching. It’s really fascinating where this may ultimately lead.”

Medable’s CEO, Michelle Longmire, has compared agents to the self-driving car. “There’s going to be a time where a lot of this stuff is stitched together by agents, and we will need agents on the site side as well,” said Tenaerts, echoing Longmire’s vision. Tenaerts said the abstract idea of an autonomous clinical trial is “certainly intriguing,” even if not a reality today, and that the concept poses a host of questions. “You wonder about human-in-the-loop oversight and all those kinds of things,” she said.

For now, the clinical vendors describe a ceiling. “Most of what the agent does isn’t autonomous,” Tenaerts said. Similarly, Mallarapu said he is unaware of any eClinical customer running fully autonomous agents in clinical trials that both process data and make decisions without human involvement. “There is a human in the loop all the time, and that is not going away anytime soon, the way I see it,” he said. Medable’s perspective was similar.

What keeps the human in that loop relates to both accuracy and architecture. “A good way to look at this is almost like agent proposes and human disposes,” Mallarapu said. “No agent directly takes an action without human approval.” Everything along the way is logged, he said: what triggered the agent, on whose behalf it acted, the input data and its version, the program and the output, which agent and which version produced it, and the human decision with its rationale and timestamp. “This gives the overall process the attributability built into the audit trail, the traceability, the transparency that’s required,” he said.

Image credit: METR, CC BY

Clinical trials carry distinct regulatory requirements, data structures and patient-safety stakes. They share with engineering a growing reliance on AI for bounded technical workflows under human review. METR’s comparison also highlights a measurement challenge relevant to both fields: recent self-reported productivity gains substantially exceed results from controlled experiments. Image credit: METR, CC BY

Swarms of agents are already a reality in clinical trials

eClinical’s Data Advisor “actually starts as one agent, but it has a swarm of agents working behind the scenes,” Mallarapu said. The user converses with a single advisor while the swarm works underneath, and the human expert remains “the one in fact taking the action, an auditable action which can stand the scrutiny of regulatory review.”

In eClinical’s SDTM [Study Data Tabulation Model] mapping automation, agents read the raw data and the mapping specification, write the transformation programs and execute them. “But it presents to the user a screen where it has clearly articulated what its confidence is on the transformations or the mappings that it has done,” Mallarapu said. “Then the human, based on the confidence, is spending more or less time on a given transformation.” The progression is consistent across these examples: the human prioritizes, the agent works while the human attends elsewhere, and agent-reported confidence helps determine how much scrutiny each output receives.

Tenaerts pointed out that automation capacity in one area tends to highlight a bottleneck somewhere else. For instance, if sponsors get dramatically more efficient at sending queries and emails to trial sites, the sites absorb the load. “If you dump all that stuff on them, that’s still a bottleneck, right?” she said. “You push the balloon and it goes somewhere else. You need to figure out the whole system.”


Filed Under: clinical trials, Drug Discovery

 

About The Author

Brian Buntz

As the pharma and biotech editor at WTWH Media, Brian has almost two decades of experience in B2B media, with a focus on healthcare and technology. While he has long maintained a keen interest in AI, more recently Brian has made making data analysis a central focus, and is exploring tools ranging from NLP and clustering to predictive analytics.

Throughout his 18-year tenure, Brian has covered an array of life science topics, including clinical trials, medical devices, and drug discovery and development. Prior to WTWH, he held the title of content director at Informa, where he focused on topics such as connected devices, cybersecurity, AI and Industry 4.0. A dedicated decade at UBM saw Brian providing in-depth coverage of the medical device sector. Engage with Brian on LinkedIn or drop him an email at [email protected].

Related Articles Read More >

Tufts model estimates 82x ROI and up to $21 million in value from Medable AI agent
How Sheba Medical Center became OpenAI’s first international hospital partner
Novo Nordisk in the Drug Discovery & Development Pharma 50
Novo Nordisk’s ziltivekimab misses Phase 3 primary endpoint 
Moderna bets on mRNA’s second act with cancer, autoimmune programs and AI research platform
“ddd
EXPAND YOUR KNOWLEDGE AND STAY CONNECTED
Get the latest news and trends happening now in the drug discovery and development industry.

MEDTECH 100 INDEX

Medtech 100 logo
Market Summary > Current Price
The MedTech 100 is a financial index calculated using the BIG100 companies covered in Medical Design and Outsourcing.
Drug Discovery and Development
  • MassDevice
  • DeviceTalks
  • Drug Delivery Business News
  • Medtech100 Index
  • Medical Design Sourcing
  • Medical Design & Outsourcing
  • Medical Tubing + Extrusion
  • Pharmaceutical Processing World
  • R&D World
  • Subscribe to our E-Newsletter
  • About Us
  • Contact Us

Copyright © 2026 Arrowfly LLC. All Rights Reserved. The material on this site may not be reproduced, distributed, transmitted, cached or otherwise used, except with the prior written permission of Arrowfly
Privacy Policy | Advertising | About Us

Search Drug Discovery & Development

  • Home Drug Discovery and Development
  • Drug Discovery
  • Women in Pharma and Biotech
  • Oncology
  • Neurological Disease
  • Infectious Disease
  • Resources
    • Video features
    • Podcast
    • Views
    • Webinars
    • PharmSci360
  • Pharma 50
    • 2026 Pharma 50
    • 2025 Pharma 50
    • 2024 Pharma 50
    • 2023 Pharma 50
    • 2022 Pharma 50
    • 2021 Pharma 50
  • Advertise
  • SUBSCRIBE