Harlan Robins, co-founder and former chief scientific officer of Adaptive Biotechnologies, has raised $15 million for a new Bellevue startup that aims to let research groups profit from the proprietary datasets they build for AI training. Fuse and Cercano Management led the round.
The company, called Harell Data, positions itself as a secure bridge between data creators and machine learning teams. It launches an early access platform this week with initial datasets from Seattle-based A-Alpha Bio and Adaptive Biotechnologies.
"The goal of the company is to connect AI modelers to proprietary training sets to enable the solution of challenging scientific problems," said Robins in an email to GeekWire. "At present, the entities that generate proprietary training data using their own technology and expense do not have a good way to commercialize the data. So they effectively sit on it."
The data silo problem
High-profile successes like Google DeepMind's AlphaFold thrived because they drew on decades of publicly available, experimentally derived protein structures. But in many areas of biology and medicine, generating high-quality datasets requires years of labor and millions of dollars. The owners of that data currently have no secure, profitable way to share it, so it remains isolated in silos.
Robins spent years at Adaptive Biotechnologies building massive datasets to decode the human immune system. As AI expanded across life sciences, he saw a divide growing: the groups spending heavily to produce research data rarely receive fair compensation when AI developers use those assets to build commercial tools.
Harell Data's approach has three components. A secure cloud platform lets data creators host proprietary datasets while machine learning groups train models without the raw data ever leaving the secure environment. A direct revenue sharing arrangement lets data owners earn a share of the compute revenue generated during each training run, rather than waiting for speculative drug royalties. And as models improve through training on those datasets, the data itself becomes more valuable.
The startup is initially focused on computational medicine, a field Robins knows well. But he said the data-silo problem spans scientific disciplines, including materials science and imaging. "If the business works right, we should be able to enable solutions to really important problems."
An 11-person team, two offices
Harell Data has an 11-person team split between Bellevue and Palo Alto, California. The team includes Chief Technology Officer Rakesh Nair, Head of Operations Saray Covey, and Head of Sales Vidal Polsky. Robins continues as a consultant focusing on scientific strategy at an Adaptive Biotechnologies.
The name Harell Data itself has a personal story. When Robins decided to start the venture, he asked his son, Ellis, for a suggestion. Within seconds, Ellis proposed the name "Harell": a combination of "Harlan" and "Ellis."
"No offense to the large cap cloud compute companies, but it sounds better to me than any of their names, and things seem to have worked out OK for them so I went with it," said Robins.
Why this matters for science and research professionals
For labs and research groups that generate specialized datasets, Harell Data suggests that the data itself may become an asset worth managing-not just a research byproduct. The plat fits into broader movement toward AI training models, and researchers can explore how to structure their data for commercial reuse. For teams funded by grants or internal budgets, the model raises a direct question: could your data generate revenue while advancing science?
Your membership also unlocks: