Google DeepMind has connected its Gemini-powered Co-Scientist system to real laboratory equipment, allowing it to design experiments, generate machine code, and grow materials with minimal human input. In a newly published 83-page paper, the company describes moving from what it calls an "in-silico hypothesis generator" to an execution-grounded research partner - a shift with direct implications for scientists who spend weeks troubleshooting equipment-specific variables.
In the most direct demonstration, researchers gave Gemini 3 Deep Think the parameters of a self-built chemical vapor deposition (CVD) furnace. Within minutes, the model produced a material growth plan tailored to that specific machine and translated the solution into code the device could execute. Three types of 2D semiconductors - MoS₂, MoSe₂ and WS₂ - grew successfully on the first attempt, with the whole process taking about one hour. MoSe₂ and WS₂ were especially notable because the research team had never grown those materials on that equipment before.
The work arrives one day after Anthropic released its Model Hardware Standard (MHS), which extends its MCP protocol beyond software tools to physical devices like robotic arms, microscopes and liquid handlers. Taken together, the two developments point toward a near-term reality: AI systems with standardized ways to manipulate lab hardware.
Finding new synthesis routes without a recipe
Parameter tuning on existing equipment is one thing. The more demanding test was whether Co-Scientist could propose a genuinely new synthesis route. Google assigned it the task of finding a safer precursor for Ti₃C₂Tₓ, a MXene 2D material commonly made with toxic or corrosive chemicals.
The AI settled on hexachloroethane C₂Cl₆ and generated 272 candidate schemes. Researchers selected the highest-ranked options, then ran 25 rounds of iterative experiments before obtaining a 2D layered crystal with characteristics highly similar to Ti₃C₂Tₓ MXene, based on XRD patterns, electron microscopy and elemental analysis.
The failure rate during reproduction told a different kind of story. The initial experimental success rate was only 11.5 percent. The culprit wasn't the AI's chemistry reasoning but oxygen leakage from poorly sealed equipment. After researchers cleaned the quartz tube, replaced sealing rings and improved maintenance procedures, the same protocol succeeded 68 percent of the time.
That distinction matters for scientists: a wrong code can be re-run, but a working scientific plan will fail if an aging O-ring introduces trace oxygen. The Google researchers acknowledge the material hasn't been definitively confirmed as Ti₃C₂Tₓ MXene, and that yield and oxidation issues remain.
Interpolation experiment reduces wet-lab workload
In a synthetic biology experiment, Google tested Co-Scientist's ability to predict unmeasured outcomes. Engineered E. coli colonies form different patterns depending on the concentration of the chemical inducer IPTG. Rather than running the full concentration range in wet experiments, the team gave the AI real colony images at some concentration points and asked it to predict intermediate states it had never seen.
Among four colony morphology indicators, three of the AI's predictions showed no significant difference from real wet-lab results. The model also correctly predicted that the control group wouldn't change with IPTG concentration. Its notable weakness: generated colonies looked rounder and more regular than real ones.
Google is careful to say the model achieved interpolation within a known concentration range, not novel prediction for an unseen genetic circuit. But the practical workflow scientists need is clear: measure a few points, let AI fill in the plausible space, then choose the most informative positions for wet-lab verification.
AI designing AI: recursive agent architecture
Computer science experiments showed the highest level of autonomy. The research team asked Co-Scientist to design an agent that answers medical questions accurately. Humans did not touch the architecture design after that.
Co-Scientist proposed its own approach, wrote the code, ran tests, analysed errors, and revised the design. It settled on a system called Agent_H that classifies each medical question by field, intended audience, and risk level; breaks complex questions into sub-questions; generates 28-48 candidate answers in parallel; uses different judges for pairwise elimination; and sends the winner through multiple clinical review and citation rounds.
The same paper describes the agent working across materials, biology and computer science, suggesting that the underlying pattern - model proposes, model executes, model iterates - is not dependent on a single field's conventions.
Why this matters for research scientists
Takeaway for scientists who work with physical equipment: the interoperability layer is arriving faster than expected. Anthropic and Google are converging on ways to connect LLMs to standard lab hardware, and the CVD demonstration shows that a well-described setup plus a capable reasoning model can produce first-try results that previously took months of manual adjustments.
The realistic limit for the next 12 to 18 months. The Google team's greatest obstacle was not chemistry but hardware maintenance - seals, oxygen leaks, contamination. The immediate opportunities are clear: likelihood of parameter fitting, interpolation of unmeasured conditions, and generation of safe synthesis candidates before stepping into the lab. For a job where the field's training data only grows if labs share AI-generated results, the incentive is to start giving the loop some real experiments.
For those in the experimental workflow, this evolution from hypothesis-only AI to execution-partner support deserves closer engagement. Scientists who begin to document their equipment and protocols in machine-readable form will be best positioned to make and direct the systems do the first round of trial-anderror work.
Your membership also unlocks: