OpenScience is an open-source AI workbench for scientific research, described by its creators as an open-source AI co-scientist. It provides one workspace for literature, code, experiments, compute, and results, replacing the usual scatter of tools a researcher juggles during a project. The agent reads papers, writes code, and runs experiments alongside the user, working inside notebooks and a terminal rather than in a single chat window. It is model agnostic: free models are included, and users can bring Claude, GPT, Gemini, or any other provider. OpenScience is free and open source, and it is backed by Synthetic Sciences and Y Combinator.
Scientific research work is fragmented by nature. A single question can require reading a stack of papers, locating measurements in a public database, writing scripts to analyze a structure, running those scripts somewhere with enough compute, and then plotting and interpreting the output. Each of those steps lives in a different tool, and long-running jobs in particular tend to break the flow of an investigation. General-purpose AI assistants help with pieces of that work but were not built around the scientific stack: they do not natively treat databases such as UniProt or PDB as tools, they do not schedule jobs onto a Slurm or PBS cluster, and their performance can vary depending on which model or provider route a request happens to take. OpenScience is positioned against exactly that problem, offering one workspace that spans literature, code, experiments, compute, and results, plus a set of models the team has tested and benchmarked specifically for scientific agents so that behaviour stays consistent across routes.
At the core is an agent that reads papers, writes code, and runs experiments with the user. It operates in notebooks and in a terminal, which matters because scientific work rarely fits inside a chat interface: notebooks are where analysis lives, and the terminal is where environments, scripts, and job submission live. The agent can execute the code it writes and return concrete artefacts rather than suggestions. In the example published on the site, it loads a research-lookup skill, searches two sources, reads a file, runs a Python script against a PDB structure, and produces a 1200 by 900 PNG plot comparing predicted and measured values. The interface shows Copy, Undo, and Fork controls along with an Explore agent, so a researcher can branch a line of work, revert it, or run investigation threads in parallel without starting over.
OpenScience is model agnostic. Free models are included, and users can supply their own keys for any provider, including Claude, GPT, and Gemini. One option described on the site is using 30+ models through a single wallet, which removes the need to maintain separate provider accounts. Users who already pay for ChatGPT Plus or Pro can sign in with OpenAI and use the subscription they already have, and a local model can be run instead where that is preferred. For a curated route, Ace gives a handpicked set of models that OpenScience has tested and benchmarked for scientific agents, plus managed search and memory, all behind one Wallet. The stated aim is to avoid provider accounts and to avoid inconsistent performance across different routes.
Scientific databases are exposed to the agent as tools rather than as something a user has to query manually. The site names UniProt, PDB, ChEMBL, PubChem, and arXiv, and says there are 37 more, putting more than forty data sources within reach of a single research session. Alongside those, OpenScience bundles 371 skills across biology, chemistry, physics, machine learning, and writing, with a curated research core. Skills act as prepared capabilities the agent loads when a task calls for them; in the published demo, the agent loads a research-lookup skill before it searches sources and reads files. The combination is useful because it lets the agent ground its reasoning in real measurements and structures, for example by scoring the same set of mutants that an external measurement set covers, rather than working only from what a language model happens to recall.
Compute is managed rather than left to the user. OpenScience builds environments and scales on demand, and it can run work on a laptop, on a cluster, or on GPUs. Long jobs are sent to Modal, to the user's own servers, or to a Slurm/PBS cluster, so a researcher can keep working while an experiment runs elsewhere. Internally the agent spends time on tasks and decides what to do next, as the demo shows with a working timer and a short reasoning step before it searches and cross-checks data. For experiment-driven work, Autoresearch takes a metric as input, runs experiments, logs every one of them, and keeps hill-climbing, so an optimisation or parameter search can proceed without manual babysitting. Multi-session support lets several agents run in parallel on the same project, which suits investigations with separate threads of analysis or comparison.
The overall approach is to put an agentic loop on top of a real scientific toolchain rather than a chat window. A session starts with a task description; the agent reasons about it, loads any skills it needs, consults databases and files, writes and runs code, and returns results with the artefacts attached. When a question calls for measurement, the agent finds the comparison set, scores the same mutants the same way, and plots predicted against measured values: the published example reports a correlation of r = 0.71 across n = 26 mutants and names the three substitutions that are stabilising under both prediction and measurement. Because the work happens in notebooks and a terminal with managed compute behind it, the same environment can carry a task from literature lookup through job execution to a finished figure. The team also publishes benchmark results to show where the agent stands: 75.7% on Terminal-Bench Science, 71.4% on Terminal-Bench 4.0 (science), 82.2 on BiomniBench-DA, and 47.3% pass@3 on OpenScience Bench, which measures end-to-end research.
The stated benefit is a single workspace in which literature, code, experiments, compute, and results live together, so a researcher spends time on the question rather than on moving data between tools. Reading papers, writing code, and running experiments are handled by one agent that can also schedule the heavy jobs. Model choice becomes a configuration detail rather than a blocker: free models are available immediately, the user's own keys work for any provider, a ChatGPT Plus or Pro subscription can be reused, or a local model can be run. Because scientific databases are available as tools, answers can be checked against real records within the same session. And because experiments are logged and the Autoresearch loop keeps climbing toward a metric, iterative work retains a record of what was tried. On the team's own measurements, the agent leads every scientific benchmark they have run.
Concrete scenarios appear throughout the site. The most fully described is a protein stability scan: the user asks which T4 lysozyme point mutants are predicted to be stabilising and asks for a comparison against ProTherm measurements. The agent identifies ProTherm as the comparison set, scores the same 26 mutants on the 2LZM structure, cross-checks the entries, and plots predicted against measured ΔΔG. The example also shows the follow-ups such a workflow implies, such as running the three candidates through FoldX for an independent estimate or drafting the methods paragraph with the ProTherm citation. Other stated uses include reading and searching literature, running code in notebooks and a terminal, sending long jobs to Modal, personal servers, or a Slurm/PBS cluster, running several agents in parallel on the same project, and using Autoresearch to iterate against a chosen metric with every experiment logged.
OpenScience is aimed at researchers and scientists who write code as part of their work, across the fields its bundled skills cover: biology, chemistry, physics, machine learning, and writing. The project is backed by Synthetic Sciences and Y Combinator, and its Product Hunt topics are Open Source, Artificial Intelligence, and Science. Integrations named in the content include scientific databases such as UniProt, PDB, ChEMBL, PubChem, arXiv and 37 more; model providers such as Claude, GPT, and Gemini; OpenAI sign-in for ChatGPT Plus or Pro subscribers; local models; the Ace model catalogue; and compute targets including Modal, the user's own servers, and Slurm/PBS clusters. Installation is shown as a shell one-liner: curl -fsSL https://openscience.sh/install | bash. The software is free and open source with free models included, and documentation is published at openscience.sh/docs.
OpenScience's value proposition is straightforward: an open-source, model-agnostic AI workbench that treats scientific research as a full workflow rather than a chat. It reads papers, writes and runs code, operates in notebooks and a terminal, reaches more than forty scientific databases as tools, brings 371 bundled skills, and manages compute from a laptop to a cluster or GPUs. Autoresearch turns a metric into a logged, iterative experiment loop, multi-session support allows parallel agents on one project, and free models or any provider, including Claude, GPT, Gemini, a ChatGPT subscription, or a local model, can drive it. The result, according to the team's benchmark reports, is an AI co-scientist that leads the scientific benchmarks they have run.