An evolving program
What evolves is a versioned program of Skills (decomposition, workflow construction, recovery) and typed Operators with explicit input, output and domain contracts.
An AI-for-Science agent that improves from verified execution.
Each task is solved as a typed, executable workflow. When a repair is reproduced under a clean replay, it becomes a linked Skill and Operator, and is kept only if independent validation tasks improve.
A repair that works once rarely survives: it lives in a transient context, or in a single tool or prompt. ScienceClaw asks a sharper question.
How can verified scientific executions drive persistent and transferable program-level self-evolution without updating foundation-model parameters?
What evolves is a versioned program of Skills (decomposition, workflow construction, recovery) and typed Operators with explicit input, output and domain contracts.
Ports carry a schema of type, shape, unit and provenance. The agent edits the graph one atomic action at a time, and nodes are fingerprinted so an edit reruns only what it affects.
Every result is regenerated from a reset environment and checked against hard scientific constraints. A failed replay and a later passing one bracket the shortest reproduced repair.
A repair splits by edit type into a Skill patch and a typed Operator, committed as one atomic bundle. Forming, ranking and selecting candidates needs no LLM judge.
An update persists only if source replay reproduces and uses it, every hard check holds on independent validation tasks, cost stays in budget, and the score strictly improves.
Updates are versioned snapshots with receipts. In the gateway they wait as candidates until you promote them, and any version can be rolled back.
Seven evolution rounds over a common stream of 23 disciplines, with 64 IID and 64 OOD instances per discipline and one fixed foundation model. OOD means an independently sourced dataset of the same discipline; OOD results never generate or select updates.
Share of the full system's gain that each ablation variant retains.
Relative improvement in each discipline's task-native metric, computed from the rounded scores in the paper's tables.
Reported trajectories are final snapshots, not uncertainty estimates over source orders or model configurations, and cost comparisons are relative to the protocol, not absolute.
ScienceClaw runs on the OpenClaw gateway as a plugin with four tools. The agent builds a workflow step by step; you decide what it keeps.
replay re-executes the whole graph from a reset state; finish returns the verified deliverable.scienceclaw_canvasopen · act · render · replay · finish · status · listscienceclaw_toolssearch · show · status · weights · setupscienceclaw_programsummary · skills · operators · show · history · rollbackscienceclaw_evolveval_add · propose · gate · run · status · candidates · show# review what the gate let through, then decide
python -m scienceclaw.cli live candidates
python -m scienceclaw.cli live show <candidate>
python -m scienceclaw.cli live promote <candidate>
python -m scienceclaw.cli live rollback <version>scilib modulesSource replay, validation, budget and strict improvement must all hold, and the gate fails closed when the model is unavailable.
Code nodes run in a separate process with a scrubbed environment, a static scan and a runtime audit guard, plus user, network and pid namespaces where the host supports them.
In the gateway, every citation must come from a tool result in the current conversation, searches cross several sources, and results are written to a file before an answer is final.
Persistent executable updates can reuse errors and widen the attack surface. Run untrusted workloads in a container as well. ScienceClaw should support, not replace, experts; high-stakes use needs provenance, licensing and privacy safeguards, and independent review.
A companion benchmark for continual self-evolution rather than single-shot ability. Systems share the foundation model, the initial program, the source order, the tools and the budget.
An instance counts as solved only if execution completes within budget, the task-native metric meets its acceptance rule, and every scientific hard constraint holds. The evaluation data, 64 IID and 64 OOD records per discipline, is on Hugging Face.
git clone https://github.com/beita6969/ScienceClaw.git cd ScienceClaw
chmod +x setup.sh && ./setup.sh
Installs Node and Python dependencies, the engine, its tools and pretrained weights, MCP servers and skills.
npx openclaw gateway
Enable the scienceclaw plugin in ~/.openclaw/openclaw.json.
The engine never assumes a model family. Point it at any OpenAI-compatible endpoint, a local program, or your own backend.
export SCIENCECLAW_API_BASE_URL=<your OpenAI-compatible endpoint> export SCIENCECLAW_API_KEY=<your key> export SCIENCECLAW_MODEL=<your model name>