Open source · Verifiable program-level self-evolution

ScienceClaw

An AI-for-Science agent that improves from verified execution.

Each task is solved as a typed, executable workflow. When a repair is reproduced under a clean replay, it becomes a linked Skill and Operator, and is kept only if independent validation tasks improve.

23Disciplines
+16.45%Mean OOD gain over the frozen agent
98.23%Hard-constraint pass rate
91.30OOD macro success rate after seven rounds
ScienceClaw-Eval spans 23 disciplines across the natural and social sciences, while ScienceClaw turns verified execution evidence into persistent Skill–Operator program updates.
Method

Persistent improvement from verified execution

A repair that works once rarely survives: it lives in a transient context, or in a single tool or prompt. ScienceClaw asks a sharper question.

How can verified scientific executions drive persistent and transferable program-level self-evolution without updating foundation-model parameters?
01

An evolving program

What evolves is a versioned program of Skills (decomposition, workflow construction, recovery) and typed Operators with explicit input, output and domain contracts.

02

Typed, execution-guided workflows

Ports carry a schema of type, shape, unit and provenance. The agent edits the graph one atomic action at a time, and nodes are fingerprinted so an edit reruns only what it affects.

03

Evidence you can replay

Every result is regenerated from a reset environment and checked against hard scientific constraints. A failed replay and a later passing one bracket the shortest reproduced repair.

04

Linked Skill–Operator updates

A repair splits by edit type into a Skill patch and a typed Operator, committed as one atomic bundle. Forming, ranking and selecting candidates needs no LLM judge.

05

A gate, not a guess

An update persists only if source replay reproduces and uses it, every hard check holds on independent validation tasks, cost stays in budget, and the score strictly improves.

06

You stay in control

Updates are versioned snapshots with receipts. In the gateway they wait as candidates until you promote them, and any version can be rolled back.

Given task specification Dt and agent program Ar, ScienceClaw produces a scientific solution Zt and retains a candidate update only after source-task replay and independent program validation.
Task formulation
Task formulation of verifiable program-level self-evolution for AI-for-Science agents.
Results

It keeps getting better — and the gain transfers

Seven evolution rounds over a common stream of 23 disciplines, with 64 IID and 64 OOD instances per discipline and one fixed foundation model. OOD means an independently sourced dataset of the same discipline; OOD results never generate or select updates.

+16.45%mean OOD gain over the frozen agent
+11.57% to +23.34% per discipline
77.72 → 91.30OOD macro success rate over seven rounds
+13.59 pp
70 / 161candidates promoted
43.48%, a selective gate that still moves
18 / 20cross-family transfer pairs positive
+1.58 pp across, +14.69 pp within a family
0.13 ppaverage forgetting
max 0.51 pp, 11.89% negative transfer
98.23%of instances pass every hard constraint
1.23% erroneous promotions

What the gain depends on

Share of the full system's gain that each ablation variant retains.

Gain over the frozen agent, by discipline

Relative improvement in each discipline's task-native metric, computed from the rounded scores in the paper's tables.

Task-native scores of the ablation variants across 23 disciplines; colours are normalised within each discipline (darker is better).
(f) Gain of the linked-mechanism variants on Commerce and Law. (g) Workflow mechanics: planner rounds, distinct Operators, feedback repair, checkpoint recovery and clean replay. (h) Hard-constraint pass rate against cost per OOD gain (ScienceClaw = 1), coloured by planner wall-time share. (i) Promoted candidates and promotion rate.

Reported trajectories are final snapshots, not uncertainty estimates over source orders or model configurations, and cost comparisons are relative to the protocol, not absolute.

In practice

Use it from the gateway

ScienceClaw runs on the OpenClaw gateway as a plugin with four tools. The agent builds a workflow step by step; you decide what it keeps.

  1. Declare the task. Objective, inputs, required output and constraints. Inputs become read-only loaders; constraints are the acceptance test.
  2. Build the workflow. Find tools, then add, modify or remove one node or edge per step, reading typed feedback after each edit.
  3. Verify. replay re-executes the whole graph from a reset state; finish returns the verified deliverable.
  4. Evolve (optional). Register independent validation tasks, propose candidates from the verified session, and gate them.
  5. You decide. Review and promote with the CLI. Every promotion is a new program version you can roll back.
scienceclaw_canvasopen · act · render · replay · finish · status · list
scienceclaw_toolssearch · show · status · weights · setup
scienceclaw_programsummary · skills · operators · show · history · rollback
scienceclaw_evolveval_add · propose · gate · run · status · candidates · show
# review what the gate let through, then decide
python -m scienceclaw.cli live candidates
python -m scienceclaw.cli live show <candidate>
python -m scienceclaw.cli live promote <candidate>
python -m scienceclaw.cli live rollback <version>
300+skills, including 36 seed Skills for the evolving program
112typed library Operators
292scientific tool functions in 42 scilib modules
28pretrained-weight assets, fetched on demand
🔒

Nothing persists without evidence

Source replay, validation, budget and strict improvement must all hold, and the gate fails closed when the model is unavailable.

🧪

Sandboxed execution

Code nodes run in a separate process with a scrubbed environment, a static scan and a runtime audit guard, plus user, network and pid namespaces where the host supports them.

📚

A citation protocol

In the gateway, every citation must come from a tool result in the current conversation, searches cross several sources, and results are written to a file before an answer is final.

Persistent executable updates can reuse errors and widen the attack surface. Run untrusted workloads in a container as well. ScienceClaw should support, not replace, experts; high-stakes use needs provenance, licensing and privacy safeguards, and independent review.

Benchmark

ScienceClaw-Eval

A companion benchmark for continual self-evolution rather than single-shot ability. Systems share the foundation model, the initial program, the source order, the tools and the budget.

Source streamthe only evolution evidence
Validation setindependent, used to select candidates
ID and OODheld-out and cross-dataset, per discipline
Replay setearlier source tasks, to measure retention
Scientific-task collection, executable instantiation, validation and reproduction, and lineage-aware evaluation splits.

An instance counts as solved only if execution completes within budget, the task-native metric meets its acceptance rule, and every scientific hard constraint holds. The evaluation data, 64 IID and 64 OOD records per discipline, is on Hugging Face.

Get started

From clone to a running agent

1

Clone

git clone https://github.com/beita6969/ScienceClaw.git
cd ScienceClaw
2

Set up

chmod +x setup.sh && ./setup.sh

Installs Node and Python dependencies, the engine, its tools and pretrained weights, MCP servers and skills.

3

Run

npx openclaw gateway

Enable the scienceclaw plugin in ~/.openclaw/openclaw.json.

Bring your own model

The engine never assumes a model family. Point it at any OpenAI-compatible endpoint, a local program, or your own backend.

export SCIENCECLAW_API_BASE_URL=<your OpenAI-compatible endpoint>
export SCIENCECLAW_API_KEY=<your key>
export SCIENCECLAW_MODEL=<your model name>