Interactive Autoresearch

Your taste,
its execution.

Agents already execute well. That was never the hard part. Sapling pairs that execution with your taste — proposing what to run and why, grounded in what you've already tried, and only moving once you say go.

Create your account → Free tier included, no card required.

runs on whatever's actually worth running

Claude Opus 5 Claude Sonnet 5 GPT-5.6 Grok 4.6 Gemini 3.7 Flash DeepSeek V4 Kimi K3 GLM-5.2
sapling · pair mode · resnet-ablation

you › what should we try next on the ablation?

✻ last 3 runs plateaued around 61% held-out once we dropped the aux head. checking whether that's a capacity or a schedule problem before proposing anything new…

worth isolating: rerun the aux-head config with a 2x longer schedule before we conclude it's capacity. Same seed, same split.

sapling wants to run: launch_experiment

config=aux_head_v3.yaml epochs=200 (was 100) budget≈$4.10

[y] once  [s] session  [p] this project  [a] every project  [n] reject

● log_write(node=run-042, dev_score=0.643, held_out_score=0.611)

✓ logged run-042 · dev 0.643 / held-out 0.611 (schedule, not capacity, aux head kept)

Two lineages, one tool

Sapling inherits its execution from autoresearch systems and its taste from coding agents, one lineage per half of the problem.

Autoresearch agents

execution

  • Tireless: runs hypothesis after hypothesis without waiting on anyone
  • SOTA methods for branching and experiment design, end to end
  • No one in the loop to say "wrong instinct, try this instead"

Coding agent tools

taste-aware

  • Proposes, reasons out loud, waits for your go-ahead
  • Keeps you in the loop by default, not as an afterthought
  • Built for shipping code, not for grounding claims in experiment history

sapling

both

  • Autoresearch-grade execution: hypothesis branching, evidence-checked claims, real literature grounding
  • A coding agent's taste-aware interaction model: propose, decide, run, log
  • Your judgment steering it the whole way

Execution was never the hard part

Agents are already excellent at running things. Sapling is built around the part that's still yours: the taste for knowing what's actually worth trying.

Grounded in your history

Every proposal is checked against what you've already tried and what the real literature supports. Not a generic best-practice guess: a claim tied to evidence.

You set the dial

Cadence, from checking in on every step to running with minimal interruption, is yours to set live, mid-session. Sapling never decides on its own how much rein to take.

An evidence-linked research log

Every outcome is logged locally, dev and held-out, tied to what backs it up. A result can't be marked "promoted" without both. It's enforced automatically.

How a turn works

Pair mode: the default and primary way to run sapling. Always in this order — propose before approve, approve before run.

01

Propose

Sapling reads your research log and proposes a next step, with its reasoning shown inline, not hidden behind a spinner.

02

Approve

You see the exact command and its cost before anything runs. Approve once, for the session, for this project, or for every project.

03

Run

Sapling kicks off the experiment and reads results back from whichever tracker you already use: W&B or MLflow, read-only.

04

Log

The outcome, dev and held-out, is written to a local, append-only research log. Nothing overwritten, nothing silently dropped.

Ready to bring your taste to the loop?

Free to start, no card required. Upgrade whenever you want the full model list.

Create your account →