Techtree

HF Tasksmith · 2026-10-04 23:26 UTC

tasksmith-starter-v1 vs No Skill

Skill name given by the publisher; not checked.

Keep this Skill change? Improved

With the Skill, the agent scored higher, and the result clears the rule this Climb set before either run.

+0.333 · 2 better, 4 same, 0 worse.

Without the Skill 0 · With the Skill 0.333

Each is the mean solved over the tasks, rounded to three decimal places. This Climb does not say what unit or range the score uses.

  • Better on 2 tasks: Task 02, Task 05
  • Worse on no tasks.
  • No change on 4 tasks.

Examples

Tasks with and without the Skill

A published Result keeps each task's fingerprint and score. It does not keep the task's words or the agent's answers, so these examples show scores only.

  • Task 02 Better with the Skill

    sha256:cbd58f4675…
    Without the Skill
    0
    With the Skill
    1
    Change
    +1
  • Task 01 No change

    sha256:51f0662127…
    Without the Skill
    0
    With the Skill
    0
    Change
    0

The Skill change

What differed between the two runs

The signed report compared the settings of the two runs and found only this difference, which is the one the Climb allows.

  • The Skill

    Without the Skill
    No Skill
    With the Skill
    sha256:974fafc5bce1d489fe332c7bdc8e9b9ecb38b873d407ac7d38c70f098362b595 1,308 bytes

Prime Intellect does not publish a build number for openai/gpt-6-luna, so both runs are known to have asked for the same model name, not shown to have used the same build of it.

The signed report names the Skill by its fingerprint, not by the name this page shows for it. Skill name given by the publisher; not checked.

Evidence

What stands behind these numbers

  • Files verified

    This site ran its 18 checks on the Result's files, including one that worked out the averages, the change and the decision again from the task scores, and every check passed. How verification works.

  • Reported by the person who ran it

    The numbers are signed with the key of the person who ran both runs on their own machine. Nobody else watched the runs.

  • Not yet reproduced

    This site has no record of anybody else running this comparison again.

Check it yourself

Run this comparison again

You need

  • macOS or Linux
  • uv 0.10.2 or later
  • Python 3.12, installed for you by uv
  • Docker, running
  • An API key for Prime Intellect, set as PRIME_API_KEY; the model calls are charged to your account
  • The Skill's files. This Result does not say where to get them.

Run it again

uv tool install --python 3.12 regents-cli==1.5.0
regents techtree setup
regents techtree doctor --climb tasksmith-climb@1
# Put the Skill's files in a folder, then prepare it:
regents techtree climb prepare tasksmith-climb@1 --skill path/to/skill
# Check that the Skill content digest it prints is sha256:974fafc5bce1d489fe332c7bdc8e9b9ecb38b873d407ac7d38c70f098362b595
# Start the draft it names. Techtree shows the most it may spend first:
regents techtree climb start DRAFT_ID
# When it finishes, check the run and read its result:
regents techtree run result RUN_ID

Limits

Each try
Stops starting model calls at 80 calls, 2,000,000 input tokens or 96,000 output tokens, whichever comes first.
Whole run
6 tasks, each tried once without the Skill and once with it: up to 12 tries. At most 960 model calls. The token limits add up to 24,000,000 input tokens and 1,152,000 output tokens.
Before it starts
Techtree shows the most the run may spend and waits for your yes.

The call that crosses a limit still finishes, so a try can go past its token limits by up to one full request and its reply.

What a new run can tell you

  • A new run is a new Result. The model may not answer the same way twice, so its numbers can differ from these.
  • Whether yours agrees is for you to judge. This site keeps no record that ties a new run to this Result.

The evidence in full

Every task and every fingerprint

01 All 6 tasks

All 6 tasks shown.

  1. Task 01 sha256:51f0662127… Without the Skill0 With the Skill0 Change0
  2. Task 02 sha256:cbd58f4675… Without the Skill0 With the Skill1 Change+1
  3. Task 03 sha256:b8c080e896… Without the Skill0 With the Skill0 Change0
  4. Task 04 sha256:969c48aff0… Without the Skill0 With the Skill0 Change0
  5. Task 05 sha256:a4b59f42e7… Without the Skill0 With the Skill1 Change+1
  6. Task 06 sha256:ef45de02c5… Without the Skill0 With the Skill0 Change0
02 Comparison conditions and fingerprints
Climb
HF Tasksmith
Tasks
6 tasks, fixed before either run
Agent host
hermes-agent v2026.9.24
Model
openai/gpt-6-luna from Prime Intellect
Climb fingerprint
sha256:e4c5943f9bb0f03340ad3ae2709a9329318dac3b4529f3029ef558d7f11e403a
Task list fingerprint
sha256:736f16777d51e38e09961b24615b45e4ff2109686ab0a72b7127465c9f780ffe
Terms fingerprint
sha256:eebb0cc785d857f47c2f0f918486aac9aba6880b992f489bc077c0fb7b5efe97
Result ID
run_4293178c2f9c4712a5fd16bd03de7545
Log sequence
9
What the report can claim
A call, signed by the person who ran it. Not repeated by anybody else.
Publisher key
sha256:bb25e94d220ba4714aa7e80f765ad3dde18d89cbdbc5da21ac47dad712cde9c2

Check this copy

Verify this Result offline.

Download the Result bundle and check it with Techtree on your own computer, without trusting this site. View the recorded data.

Verify offline

regents techtree proof verify techtree-result.json

All Results · How verification works