HF Tasksmith · 2026-10-04 23:26 UTC
tasksmith-starter-v1 vs No Skill
Skill name given by the publisher; not checked.
Keep this Skill change? Improved
With the Skill, the agent scored higher, and the result clears the rule this Climb set before either run.
+0.333 · 2 better, 4 same, 0 worse.
Without the Skill 0 · With the Skill 0.333
Each is the mean solved
over the tasks, rounded to three decimal places. This Climb does not say what unit or range the score uses.
- Better on 2 tasks: Task 02, Task 05
- Worse on no tasks.
- No change on 4 tasks.
Examples
Tasks with and without the Skill
A published Result keeps each task's fingerprint and score. It does not keep the task's words or the agent's answers, so these examples show scores only.
-
Task 02 Better with the Skill
sha256:cbd58f4675…- Without the Skill
- 0
- With the Skill
- 1
- Change
- +1
-
Task 01 No change
sha256:51f0662127…- Without the Skill
- 0
- With the Skill
- 0
- Change
- 0
The Skill change
What differed between the two runs
The signed report compared the settings of the two runs and found only this difference, which is the one the Climb allows.
-
The Skill
- Without the Skill
- No Skill
- With the Skill
- sha256:974fafc5bce1d489fe332c7bdc8e9b9ecb38b873d407ac7d38c70f098362b595 1,308 bytes
Prime Intellect does not publish a build number for openai/gpt-6-luna, so both runs are known to have asked for the same model name, not shown to have used the same build of it.
The signed report names the Skill by its fingerprint, not by the name this page shows for it. Skill name given by the publisher; not checked.
Evidence
What stands behind these numbers
-
Files verified
This site ran its 18 checks on the Result's files, including one that worked out the averages, the change and the decision again from the task scores, and every check passed. How verification works.
-
Reported by the person who ran it
The numbers are signed with the key of the person who ran both runs on their own machine. Nobody else watched the runs.
-
Not yet reproduced
This site has no record of anybody else running this comparison again.
Check it yourself
Run this comparison again
You need
- macOS or Linux
- uv 0.10.2 or later
- Python 3.12, installed for you by uv
- Docker, running
-
An API key for Prime Intellect, set as
PRIME_API_KEY; the model calls are charged to your account - The Skill's files. This Result does not say where to get them.
Run it again
uv tool install --python 3.12 regents-cli==1.5.0
regents techtree setup
regents techtree doctor --climb tasksmith-climb@1
# Put the Skill's files in a folder, then prepare it:
regents techtree climb prepare tasksmith-climb@1 --skill path/to/skill
# Check that the Skill content digest it prints is sha256:974fafc5bce1d489fe332c7bdc8e9b9ecb38b873d407ac7d38c70f098362b595
# Start the draft it names. Techtree shows the most it may spend first:
regents techtree climb start DRAFT_ID
# When it finishes, check the run and read its result:
regents techtree run result RUN_ID
Limits
- Each try
- Stops starting model calls at 80 calls, 2,000,000 input tokens or 96,000 output tokens, whichever comes first.
- Whole run
- 6 tasks, each tried once without the Skill and once with it: up to 12 tries. At most 960 model calls. The token limits add up to 24,000,000 input tokens and 1,152,000 output tokens.
- Before it starts
- Techtree shows the most the run may spend and waits for your yes.
The call that crosses a limit still finishes, so a try can go past its token limits by up to one full request and its reply.
What a new run can tell you
- A new run is a new Result. The model may not answer the same way twice, so its numbers can differ from these.
- Whether yours agrees is for you to judge. This site keeps no record that ties a new run to this Result.
The evidence in full
Every task and every fingerprint
01 All 6 tasks
All 6 tasks shown.
-
Task 01
sha256:51f0662127…Without the Skill0 With the Skill0 Change0 -
Task 02
sha256:cbd58f4675…Without the Skill0 With the Skill1 Change+1 -
Task 03
sha256:b8c080e896…Without the Skill0 With the Skill0 Change0 -
Task 04
sha256:969c48aff0…Without the Skill0 With the Skill0 Change0 -
Task 05
sha256:a4b59f42e7…Without the Skill0 With the Skill1 Change+1 -
Task 06
sha256:ef45de02c5…Without the Skill0 With the Skill0 Change0
02 Comparison conditions and fingerprints
- Climb
- HF Tasksmith
- Tasks
- 6 tasks, fixed before either run
- Agent host
- hermes-agent v2026.9.24
- Model
- openai/gpt-6-luna from Prime Intellect
- Climb fingerprint
- sha256:e4c5943f9bb0f03340ad3ae2709a9329318dac3b4529f3029ef558d7f11e403a
- Task list fingerprint
- sha256:736f16777d51e38e09961b24615b45e4ff2109686ab0a72b7127465c9f780ffe
- Terms fingerprint
- sha256:eebb0cc785d857f47c2f0f918486aac9aba6880b992f489bc077c0fb7b5efe97
- Result ID
- run_4293178c2f9c4712a5fd16bd03de7545
- Log sequence
- 9
- What the report can claim
- A call, signed by the person who ran it. Not repeated by anybody else.
- Publisher key
- sha256:bb25e94d220ba4714aa7e80f765ad3dde18d89cbdbc5da21ac47dad712cde9c2
Check this copy
Verify this Result offline.
Download the Result bundle and check it with Techtree on your own computer, without trusting this site. View the recorded data.
Verify offline
regents techtree proof verify techtree-result.json