Frontier-CS Open-Ended · 2026-09-29 21:45 UTC
frontier-cs-v2 vs No Skill
Skill name given by the publisher; not checked.
Keep this Skill change? Regressed
With the Skill, the agent scored lower overall.
-0.031 · 2 better, 5 same, 3 worse.
Without the Skill 0.131 · With the Skill 0.1
Each is the mean case_score_mean
over the tasks, rounded to three decimal places. This Climb does not say what unit or range the score uses.
- Better on 2 tasks: Task 01, Task 06
- Worse on 3 tasks: Task 02, Task 07, Task 08
- No change on 5 tasks.
Examples
Tasks with and without the Skill
A published Result keeps each task's fingerprint and score. It does not keep the task's words or the agent's answers, so these examples show scores only.
-
Task 01 Better with the Skill
sha256:ea3ce53369…- Without the Skill
- 0.167
- With the Skill
- 0.186
- Change
- +0.019
-
Task 02 Worse with the Skill
sha256:7c1214cd1c…- Without the Skill
- 0.444
- With the Skill
- 0.3
- Change
- -0.144
-
Task 03 No change
sha256:4ea68dc70e…- Without the Skill
- 0
- With the Skill
- 0
- Change
- 0
The Skill change
What differed between the two runs
The signed report compared the settings of the two runs and found only this difference, which is the one the Climb allows.
-
The Skill
- Without the Skill
- No Skill
- With the Skill
- sha256:e2a6c529459c265e7470ed5bbf168d6ab3e1ee82ae0abfcd8b3abdf4c53e3087 2,855 bytes
Prime Intellect does not publish a build number for qwen/qwen3.7-flash, so both runs are known to have asked for the same model name, not shown to have used the same build of it.
The signed report names the Skill by its fingerprint, not by the name this page shows for it. Skill name given by the publisher; not checked.
Evidence
What stands behind these numbers
-
Files verified
This site ran its 18 checks on the Result's files, including one that worked out the averages, the change and the decision again from the task scores, and every check passed. How verification works.
-
Reported by the person who ran it
The numbers are signed with the key of the person who ran both runs on their own machine. Nobody else watched the runs.
-
Not yet reproduced
This site has no record of anybody else running this comparison again.
Check it yourself
Run this comparison again
You need
- macOS or Linux
- uv 0.10.2 or later
- Python 3.12, installed for you by uv
- Docker, running
-
An API key for Prime Intellect, set as
PRIME_API_KEY; the model calls are charged to your account - The Skill's files. This Result does not say where to get them.
Run it again
uv tool install --python 3.12 regents-cli==1.2.1
regents techtree setup
regents techtree doctor --climb frontier-cs-open-ended-climb@1
# Put the Skill's files in a folder, then prepare it:
regents techtree climb prepare frontier-cs-open-ended-climb@1 --skill path/to/skill
# Check that the Skill content digest it prints is sha256:e2a6c529459c265e7470ed5bbf168d6ab3e1ee82ae0abfcd8b3abdf4c53e3087
# Start the draft it names. Techtree shows the most it may spend first:
regents techtree climb start DRAFT_ID
# When it finishes, check the run and read its result:
regents techtree run result RUN_ID
Limits
- Each try
- Stops starting model calls at 60 calls, 1,300,000 input tokens or 32,000 output tokens, whichever comes first.
- Whole run
- 10 tasks, each tried once without the Skill and once with it: up to 20 tries. At most 1,200 model calls. The token limits add up to 26,000,000 input tokens and 640,000 output tokens.
- Before it starts
- Techtree shows the most the run may spend and waits for your yes.
The call that crosses a limit still finishes, so a try can go past its token limits by up to one full request and its reply.
What a new run can tell you
- A new run is a new Result. The model may not answer the same way twice, so its numbers can differ from these.
- Whether yours agrees is for you to judge. This site keeps no record that ties a new run to this Result.
The evidence in full
Every task and every fingerprint
01 All 10 tasks
All 10 tasks shown.
-
Task 01
sha256:ea3ce53369…Without the Skill0.167 With the Skill0.186 Change+0.019 -
Task 02
sha256:7c1214cd1c…Without the Skill0.444 With the Skill0.3 Change-0.144 -
Task 03
sha256:4ea68dc70e…Without the Skill0 With the Skill0 Change0 -
Task 04
sha256:5de9efe603…Without the Skill0 With the Skill0 Change0 -
Task 05
sha256:a0803a397f…Without the Skill0 With the Skill0 Change0 -
Task 06
sha256:1dc1d0df21…Without the Skill0.327 With the Skill0.517 Change+0.19 -
Task 07
sha256:e743fd173b…Without the Skill0.1 With the Skill0 Change-0.1 -
Task 08
sha256:e88538664d…Without the Skill0.273 With the Skill0 Change-0.273 -
Task 09
sha256:805a897995…Without the Skill0 With the Skill0 Change0 -
Task 10
sha256:bc76adb754…Without the Skill0 With the Skill0 Change0
02 Comparison conditions and fingerprints
- Climb
- Frontier-CS Open-Ended
- Tasks
- 10 tasks, fixed before either run
- Agent host
- hermes-agent v2026.7.20
- Model
- qwen/qwen3.7-flash from Prime Intellect
- Climb fingerprint
- sha256:7a170f9634b4d9407d1c3817189f4f8c69f0ba314d1cbce41fd8e66571e48c83
- Task list fingerprint
- sha256:a3ad61566551f6d7a1061dc562df06afc5b19ee097b7db364a782051753f7211
- Terms fingerprint
- sha256:36d2689f1e0971456f98b65b2bf853f742f596e8559c611ba406683ad7b0315f
- Result ID
- run_202665f4d7654e5f9ee59058dd69b4e7
- Log sequence
- 8
- What the report can claim
- A call, signed by the person who ran it. Not repeated by anybody else.
- Publisher key
- sha256:bb25e94d220ba4714aa7e80f765ad3dde18d89cbdbc5da21ac47dad712cde9c2
Check this copy
Verify this Result offline.
Download the Result bundle and check it with Techtree on your own computer, without trusting this site. View the recorded data.
Verify offline
regents techtree proof verify techtree-result.json