Techtree

Frontier-CS Open-Ended · 2026-09-29 21:45 UTC

frontier-cs-v2 vs No Skill

Skill name given by the publisher; not checked.

Keep this Skill change? Regressed

With the Skill, the agent scored lower overall.

-0.031 · 2 better, 5 same, 3 worse.

Without the Skill 0.131 · With the Skill 0.1

Each is the mean case_score_mean over the tasks, rounded to three decimal places. This Climb does not say what unit or range the score uses.

  • Better on 2 tasks: Task 01, Task 06
  • Worse on 3 tasks: Task 02, Task 07, Task 08
  • No change on 5 tasks.

Examples

Tasks with and without the Skill

A published Result keeps each task's fingerprint and score. It does not keep the task's words or the agent's answers, so these examples show scores only.

  • Task 01 Better with the Skill

    sha256:ea3ce53369…
    Without the Skill
    0.167
    With the Skill
    0.186
    Change
    +0.019
  • Task 02 Worse with the Skill

    sha256:7c1214cd1c…
    Without the Skill
    0.444
    With the Skill
    0.3
    Change
    -0.144
  • Task 03 No change

    sha256:4ea68dc70e…
    Without the Skill
    0
    With the Skill
    0
    Change
    0

The Skill change

What differed between the two runs

The signed report compared the settings of the two runs and found only this difference, which is the one the Climb allows.

  • The Skill

    Without the Skill
    No Skill
    With the Skill
    sha256:e2a6c529459c265e7470ed5bbf168d6ab3e1ee82ae0abfcd8b3abdf4c53e3087 2,855 bytes

Prime Intellect does not publish a build number for qwen/qwen3.7-flash, so both runs are known to have asked for the same model name, not shown to have used the same build of it.

The signed report names the Skill by its fingerprint, not by the name this page shows for it. Skill name given by the publisher; not checked.

Evidence

What stands behind these numbers

  • Files verified

    This site ran its 18 checks on the Result's files, including one that worked out the averages, the change and the decision again from the task scores, and every check passed. How verification works.

  • Reported by the person who ran it

    The numbers are signed with the key of the person who ran both runs on their own machine. Nobody else watched the runs.

  • Not yet reproduced

    This site has no record of anybody else running this comparison again.

Check it yourself

Run this comparison again

You need

  • macOS or Linux
  • uv 0.10.2 or later
  • Python 3.12, installed for you by uv
  • Docker, running
  • An API key for Prime Intellect, set as PRIME_API_KEY; the model calls are charged to your account
  • The Skill's files. This Result does not say where to get them.

Run it again

uv tool install --python 3.12 regents-cli==1.2.1
regents techtree setup
regents techtree doctor --climb frontier-cs-open-ended-climb@1
# Put the Skill's files in a folder, then prepare it:
regents techtree climb prepare frontier-cs-open-ended-climb@1 --skill path/to/skill
# Check that the Skill content digest it prints is sha256:e2a6c529459c265e7470ed5bbf168d6ab3e1ee82ae0abfcd8b3abdf4c53e3087
# Start the draft it names. Techtree shows the most it may spend first:
regents techtree climb start DRAFT_ID
# When it finishes, check the run and read its result:
regents techtree run result RUN_ID

Limits

Each try
Stops starting model calls at 60 calls, 1,300,000 input tokens or 32,000 output tokens, whichever comes first.
Whole run
10 tasks, each tried once without the Skill and once with it: up to 20 tries. At most 1,200 model calls. The token limits add up to 26,000,000 input tokens and 640,000 output tokens.
Before it starts
Techtree shows the most the run may spend and waits for your yes.

The call that crosses a limit still finishes, so a try can go past its token limits by up to one full request and its reply.

What a new run can tell you

  • A new run is a new Result. The model may not answer the same way twice, so its numbers can differ from these.
  • Whether yours agrees is for you to judge. This site keeps no record that ties a new run to this Result.

The evidence in full

Every task and every fingerprint

01 All 10 tasks

All 10 tasks shown.

  1. Task 01 sha256:ea3ce53369… Without the Skill0.167 With the Skill0.186 Change+0.019
  2. Task 02 sha256:7c1214cd1c… Without the Skill0.444 With the Skill0.3 Change-0.144
  3. Task 03 sha256:4ea68dc70e… Without the Skill0 With the Skill0 Change0
  4. Task 04 sha256:5de9efe603… Without the Skill0 With the Skill0 Change0
  5. Task 05 sha256:a0803a397f… Without the Skill0 With the Skill0 Change0
  6. Task 06 sha256:1dc1d0df21… Without the Skill0.327 With the Skill0.517 Change+0.19
  7. Task 07 sha256:e743fd173b… Without the Skill0.1 With the Skill0 Change-0.1
  8. Task 08 sha256:e88538664d… Without the Skill0.273 With the Skill0 Change-0.273
  9. Task 09 sha256:805a897995… Without the Skill0 With the Skill0 Change0
  10. Task 10 sha256:bc76adb754… Without the Skill0 With the Skill0 Change0
02 Comparison conditions and fingerprints
Climb
Frontier-CS Open-Ended
Tasks
10 tasks, fixed before either run
Agent host
hermes-agent v2026.7.20
Model
qwen/qwen3.7-flash from Prime Intellect
Climb fingerprint
sha256:7a170f9634b4d9407d1c3817189f4f8c69f0ba314d1cbce41fd8e66571e48c83
Task list fingerprint
sha256:a3ad61566551f6d7a1061dc562df06afc5b19ee097b7db364a782051753f7211
Terms fingerprint
sha256:36d2689f1e0971456f98b65b2bf853f742f596e8559c611ba406683ad7b0315f
Result ID
run_202665f4d7654e5f9ee59058dd69b4e7
Log sequence
8
What the report can claim
A call, signed by the person who ran it. Not repeated by anybody else.
Publisher key
sha256:bb25e94d220ba4714aa7e80f765ad3dde18d89cbdbc5da21ac47dad712cde9c2

Check this copy

Verify this Result offline.

Download the Result bundle and check it with Techtree on your own computer, without trusting this site. View the recorded data.

Verify offline

regents techtree proof verify techtree-result.json

All Results · How verification works