Techtree

Example comparison · 27 September 2026

tdd vs No Skill

A local comparison of the tdd Skill against no Skill, reported by the person who ran it.

Keep this Skill change? Improved

On the 3 held-out tasks, the agent did better with the Skill on 1 task and worse on none.

Judged on the held-out tasks alone: tasks kept from any agent that tries to improve the Skill, so the answer does not rest on tasks it could learn from.

On average 0.667 without the Skill, 1 with it (+0.333). Averages are rounded to three decimal places.

  • Better on 1 task: retry-backoff-boundary
  • Worse on no tasks.
  • No change on 2 tasks: reroute-public-seam, shipping-oracle

Tasks the improving agent could see

Improved on the 3 other tasks

An agent that tries to improve the Skill from this comparison is shown these tasks, so they are reported here on their own and do not decide the answer above.

On average 0.333 without the Skill, 1 with it (+0.667). Averages are rounded to three decimal places.

  • Better on 2 tasks: exchange-rate-boundary, deactivate-account-interface
  • Worse on no tasks.
  • No change on 1 task: theater-cents-adr

Claims

What the Skill claims, and the tasks that test it

A positive case is one where following the Skill should give the right result; a boundary case sits at the edge of where the claim applies; a counterexample checks that the Skill is not overused where it would give a wrong result or should change nothing.

C1. Tests target public behavioral seams rather than private methods, internal collaborators, or side channels, so they remain valid across behavior-preserving refactors.

What would show it: The submitted tests exercise and observe the feature only through exported interfaces and still pass when the implementation is replaced by an equivalent version with different private structure and storage.

  • reroute-public-seam No change

    Implement and test rerouting an undispatched parcel to a new destination in a small Python parcel service.

    Positive case · Held out · time limit 30 minutes

    Without the Skill
    1
    With the Skill
    1
    Change
    0
  • deactivate-account-interface Better with the Skill

    Add account deactivation to a small Python accounts package backed by SQLite, with tests.

    Boundary case · Seen by the improving agent · time limit 30 minutes

    Without the Skill
    0
    With the Skill
    1
    Change
    +1

C2. Tests use an independent source of truth for expected results instead of recomputing expectations with the production algorithm.

What would show it: The submitted tests pass against the correct implementation and fail against plausible algorithm mutants, including mutants that would also affect expectations obtained by calling production code twice.

  • shipping-oracle No change

    Implement and test the shipping quote rules described for a local command-line fulfillment package.

    Positive case · Held out · time limit 30 minutes

    Without the Skill
    1
    With the Skill
    1
    Change
    0

C3. Tests substitute dependencies at system boundaries such as time or external services while using real internal collaborators.

What would show it: The tests control the external boundary deterministically, contain no doubles for project-owned collaborators, and continue to pass after those collaborators are reorganized without changing public behavior.

  • exchange-rate-boundary Better with the Skill

    Add a converted price quote with a conversion fee to a small Python pricing package whose exchange rates come from a remote rates service, with tests.

    Positive case · Seen by the improving agent · time limit 30 minutes

    Without the Skill
    0
    With the Skill
    1
    Change
    +1
  • retry-backoff-boundary Better with the Skill

    Add retries with growing waits to a small Python notifier that sends messages through a mail service client, with tests.

    Positive case · Held out · time limit 30 minutes

    Without the Skill
    0
    With the Skill
    1
    Change
    +1

C4. Code and tests follow the repository's established domain vocabulary and applicable architectural decisions.

What would show it: The resulting public API, test descriptions, values, and persisted representations use the terms and constraints defined by the repository guidance rather than generic alternatives.

  • theater-cents-adr No change

    Add seat holds and the amount owed for them to a small Python theater-sales package, with tests.

    Positive case · Seen by the improving agent · time limit 30 minutes

    Without the Skill
    1
    With the Skill
    1
    Change
    0

The Skill change

What differed between the two runs

The two runs used the same tasks, model and limits. One ran with no Skill; the other ran with this Skill, the one the tasks were written from.

Without the Skill
No Skill
With the Skill
tdd
Licence
MIT License. Copyright (c) 2026 Matt Pocock
Fingerprint
sha256:7dc0ee968fc1f3717b653a11b172f4c89b080edb0d4290d61b21135ffad14099

The Skill's files, its licence among them, are below, exactly as the run with the Skill used them. Read the Skill's files.

Evidence

What stands behind these numbers

  • Files checked

    This site checked the files behind this page against each other: each of the Skill's files and the Skill's fingerprint, both runs' records, the export's tasks and instructions, and the decision, worked out again from the task scores. Every check passed.

  • Reported by the person who ran it

    One local comparison, run on one computer by the person who reported it. It is not a published Result, nothing about it is signed, and nobody else watched the runs.

  • Not yet reproduced

    This site has no record of anybody else running this comparison again.

Method

How the numbers were made

What this is
One local comparison on one computer. It is not a published or signed Result.
Tries
Each task was tried once without the Skill and once with it.
Model
gpt-5.6-sol from OpenAI, with medium reasoning, as both runs asked for it
What limited the runs
Time: 30 minutes for each try. Each try ran in its own container with 2 processor cores, 4 GB of memory and no network. Nothing limited the agent's turns, model calls or tokens.
Model calls and tokens
Without the Skill 41 model calls, 338,454 tokens. With the Skill 95 model calls, 1,026,256 tokens.

What these runs cannot show

  • Which model the provider actually served. Only the provider's own report, passed on by Hermes, says which model answered.
  • That the Hermes agent program was unmodified. Its version is only what it reports.
  • How the model chose its answers. Those settings cannot be changed, so every try used the provider's defaults.

Check it yourself

Check the tasks and run this comparison again

Fingerprint of the tasks
sha256:5ce52d4aa2b5ba23ffdc5e7982a508b7ad9bc938f7f46a1b06feb324a7bab808

The first command below prints the fingerprint of the tasks it checks. They are the same tasks as these only if it matches this one.

You need

  • The export folder of these tasks, with its tests and records. Use it only if the first command prints the fingerprint above; a folder with any other fingerprint holds other tasks.
  • The Skill's files, shown on this page, in a folder of their own
  • Docker, Techtree and Hermes, as the export folder's own instructions describe
  • An account with a model provider you choose; its calls may cost money

Check it and run it again

techtree forge verify-export EXPORT_FOLDER
techtree forge import EXPORT_FOLDER
techtree forge inspect-skill SKILL_FOLDER
techtree forge run --arm baseline --collection forgecol_ffa133338756407eba327098051f996c --provider PROVIDER --model MODEL
techtree forge run --arm candidate --collection forgecol_ffa133338756407eba327098051f996c --provider PROVIDER --model MODEL --skill SKILL_FOLDER
techtree forge compare BASELINE_RUN_ID CANDIDATE_RUN_ID

Words in capitals stand for what only you know: your folders, the provider and model you choose, and the ids the two runs print. Each run shows what it will do and asks before it starts.

What a new run can tell you

  • A new run is a new comparison. The model may not answer the same way twice, so its numbers can differ from these.
  • Whether yours agrees is for you to judge. This site keeps no record that ties a new run to this one.
  • To test your own change to this Skill, give the baseline run the earlier version with --skill instead of no Skill, and the candidate run your new version. The comparison then says whether the change is worth keeping. How to compare two versions →

The evidence in full

Every task, the Skill's files and the fingerprints

01 All 6 tasks
  1. reroute-public-seam Held out Without the Skill1 With the Skill1 Change0
  2. shipping-oracle Held out Without the Skill1 With the Skill1 Change0
  3. exchange-rate-boundary Seen by the improving agent Without the Skill0 With the Skill1 Change+1
  4. deactivate-account-interface Seen by the improving agent Without the Skill0 With the Skill1 Change+1
  5. theater-cents-adr Seen by the improving agent Without the Skill1 With the Skill1 Change0
  6. retry-backoff-boundary Held out Without the Skill0 With the Skill1 Change+1
02 The Skill's file LICENSE.txt
MIT License

Copyright (c) 2026 Matt Pocock

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
03 The Skill's file SKILL.md
---
name: tdd
description: Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions "red-green-refactor", or wants integration tests.
---

# Test-Driven Development

TDD is the red → green loop. This skill is the reference that makes that loop produce tests worth keeping: what a good test is, where tests go, the anti-patterns, and the rules of the loop. Every section applies on every cycle: consult them before and during the loop, not after.

When exploring the codebase, read `CONTEXT.md` (if it exists) so test names and interface vocabulary match the project's domain language, and respect ADRs in the area you're touching.

## What a good test is

Tests verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't. A good test reads like a specification: "user can checkout with valid cart" tells you exactly what capability exists, and it survives refactors because it doesn't care about internal structure.

See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines.

## Seams: where tests go

A **seam** is the public boundary you test at: the interface where you observe behavior without reaching inside. Tests live at seams, never against internals.

**Test only at pre-agreed seams.** Before writing any test, write down the seams under test and confirm them with the user. No test is written at an unconfirmed seam. You can't test everything, so agreeing the seams up front is how testing effort lands on the critical paths and complex logic instead of every edge case.

Ask: "What's the public interface, and which seams should we test?"

When the shape of that interface is itself in question (how deep the module is, where the seam belongs, what the interface should expose), call the Skill tool with "codebase-design" for the vocabulary. It is the shared source of the module, interface, depth, seam, adapter, leverage and locality terms, and it is a reference to consult, not a session to run.

## Anti-patterns

- **Implementation-coupled**: mocks internal collaborators, tests private methods, or verifies through a side channel (querying the database instead of using the interface). The tell: the test breaks when you refactor but behavior hasn't changed.
- **Tautological**: the assertion recomputes the expected value the way the code does (`expect(add(a, b)).toBe(a + b)`, a snapshot derived by hand the same way, a constant asserted equal to itself), so it passes by construction and can never disagree with the code. Expected values must come from an independent source of truth: a known-good literal, a worked example, the spec.
- **Horizontal slicing**: writing all tests first, then all implementation. Bulk tests verify _imagined_ behavior: you test the _shape_ of things rather than user-facing behavior, the tests go insensitive to real changes, and you commit to test structure before understanding the implementation. Work in **vertical slices** instead: one test → one implementation → repeat, each test a **tracer bullet** that responds to what the last cycle taught you.

## Rules of the loop

- **Red before green.** Write the failing test first, then only enough code to pass it. Don't anticipate future tests or add speculative features.
- **One slice at a time.** One seam, one test, one minimal implementation per cycle.
- **Refactoring is not part of the loop.** It belongs to the review stage (see the `code-review` skill), not the red → green implementation cycle.
04 The Skill's file agents/openai.yaml
interface:
  display_name: "TDD"
  short_description: "Test-driven red-green-refactor"
05 The Skill's file mocking.md
# When to Mock

Mock at **system boundaries** only:

- External APIs (payment, email, etc.)
- Databases (sometimes - prefer test DB)
- Time/randomness
- File system (sometimes)

Don't mock:

- Your own classes/modules
- Internal collaborators
- Anything you control

## Designing for Mockability

At system boundaries, design interfaces that are easy to mock:

**1. Use dependency injection**

Pass external dependencies in rather than creating them internally:

```typescript
// Easy to mock
function processPayment(order, paymentClient) {
  return paymentClient.charge(order.total);
}

// Hard to mock
function processPayment(order) {
  const client = new StripeClient(process.env.STRIPE_KEY);
  return client.charge(order.total);
}
```

**2. Prefer SDK-style interfaces over generic fetchers**

Create specific functions for each external operation instead of one generic function with conditional logic:

```typescript
// GOOD: Each function is independently mockable
const api = {
  getUser: (id) => fetch(`/users/${id}`),
  getOrders: (userId) => fetch(`/users/${userId}/orders`),
  createOrder: (data) => fetch('/orders', { method: 'POST', body: data }),
};

// BAD: Mocking requires conditional logic inside the mock
const api = {
  fetch: (endpoint, options) => fetch(endpoint, options),
};
```

The SDK approach means:
- Each mock returns one specific shape
- No conditional logic in test setup
- Easier to see which endpoints a test exercises
- Type safety per endpoint
06 The Skill's file tests.md
# Good and Bad Tests

## Good Tests

**Integration-style**: Test through real interfaces, not mocks of internal parts.

```typescript
// GOOD: Tests observable behavior
test("user can checkout with valid cart", async () => {
  const cart = createCart();
  cart.add(product);
  const result = await checkout(cart, paymentMethod);
  expect(result.status).toBe("confirmed");
});
```

Characteristics:

- Tests behavior users/callers care about
- Uses public API only
- Survives internal refactors
- Describes WHAT, not HOW
- One logical assertion per test

## Bad Tests

**Implementation-detail tests**: Coupled to internal structure.

```typescript
// BAD: Tests implementation details
test("checkout calls paymentService.process", async () => {
  const mockPayment = jest.mock(paymentService);
  await checkout(cart, payment);
  expect(mockPayment.process).toHaveBeenCalledWith(cart.total);
});
```

Red flags:

- Mocking internal collaborators
- Testing private methods
- Asserting on call counts/order
- Test breaks when refactoring without behavior change
- Test name describes HOW not WHAT
- Verifying through external means instead of interface

```typescript
// BAD: Bypasses interface to verify
test("createUser saves to database", async () => {
  await createUser({ name: "Alice" });
  const row = await db.query("SELECT * FROM users WHERE name = ?", ["Alice"]);
  expect(row).toBeDefined();
});

// GOOD: Verifies through interface
test("createUser makes user retrievable", async () => {
  const user = await createUser({ name: "Alice" });
  const retrieved = await getUser(user.id);
  expect(retrieved.name).toBe("Alice");
});
```

**Tautological tests**: Expected value restates the implementation, so the test passes by construction.

```typescript
// BAD: Expected value is recomputed the way the code computes it
test("calculateTotal sums line items", () => {
  const items = [{ price: 10 }, { price: 5 }];
  const expected = items.reduce((sum, i) => sum + i.price, 0);
  expect(calculateTotal(items)).toBe(expected);
});

// GOOD: Expected value is an independent, known literal
test("calculateTotal sums line items", () => {
  expect(calculateTotal([{ price: 10 }, { price: 5 }])).toBe(15);
});
```
07 Records and fingerprints
Comparison
forgecmp_93a35d063f2b45d9b551f13ab6f4f0f9
Compared on
27 September 2026
Tasks
forgecol_ffa133338756407eba327098051f996c
Fingerprint of the tasks
sha256:5ce52d4aa2b5ba23ffdc5e7982a508b7ad9bc938f7f46a1b06feb324a7bab808
Fingerprint of the Skill
sha256:7dc0ee968fc1f3717b653a11b172f4c89b080edb0d4290d61b21135ffad14099

Published Results · Start