Example comparison · 27 September 2026
tdd vs No Skill
A local comparison of the tdd Skill against no Skill, reported by the person who ran it.
Keep this Skill change? Improved
On the 3 held-out tasks, the agent did better with the Skill on 1 task and worse on none.
Judged on the held-out tasks alone: tasks kept from any agent that tries to improve the Skill, so the answer does not rest on tasks it could learn from.
On average 0.667 without the Skill, 1 with it (+0.333). Averages are rounded to three decimal places.
- Better on 1 task: retry-backoff-boundary
- Worse on no tasks.
- No change on 2 tasks: reroute-public-seam, shipping-oracle
Tasks the improving agent could see
Improved on the 3 other tasks
An agent that tries to improve the Skill from this comparison is shown these tasks, so they are reported here on their own and do not decide the answer above.
On average 0.333 without the Skill, 1 with it (+0.667). Averages are rounded to three decimal places.
- Better on 2 tasks: exchange-rate-boundary, deactivate-account-interface
- Worse on no tasks.
- No change on 1 task: theater-cents-adr
Claims
What the Skill claims, and the tasks that test it
A positive case is one where following the Skill should give the right result; a boundary case sits at the edge of where the claim applies; a counterexample checks that the Skill is not overused where it would give a wrong result or should change nothing.
C1. Tests target public behavioral seams rather than private methods, internal collaborators, or side channels, so they remain valid across behavior-preserving refactors.
What would show it: The submitted tests exercise and observe the feature only through exported interfaces and still pass when the implementation is replaced by an equivalent version with different private structure and storage.
-
reroute-public-seam No change
Implement and test rerouting an undispatched parcel to a new destination in a small Python parcel service.
Positive case · Held out · time limit 30 minutes
- Without the Skill
- 1
- With the Skill
- 1
- Change
- 0
-
deactivate-account-interface Better with the Skill
Add account deactivation to a small Python accounts package backed by SQLite, with tests.
Boundary case · Seen by the improving agent · time limit 30 minutes
- Without the Skill
- 0
- With the Skill
- 1
- Change
- +1
C2. Tests use an independent source of truth for expected results instead of recomputing expectations with the production algorithm.
What would show it: The submitted tests pass against the correct implementation and fail against plausible algorithm mutants, including mutants that would also affect expectations obtained by calling production code twice.
-
shipping-oracle No change
Implement and test the shipping quote rules described for a local command-line fulfillment package.
Positive case · Held out · time limit 30 minutes
- Without the Skill
- 1
- With the Skill
- 1
- Change
- 0
C3. Tests substitute dependencies at system boundaries such as time or external services while using real internal collaborators.
What would show it: The tests control the external boundary deterministically, contain no doubles for project-owned collaborators, and continue to pass after those collaborators are reorganized without changing public behavior.
-
exchange-rate-boundary Better with the Skill
Add a converted price quote with a conversion fee to a small Python pricing package whose exchange rates come from a remote rates service, with tests.
Positive case · Seen by the improving agent · time limit 30 minutes
- Without the Skill
- 0
- With the Skill
- 1
- Change
- +1
-
retry-backoff-boundary Better with the Skill
Add retries with growing waits to a small Python notifier that sends messages through a mail service client, with tests.
Positive case · Held out · time limit 30 minutes
- Without the Skill
- 0
- With the Skill
- 1
- Change
- +1
C4. Code and tests follow the repository's established domain vocabulary and applicable architectural decisions.
What would show it: The resulting public API, test descriptions, values, and persisted representations use the terms and constraints defined by the repository guidance rather than generic alternatives.
-
theater-cents-adr No change
Add seat holds and the amount owed for them to a small Python theater-sales package, with tests.
Positive case · Seen by the improving agent · time limit 30 minutes
- Without the Skill
- 1
- With the Skill
- 1
- Change
- 0
The Skill change
What differed between the two runs
The two runs used the same tasks, model and limits. One ran with no Skill; the other ran with this Skill, the one the tasks were written from.
- Without the Skill
- No Skill
- With the Skill
- tdd
- Licence
- MIT License. Copyright (c) 2026 Matt Pocock
- Fingerprint
- sha256:7dc0ee968fc1f3717b653a11b172f4c89b080edb0d4290d61b21135ffad14099
The Skill's files, its licence among them, are below, exactly as the run with the Skill used them. Read the Skill's files.
Evidence
What stands behind these numbers
-
Files checked
This site checked the files behind this page against each other: each of the Skill's files and the Skill's fingerprint, both runs' records, the export's tasks and instructions, and the decision, worked out again from the task scores. Every check passed.
-
Reported by the person who ran it
One local comparison, run on one computer by the person who reported it. It is not a published Result, nothing about it is signed, and nobody else watched the runs.
-
Not yet reproduced
This site has no record of anybody else running this comparison again.
Method
How the numbers were made
- What this is
- One local comparison on one computer. It is not a published or signed Result.
- Tries
- Each task was tried once without the Skill and once with it.
- Model
- gpt-5.6-sol from OpenAI, with medium reasoning, as both runs asked for it
- What limited the runs
- Time: 30 minutes for each try. Each try ran in its own container with 2 processor cores, 4 GB of memory and no network. Nothing limited the agent's turns, model calls or tokens.
- Model calls and tokens
- Without the Skill 41 model calls, 338,454 tokens. With the Skill 95 model calls, 1,026,256 tokens.
What these runs cannot show
- Which model the provider actually served. Only the provider's own report, passed on by Hermes, says which model answered.
- That the Hermes agent program was unmodified. Its version is only what it reports.
- How the model chose its answers. Those settings cannot be changed, so every try used the provider's defaults.
Check it yourself
Check the tasks and run this comparison again
- Fingerprint of the tasks
- sha256:5ce52d4aa2b5ba23ffdc5e7982a508b7ad9bc938f7f46a1b06feb324a7bab808
The first command below prints the fingerprint of the tasks it checks. They are the same tasks as these only if it matches this one.
You need
- The export folder of these tasks, with its tests and records. Use it only if the first command prints the fingerprint above; a folder with any other fingerprint holds other tasks.
- The Skill's files, shown on this page, in a folder of their own
- Docker, Techtree and Hermes, as the export folder's own instructions describe
- An account with a model provider you choose; its calls may cost money
Check it and run it again
techtree forge verify-export EXPORT_FOLDER
techtree forge import EXPORT_FOLDER
techtree forge inspect-skill SKILL_FOLDER
techtree forge run --arm baseline --collection forgecol_ffa133338756407eba327098051f996c --provider PROVIDER --model MODEL
techtree forge run --arm candidate --collection forgecol_ffa133338756407eba327098051f996c --provider PROVIDER --model MODEL --skill SKILL_FOLDER
techtree forge compare BASELINE_RUN_ID CANDIDATE_RUN_ID
Words in capitals stand for what only you know: your folders, the provider and model you choose, and the ids the two runs print. Each run shows what it will do and asks before it starts.
What a new run can tell you
- A new run is a new comparison. The model may not answer the same way twice, so its numbers can differ from these.
- Whether yours agrees is for you to judge. This site keeps no record that ties a new run to this one.
-
To test your own change to this Skill, give the baseline run the earlier version with
--skillinstead of no Skill, and the candidate run your new version. The comparison then says whether the change is worth keeping. How to compare two versions →
The evidence in full
Every task, the Skill's files and the fingerprints
01 All 6 tasks
- reroute-public-seam Held out Without the Skill1 With the Skill1 Change0
- shipping-oracle Held out Without the Skill1 With the Skill1 Change0
- exchange-rate-boundary Seen by the improving agent Without the Skill0 With the Skill1 Change+1
- deactivate-account-interface Seen by the improving agent Without the Skill0 With the Skill1 Change+1
- theater-cents-adr Seen by the improving agent Without the Skill1 With the Skill1 Change0
- retry-backoff-boundary Held out Without the Skill0 With the Skill1 Change+1
02 The Skill's file LICENSE.txt
MIT License Copyright (c) 2026 Matt Pocock Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
03 The Skill's file SKILL.md
--- name: tdd description: Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions "red-green-refactor", or wants integration tests. --- # Test-Driven Development TDD is the red → green loop. This skill is the reference that makes that loop produce tests worth keeping: what a good test is, where tests go, the anti-patterns, and the rules of the loop. Every section applies on every cycle: consult them before and during the loop, not after. When exploring the codebase, read `CONTEXT.md` (if it exists) so test names and interface vocabulary match the project's domain language, and respect ADRs in the area you're touching. ## What a good test is Tests verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't. A good test reads like a specification: "user can checkout with valid cart" tells you exactly what capability exists, and it survives refactors because it doesn't care about internal structure. See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines. ## Seams: where tests go A **seam** is the public boundary you test at: the interface where you observe behavior without reaching inside. Tests live at seams, never against internals. **Test only at pre-agreed seams.** Before writing any test, write down the seams under test and confirm them with the user. No test is written at an unconfirmed seam. You can't test everything, so agreeing the seams up front is how testing effort lands on the critical paths and complex logic instead of every edge case. Ask: "What's the public interface, and which seams should we test?" When the shape of that interface is itself in question (how deep the module is, where the seam belongs, what the interface should expose), call the Skill tool with "codebase-design" for the vocabulary. It is the shared source of the module, interface, depth, seam, adapter, leverage and locality terms, and it is a reference to consult, not a session to run. ## Anti-patterns - **Implementation-coupled**: mocks internal collaborators, tests private methods, or verifies through a side channel (querying the database instead of using the interface). The tell: the test breaks when you refactor but behavior hasn't changed. - **Tautological**: the assertion recomputes the expected value the way the code does (`expect(add(a, b)).toBe(a + b)`, a snapshot derived by hand the same way, a constant asserted equal to itself), so it passes by construction and can never disagree with the code. Expected values must come from an independent source of truth: a known-good literal, a worked example, the spec. - **Horizontal slicing**: writing all tests first, then all implementation. Bulk tests verify _imagined_ behavior: you test the _shape_ of things rather than user-facing behavior, the tests go insensitive to real changes, and you commit to test structure before understanding the implementation. Work in **vertical slices** instead: one test → one implementation → repeat, each test a **tracer bullet** that responds to what the last cycle taught you. ## Rules of the loop - **Red before green.** Write the failing test first, then only enough code to pass it. Don't anticipate future tests or add speculative features. - **One slice at a time.** One seam, one test, one minimal implementation per cycle. - **Refactoring is not part of the loop.** It belongs to the review stage (see the `code-review` skill), not the red → green implementation cycle.
04 The Skill's file agents/openai.yaml
interface: display_name: "TDD" short_description: "Test-driven red-green-refactor"
05 The Skill's file mocking.md
# When to Mock
Mock at **system boundaries** only:
- External APIs (payment, email, etc.)
- Databases (sometimes - prefer test DB)
- Time/randomness
- File system (sometimes)
Don't mock:
- Your own classes/modules
- Internal collaborators
- Anything you control
## Designing for Mockability
At system boundaries, design interfaces that are easy to mock:
**1. Use dependency injection**
Pass external dependencies in rather than creating them internally:
```typescript
// Easy to mock
function processPayment(order, paymentClient) {
return paymentClient.charge(order.total);
}
// Hard to mock
function processPayment(order) {
const client = new StripeClient(process.env.STRIPE_KEY);
return client.charge(order.total);
}
```
**2. Prefer SDK-style interfaces over generic fetchers**
Create specific functions for each external operation instead of one generic function with conditional logic:
```typescript
// GOOD: Each function is independently mockable
const api = {
getUser: (id) => fetch(`/users/${id}`),
getOrders: (userId) => fetch(`/users/${userId}/orders`),
createOrder: (data) => fetch('/orders', { method: 'POST', body: data }),
};
// BAD: Mocking requires conditional logic inside the mock
const api = {
fetch: (endpoint, options) => fetch(endpoint, options),
};
```
The SDK approach means:
- Each mock returns one specific shape
- No conditional logic in test setup
- Easier to see which endpoints a test exercises
- Type safety per endpoint
06 The Skill's file tests.md
# Good and Bad Tests
## Good Tests
**Integration-style**: Test through real interfaces, not mocks of internal parts.
```typescript
// GOOD: Tests observable behavior
test("user can checkout with valid cart", async () => {
const cart = createCart();
cart.add(product);
const result = await checkout(cart, paymentMethod);
expect(result.status).toBe("confirmed");
});
```
Characteristics:
- Tests behavior users/callers care about
- Uses public API only
- Survives internal refactors
- Describes WHAT, not HOW
- One logical assertion per test
## Bad Tests
**Implementation-detail tests**: Coupled to internal structure.
```typescript
// BAD: Tests implementation details
test("checkout calls paymentService.process", async () => {
const mockPayment = jest.mock(paymentService);
await checkout(cart, payment);
expect(mockPayment.process).toHaveBeenCalledWith(cart.total);
});
```
Red flags:
- Mocking internal collaborators
- Testing private methods
- Asserting on call counts/order
- Test breaks when refactoring without behavior change
- Test name describes HOW not WHAT
- Verifying through external means instead of interface
```typescript
// BAD: Bypasses interface to verify
test("createUser saves to database", async () => {
await createUser({ name: "Alice" });
const row = await db.query("SELECT * FROM users WHERE name = ?", ["Alice"]);
expect(row).toBeDefined();
});
// GOOD: Verifies through interface
test("createUser makes user retrievable", async () => {
const user = await createUser({ name: "Alice" });
const retrieved = await getUser(user.id);
expect(retrieved.name).toBe("Alice");
});
```
**Tautological tests**: Expected value restates the implementation, so the test passes by construction.
```typescript
// BAD: Expected value is recomputed the way the code computes it
test("calculateTotal sums line items", () => {
const items = [{ price: 10 }, { price: 5 }];
const expected = items.reduce((sum, i) => sum + i.price, 0);
expect(calculateTotal(items)).toBe(expected);
});
// GOOD: Expected value is an independent, known literal
test("calculateTotal sums line items", () => {
expect(calculateTotal([{ price: 10 }, { price: 5 }])).toBe(15);
});
```
07 Records and fingerprints
- Comparison
- forgecmp_93a35d063f2b45d9b551f13ab6f4f0f9
- Compared on
- 27 September 2026
- Tasks
- forgecol_ffa133338756407eba327098051f996c
- Fingerprint of the tasks
- sha256:5ce52d4aa2b5ba23ffdc5e7982a508b7ad9bc938f7f46a1b06feb324a7bab808
- Fingerprint of the Skill
- sha256:7dc0ee968fc1f3717b653a11b172f4c89b080edb0d4290d61b21135ffad14099