Not a leaderboard

The Harness Ledger.

Everyone here is building an AI system: memory, autonomy, retrieval, agents, security, the rails that connect it to real work. Nobody's is shaped like anyone else's. This is the record of those shapes, graded the same way, receipts first. Nobody wins a scorecard. The differences are the point.

Twelve dimensions. Every harness, the same grid.

  1. D1Ingress + voice

    Channels a human can reach it through; live voice loop.

  2. D2Autonomy + scheduling

    Useful work with nobody in the chair, and dead work gets loud.

  3. D3Memory + provenance

    Can it prove where a fact came from; does history survive updates.

  4. D4Retrieval, measured

    Hybrid search quality, and whether anyone actually benchmarked it.

  5. D5Skill library

    Packaged capability, kept trustworthy with linting and install gates.

  6. D6Self-improvement

    Does it get better on its own, and is that safe.

  7. D7Evals + dispatch

    Is 'the right capability fired and the output was good' measured or vibes.

  8. D8Multi-agent

    Fan-out, pipelines, cross-model adjudication, quarantine tiers.

  9. D9Security

    Injection defense, blast-radius control, secrets, supply chain.

  10. D10Compute sovereignty

    Local models, own hardware, no single vendor able to switch it off.

  11. D11Business integration

    Finance, CRM, legal, client delivery, approval-gated publishing.

  12. D12Ecosystem + momentum

    Who else improves it while the operator sleeps.

The record

One harness on the board. Yours is next.

#001

Lucas Cooper-Bey

8/12 leadsavg 7.92026-07-22

Lucface IOS · one-person Intelligent OS

5
D1
9
D2
9
D3
9
D4
9
D5
8
D6
7
D7
9
D8
9
D9
9
D10
10
D11
2
D12

Full-stack institution: provenance, security, and business rails run deep; voice and ecosystem are thin by choice.

Ghost vs now · Scored 7/12 that morning (evals were a 4). Closed the biggest gap the same day and re-graded that night. The delta is the point of the method.

How it's graded

Receipts before scores. Measured beats claimed.

  1. 01

    Receipts before scores. Inventory the system by looking, not by asking.

  2. 02

    Measured beats claimed. A benchmark number outranks an architecture that should work.

  3. 03

    N/A is never zero. Out-of-scope-by-design is graded N/A and left out of the lead count.

  4. 04

    No flattery. Every trailing dimension is either a deliberate choice or a real gap with the smallest artifact that would close it.

  5. 05

    Zero-regression honesty. Re-grades never quietly lower a number.

Grade your own harness. It takes one prompt and a bit of honesty.

Run the paste-in prompt against your machine, get a dossier of your own, and post the scorecard. It lands on this board next to everyone else's. Bring it to a Sharathon and we'll read the shapes together.