We measure, we don't guess
Small models. Local hardware. Measured against frontier systems.
Every number here comes from a benchmark I can reproduce on machines I own — and the failed approaches are reported next to the results, because a record that only lists wins is marketing. Same instrumentation habit behind a study of chain-of-thought faithfulness across 19 frontier models.
Four claims, each measured against the system it's up against. Open one for the method, the head-to-head, and what failed on the way.
Legal reasoning · MBE bar exam · Harvey AI has never published a score
Cadence
MBE bar exam under an adversarial protocol — a ~14B open model on one 4090.
98.3%
Legal reasoning · MBE bar exam · Harvey AI has never published a score
Cadence
MBE bar exam under an adversarial protocol — a ~14B open model on one 4090.
A ~14B open model wrapped in retrieval, deterministic citation extraction, and an adversarial checker, over a million-opinion, 50-state corpus — running locally on one RTX 4090, not on rented frontier inference. The 28-point lift over the base model came from pipeline architecture, not parameters.
| System | MBE score | Hardware | Funding |
|---|---|---|---|
| Cadence (local) | 98.3% | 1× RTX 4090 ($2K) | $0 |
| DescrybeLM * | 100% | Cloud API | startup VC |
| ChatGPT 5.2 | 93.4% | Azure datacenters | $13B+ |
| Gemini 3 Pro | 91.5% | Google TPU fleet | |
| Claude Opus 4.5 | 89.0% | AWS clusters | $8B+ |
| Harvey AI | no published score | Cloud API | $1B+ · $11B val |
| CoCounsel | no published score | Cloud API | $200M+ acq. |
| Lexis+AI | no published score | Cloud API | LexisNexis |
| GPT-4 (2023) | 75.7% | Azure datacenters | $13B+ |
| Kimi K2.5 (1T params) | ~70% | 240GB+ RAM min | Moonshot AI |
| Human avg (law grads) | 68% | a brain | $200K+ debt |
Cadence, ChatGPT 5.2, Gemini 3 Pro and Opus 4.5 are my own run on one 117-question MBE set (March 2026). Independently, a March-2026 study put the same general-purpose models at 88.5–93.5%. * DescrybeLM's 100% in that study was self-evaluated and AI-judged on its own question set, with training-data overlap not ruled out. GPT-4 scored 75.7% on the same exam in 2023; the human average is 68%. Harvey / CoCounsel / Lexis have never published an MBE score.
Execution latency · on-chain verified against competitor marketing claims
Same-slot execution engine
of 1,817 trades land in the leader's own block — at a $0.001 fee, over a home connection.
40.2%
Execution latency · on-chain verified against competitor marketing claims
Same-slot execution engine
of 1,817 trades land in the leader's own block — at a $0.001 fee, over a home connection.
A Solana copy-execution pipeline built and instrumented end-to-end — signal, decision, transaction build, network — running on consumer hardware over a home connection. No VPS, no colocation. A slot is ~400ms; a same-slot fill lands in the very block of the trade it copies — the fastest fill physically possible. Every figure below settles on a public ledger; verification requires a block explorer, not my word.
| Bot | Effective cost | Priority fee | Latency | vs this engine |
|---|---|---|---|---|
| GMGN AI | ~3–5% | ~$1.20 | 1–3s (claim) | 5.3× slower |
| Trojan | ~4–5% | ~$3.50 | <500ms (claim) | 2.6× slower |
| Axiom | ~3–6% | ~$2–10 | ≤400ms (claim) | 2.1× slower |
| This engine | 0% | $0.001 | 190ms (on-chain) | 40.2% same-slot |
Platform fees run ~1% on top of inflated priority fees and Jito tips — 3–6% effective on small trades. Competitor latencies are the bots' own marketing claims; mine settle on-chain. No Telegram bot reliably lands in the leader's slot — this one does, 40.2% of the time.
Poker decision quality · solver-graded head-to-head against frontier models, identical spots
Qwen3-4B exploit adapters
same-spot record vs Gemini 3 Flash — every model graded on the same hands by the same solver.
72–5
Poker decision quality · solver-graded head-to-head against frontier models, identical spots
Qwen3-4B exploit adapters
same-spot record vs Gemini 3 Flash — every model graded on the same hands by the same solver.
Per-street LoRA adapters on a 4B open model, trained on solver-graded contested spots — 856 converged postflop decisions, every model graded on the same hands by the same solver. Trained on one consumer 4090.
Animal pose estimation · benchmarked against DeepLabCut SuperAnimal (EPFL)
26-keypoint canine pose model
mAP50 on the domain benchmark — a specialist beating the award-winning generalist.
.858
Animal pose estimation · benchmarked against DeepLabCut SuperAnimal (EPFL)
26-keypoint canine pose model
mAP50 on the domain benchmark — a specialist beating the award-winning generalist.
A custom 26-keypoint canine pose model, trained for a production behavior-telemetry pipeline that tracks dog and handler skeletons in the same frame — the interaction, not just the animal.
Pawly behavior telemetry
dual-skeleton · liveConsumer perception pipeline: pose, segmentation, and bark-event detection over ordinary phone video, with a certified trainer supplying ground-truth labels in the loop. Demand side grounded in a published 228k-discussion study.
Self-hosted semantic memory
86,942 chunks · sub-second recallSemantic memory over two years of working sessions — local 4090 inference, queryable from any session. The continuity layer everything else runs on.
Garden rover
<200ms glass-to-glassGarage-built garden rover: open-vocabulary detection, tracked segmentation, custom-trained probe detector, geometric behavior states — perception to display in under 200 milliseconds on a Raspberry Pi camera feed.
Unified inference daemon
owned GPUs · one HTTP surfacePose, segmentation, depth, 3D body, face analysis, and embedding models consolidated behind a single local endpoint — the reason marginal inference cost across every project above rounds to zero.
Empirical study · DOI 10.5281/zenodo.21495752 · 6.62M decisions
The Say/Do Gap of Frontier Language Models
How much a model's visible reasoning predicts its committed action — measured judge-free across 19 models and 27,750 traces, with match-grouped splits, permutation nulls, and bootstrap intervals. The link is strong for a few models, weak for many, and it is a per-model number nobody publishes.
Essay · DOI 10.5281/zenodo.20768819
Robots Aren't Ten Years Away
The bottleneck was never the body — it was perception, and it broke. Deployment timelines offered as falsifiable predictions, with a stated note on what the numbers are and aren't.
Empirical study · 228,753 discussions
What Dog Owners Are Actually Dealing With
A frozen-instruction LLM read over 87 public communities, with per-batch grounding audits (85–96% exact-match, reported not assumed) and a full statement of limitations.