Cooked runs on a proprietary stack that pairs Anthropic frontier models with classical machine learning and fine-tuned local models on our own hardware. What follows is the architecture. The thresholds, features, and signals stay ours.
Before the bell, multiple independent frontier-model calls read the same tape, each biased toward a different way of seeing it: momentum, mean reversion, macro, structure. None of them sees the others' work. A final synthesis call argues them into one plan.
One model with one opinion is a guess. A panel that has to survive disagreement is a process.
Every job gets routed to the cheapest model tier that can actually do it. Fable-class reasoning takes the panel and the audits. Sonnet-class handles session work. Haiku-class burns through classification at speed. The governor tracks spend per call and refuses waste.
Bulk work never leaves the building. Local models served on our own hardware handle summarization, extraction, and classification around the clock at zero marginal cost, and we fine-tune adapters on our own trade history so the fast lane keeps learning our tape, not the internet's.
Classical machine learning does the unglamorous work: engineered features off every session's tape, nightly dataset builds, and pattern reports that get re-scored each morning against what actually happened. No feature survives on a story. It survives on out-of-sample behavior.
Every trading idea enters as a logged hypothesis with a testable claim. It earns its way from hypothesis to simulated rulebook to promotion, or it dies in the log with the reason written down. The log never forgets, which keeps us from falling for the same pretty idea twice.
Proposed rule changes get stress-tested against stored market history with pessimistic assumptions: worst-touch fills, modeled spread costs, never mid-marked. If a change only works with generous assumptions, it doesn't work. The tuner has killed more good-looking ideas than any human on the desk.
Every session is recorded, replayed, and audited by a frontier model against the locked spec: what the rules said, what the operator did, where the two disagreed, and what it cost. The audit is adversarial by design. It gets graded on what it catches, not on how nice it is.
An idea can fail out at any stage, and most do. The research board shows the survivors and the bodies.
Real lessons from the audit log. Specifics are redacted because the lessons are the product.
We ran a mirror book that took the exact opposite side of every core signal. If our signals were noise it would have gone sideways. It lost ██% of its seed in five weeks. Direction confirmed; the edge was in the exits, so that's where the work went.
The audit showed we were capturing only ██ to ██% of each trade's maximum favorable move. The entry engine wasn't the problem. The exit engine was rebuilt around ███████ ████████ and the capture rate is now the desk's most-watched number.
Removing a daily throttle looked brilliant for ██ sessions in simulation. The stored-history replay with pessimistic fills said the uncapped version was net negative. The throttle stayed. The lesson is that the boring rail was load-bearing.
A stale quote once printed a phantom exit worth +$█,███. It never happened; the book was just old. Every mark now carries a freshness gate of ███, and no fill prints off a ghost. Data hygiene is a trading edge nobody brags about.
Positions priced above a certain premium bled an average of -$███ per trade to time decay before the thesis could even play. A hard premium cap fixed what better timing never could. Sometimes the edge is refusing to pay the wrong price for being right.
You can read every note on this page and still not have the system. That's on purpose.