Skip to the content
Georgi DimitrovdaTuzzo

Polymarket research desk

A fail-closed trading desk where every strategy meets frozen rules and live order books

Role
Owner and research lead; agents built it under fail-closed rules
Status
In progress
Source
Private repository
Stack
PythonTypeScriptNext.jsSQLitesystemdDockerHermes Agent (NousResearch)GitHub ActionsChromium extension (MV3)
Private

In numbers

15

strategy families tested against live order books by parallel research lanes

3,072

configurations searched in one sweep, each candidate frozen before it was scored

3

conditions that must hold at once before a live order: live mode in config, an ARMED flag file and no HALT file

0

language models in the money loop

1s

tick of the money-loop daemon, which records spot prices, order books and the trade tape

The problem

Short-duration prediction markets settle every few minutes, which invites a quick conclusion from one good streak. I wanted a desk where a strategy reaches real money only after it passes rules written before the test, and where the part that trades is safe from the agents that help build and run it, including when I push them hardest.

The approach

I split the system by authority. The money loop is a deterministic daemon with no language model in it. Agents work on research and on a control plane that can observe and advise and holds no trading power by default. The statistics were written down before any learning loop existed, and candidate rules are frozen before they are scored. Claude and Codex agents built it in parallel worktrees, each run ending with a written report, test counts and an independent read-only review. On one handoff, Codex treated five Claude-built PRs as work from another author instead of trusting their completion reports: it merged them in a disposable worktree and ran the combined suite, 1,117 passed and 90 skipped, before anything reached main.

How it works

  1. A money loop with no model in it

    A Python daemon under systemd ticks once a second and records spot prices, order books and the trade tape to SQLite. In shadow mode it quotes virtual orders, counts a fill only when the tape prints through its price, and settles at resolution. Going live needs three things at once: live mode in config, an ARMED flag file and no HALT file. Touching HALT cancels quotes, a daily-loss stop halts on its own, and cancels go by order id because the account can hold manual orders. The operator extension arms a pod only after the exact typed phrase ARM followed by the pod name, and a single-instance guard went in after two daemons once drove the same pod.

  2. A control plane that fails closed

    My agent, Jeisan (it has its own page), runs on NousResearch's Hermes Agent on a separate host and reaches the pods only through a chain of least-privilege hops that ends at a strict server-side parser. The only production mutation is performed by a root systemd timer with no model in it. A live start would need an expiring, non-replayable mandate, root-attested config hashes and fresh signed exchange evidence, and the policy and mandate ship disabled, so installing the control plane cannot arm or trade. Separate watch jobs hold a sticky HALT that overrides the manager.

  3. Statistics before the learning loop

    Before building a loop that retunes itself, I wrote the statistical guardrails down as an issue and a document. They set how much evidence a change must carry before it counts, and the checks are pure functions in code. The decision layer is limited to operational facts and pre-registered experiments.

  4. A research fleet that checks itself

    Four parallel Opus lanes researched at once and checked each other's work. A backtest from one lane was tested against another lane's live, millisecond-timestamped book recording, and the check exposed two data traps that then went to every lane: stale book snapshots, and trade timestamps that are block time, about two seconds late. Candidate rules were frozen and hash-stamped before scoring, then scored once against a pre-registered bar.

  5. A simulator that has to prove its fills

    A Codex research sprint reworked the simulator so that each simulated fill has to prove its price was tradable, using exchange timestamps and trading status. The fill, quote-age, fee and settlement accounting are pinned by 70 execution assertions, 50 trading-invariant scenarios and 26 tests for time-weighted average prices. A secret scanner in CI blocked two research bundles from publishing.

What I chose, and what lost

Chose

No language model in the money loop

Over

An agent that trades and retunes itself

A model in the loop adds a failure mode with real money behind it, and the written statistics hold any retuning to pre-registered experiments. Agents research, review and advise; a deterministic daemon trades.

Chose

Freeze and hash candidate rules, then score them once against a pre-registered bar

Over

Iterate on the same data until a backtest looks good

Iterating on one data set until a backtest looks good fits noise. One sprint searched 3,072 configurations, and each candidate that went forward was frozen first and scored once, which is what makes any answer believable.

Outcome

Results stay private: balances, positions, fills and which strategies survived are not on this page, and no figure here is a return. The strategies that passed run in pods behind the kill switches above: shadow until armed, armed only by the typed phrase, halted by a file or by the daily-loss stop. The engineering is checkable: a control plane that ships disarmed, a simulator whose fills must prove their prices, and a written method for testing the next idea before it touches money.

Field notes

From the trading desk

Two copies of the same AI, fighting

I once lost an afternoon to a fight between two copies of the same AI, and I had set up both.

Two Claude Code sessions shared one server running my automated trading daemons. The first had started as a question about reaching a site from my phone, then appointed itself guard of the trading directory and armed its own monitors. I had stopped using it hours earlier and moved to a fresh session to keep building. The new session knew a parallel one existed. It did not know the old one still had alarms feeding it, or what that one counted as an intrusion.

The new session deployed a websocket module. The old guard found a file it had not written and moved it into a quarantine folder. Hours later the new session redeployed the same file and stood up a second pod. An alarm woke the guard, which decided an intruder was deploying live and halted the fleet. I pressed START in the console, and it read that as the intruder fighting back. So it stopped the services, killed the processes, folded the new pod's directory into a timestamped folder and deleted its service unit.

No money moved and nothing was lost. When I found the culprit, my new session sent it a stand-down message. It apologised and left an exact list of what it had moved. Everything was restored and hash-verified. The repo now carries one rule: one operator per box.

One operator per box.