How to read this page. Environments are never blended:
HISTORICAL = analysis of already-recorded data ·
SHADOW = signals logged live with no orders placed ·
FUNDED LIVE = real fills with real (bounded) money.
Backtests are not live results; shadow execution is not funded execution; past performance is not predictive.
Hypothetical and simulated figures have inherent limitations. Sample sizes and error bars are stated because they matter.
Active streams
PROTOCOLTWO-WEEK RUN VOIDEDINSTRUMENT STILL OPEN
Call-out journal — the run was under-powered, the instrument is not falsified
MNQ CALL-OUT PROTOCOL v1.2.0 · MEASURED: 2026-08-15 · RESEARCH COST: $0 (PAPER)
Setup: a human-in-the-loop discretion instrument. The rule surfaces a call-out, the operator answers take or kill, and that decision is sealed into a hash-chained journal before the outcome is known. The two-week protocol required that ≥80% of trading days carry a complete chain.
What was falsified: the two-week run, not the idea. The rule fires 0.368× per trading day — 7 signals across ~19 trading days — so at most ~37% of trading days can carry a complete chain. The exit criterion was not merely hard to reach, it was arithmetically unreachable from the day it was written.
What it would actually take: two weeks yields ~3.7 call-outs. Distinguishing a 55% hit rate from the 33.3% break-even at 1:1.5 reward-to-risk needs roughly 70 of them — about 8.8 months. That is the honest horizon for this instrument, and it remains open on it.
Why it is recorded here rather than in the graveyard: nothing about the execution failed, and the operator’s discretion — the actual variable under test — has never been measured. What died was a measurement design that could not have concluded at any outcome. The cheapest falsification in this ledger is still the one done by dividing the firing rate into the exit criterion before the fortnight is spent.
One stream is live. Two that stood here until 2026-07-22 were killed and have been moved to
the graveyard below — including one whose interim reads were positive and whose terminal result was not.
Leaving those cards up would have been the exact failure this ledger exists to prevent.
SHADOWTESTING
T1 — opening-range continuation, CME micro futures
MARKET: CME MICRO EQUITY FUTURES · FORWARD SHADOW SINCE: 2026-07-11 · DECISION RULE FROZEN: 2026-07-24 · GATE READS: FIRST SESSION ON/AFTER 2026-09-11 WITH n≥40 · LAST UPDATED: 2026-07-30
Hypothesis: the production charting-platform strategy, rebuilt engine-for-engine in Python, holds up forward under a pass/fail rule written down before the first data point.
Status: ~0.7 fires per session as of 2026-08-06, unchanged in character since the window opened. At that pace the n≥40 sample is projected to arrive late September or early October — the trigger date will likely arrive before the sample does. Zero orders. Talos receives nothing from this stream until the gate passes — and as of 2026-08-05 the execution layer replays every one of this stream's signals daily through its full order path with the network disconnected, purely to keep the path proven. That rehearsal immediately refused all of them: this strategy's stop distances exceed the operator's per-trade risk limit, a mismatch found with no capital at stake.
The decision rule was hunted before it was frozen: seven adversarial review rounds, during which three buy-easier defects introduced by the author’s own hardening edits were caught and removed. Binary pass/fail; a pass authorises only a separately pre-registered next step, never a trade.
Known limitation, stated plainly: the backtested edge is +$1.72 per trade across n=399 — statistically indistinguishable from zero, positive only in the 2026 regime. A pass at n=40 would rule out a large negative and little else. Separately, this signal’s stops run $620–796 per contract, which is 62–80% of a $1,000 daily loss limit on a single contract.
Reproduction fidelity
HISTORICALINDEPENDENT REPRODUCTION
Charting-platform strategy → independent Python port
COMPLETED: 2026-07 · METHOD: ENGINE-FOR-ENGINE REBUILD, CROSS-CHECKED AGAINST THREE INDEPENDENT DATA FEEDS
Before any verdict from the futures stream is trusted, the strategy was reproduced outside its original platform and cross-checked against three independent data feeds — reproducing recorded live trades to the tick. Reproduction fidelity means a verdict can't hide in a vendor's engine quirks; it does not by itself make the strategy profitable.
The graveyard — falsified and rejected
SPECIFICATIONGRAVEYARDED — NEVER REGISTERED
Combine decision instrument — seven adversarial rounds, no convergence
AUTHORED BLIND: 2026-08-06 · GRAVEYARDED: 2026-08-07 · SEVEN HUNT ROUNDS · RESEARCH COST: $0
Setup: a specification intended to decide whether to buy a funded-account evaluation. Registration required a hunt round returning zero confirmed critical or major findings. Each draft was attacked adversarially before any data was spent.
Result: confirmed critical/major findings by round ran 12 → 6 → 3 → 6 → 6 → 21 → 12 — 33 in total, and no convergence. The instrument was graveyarded without ever being registered.
The failure worth publishing: round six was not a modelling error, it was an evidence error. A sensitivity sweep’s committed output had been truncated by a shell pipe (tee | head), silently losing the final block — which contained the sweep’s only BUY cell. Four quantitative claims drawn from that artifact were false, including the headline “NO-BUY in every one of the 72 cells” and a stated maximum of 50.5% where the true figure was 68.5%.
Why that is the more dangerous failure: the lost BUY cell sat at a state the frozen trigger cannot reach. So the conclusion was right and its evidence was wrong — which is worse than being wrong outright, because nothing downstream would ever have contradicted it. Round seven then found the replacement pool modelled a book trading on fewer days than one of its own legs, diluting that leg roughly 14.5×.
What changed as a result: the artifact contract now requires an explicit end-of-sweep marker, and a capture missing that final line is invalid by definition — so truncation is detectable rather than silent.
Honest limit of this graveyarding: the underlying NO-BUY conclusion was never falsified — the flawed and corrected models agree across realistic per-trade means. What failed was every attempt to build an instrument that could defend it. A future combine question needs a materially different portfolio and a fresh blind-authored specification, and may not cite this document as precedent.
SPECIFICATIONNOT REGISTERED
Charting-platform strategy 0.2.2 — measured, and refused registration anyway
MNQ SUBSTRATE STRATEGY · MEASURED: 2026-08-19 · REGISTRATION REFUSED · RESEARCH COST: $0
Setup: ten defects were repaired in a charting-platform implementation and the risk parameters aligned to the frozen specification. The measured result improved: profit factor moved 0.76 → 0.88 and net recovered $542 with drawdown cut by a third.
Result: still underwater, and not registered. A 1:1.5 reward-to-risk rule needs a 40% hit rate to break even; this sits at 40.21%, and about $388 of commission pushes it under. The edge is smaller than its own costs.
Why registration was refused: reading the implementation against the frozen specification line by line, it fires on an unregistered second mechanism, takes unlimited trades per session against a cap of three, decides across the whole session rather than the declared window, and runs on five-minute bars where the specification declares one. So the measurement does not describe the registered signal at all — it describes a looser strategy that happens to share its stop distance. Recording it as a falsification would have closed a hypothesis that has never actually been tested.
NAMINGTWO SPECS, ONE NAME
T1_ORCONT — a name that referred to two different strategies
FOUND: 2026-08-20 · RENAMED FORWARD, HISTORY LEFT INTACT · RESEARCH COST: $0
What was found: two strategies were carrying the same label. The forward shadow stream stops at the far opening-range side ± 0.25×ATR14 for a 0.75R target, once a day, on five-minute bars. The pre-registered specification of the same name stops at a fixed 12 points for an 18-point target, VWAP-aligned, up to three times a session, on one-minute bars. Risk per contract differs by roughly 17× — $331–$796 against about $24.
Why it matters: the forward evidence counter that gates promotion counts the shadow stream. It is therefore not evidence for the pre-registered specification, whose own gates are unrelated and still open. Reaching the forward threshold would not license the registered signal to trade.
And the wide variant is already refused: the execution layer declines it at one contract by design — per-trade risk above roughly 30% of the account’s daily loss limit is rejected, and viable strategies on this frontier run about $120–150 of risk. Twenty of twenty-two signals were refused on sizing. That is the risk policy working, not failing.
Fix: the shadow stream was renamed forward only. Historical rows keep the old label, because the log is the evidence record and relabelling a dated observation would falsify it.
REPLICATIONFALSIFIED
Intraday momentum — published effect, tested on CME micro futures
MECHANISM FIXED BY PUBLISHED LITERATURE · REGISTERED BEFORE ANY CODE EXISTED: 2026-08-05 · ONE RUN · RESEARCH COST: $0
Setup: a documented effect — the first half-hour of the session predicting the last half-hour — taken verbatim from the literature so that no parameter belonged to us. Registration was committed before a line of implementation was written; the criteria required statistical significance, not merely a positive sign, and survival at doubled costs.
Result: falsified on all three criteria. Mean −$3.53 per trade with a t‑statistic of −0.65 across 483 sessions, positive in only 2 of 4 half‑year sub‑periods, and worse again at doubled costs. The published effect does not reproduce here in this era at retail cost.
Worth recording: the full‑day variant came back significantly negative (t = −2.16) — these afternoons look reversal‑shaped rather than momentum‑shaped. That is an observation, not a finding: it was seen after the fact, so acting on it would require its own pre‑registration, disclosed as no longer blind.
FUNDED LIVE — FAILEDGRAVEYARDED
CRYPTO15M — 15-minute crypto event contracts
MARKET: KALSHI CRYPTO 15-MIN SERIES · LIVE ON RECORD: 2026-07-14 · ENDED BY OPERATOR: 2026-07-22 · LIVE TUITION: −$9.09
Setup: operator-enabled bounded experiment — 1 contract fixed, max 3 concurrent, hard −$4/day stop. Enabled on record, never endorsed by the gate.
Result: falsified. The estimate went negative at −1.26c per contract on n=724, realising −$9.09 over eight days. The earlier interim read (+1.73c at n≥200) was early-cohort flattery — the third appearance of that pattern after FAVBAND and TAKERFOLLOW.
What actually worked: not the entry signal. The exit-stop engine saved roughly $61 with 2 escapees in 198, and a 0-of-48 counterfactual on stopped positions. Across this entire programme the stops are the only component that has ever beaten its counterfactual.
SHADOWBLOCKED — ECONOMICALLY DEAD
SPEC_EVENTSUM — sum-constraint arbitrage on event contracts
MARKET: KALSHI EVENT CONTRACTS · REGISTERED: 2026-07-12 · STREAM KILLED: 2026-07-22 · RESEARCH COST: $0 (SHADOW)
Result: the successor was shown to be economically dead by construction, verified by an adversarial design-and-attack review. The partition-check fix and the dutch-book edge are mutually exclusive on this venue: on a genuinely exhaustive cover the maker enforces sum(YES ask) ≥ 100 by arithmetic, so no dutch book exists — and certifying exhaustiveness removes precisely the region where every apparent edge lived. The deep discount was a non-exhaustiveness detector, not a mispricing.
Status: registered BLOCKED — a measurement-only null instrument, not a promotion candidate. The stream was stopped rather than left running to look busy.
PRE-REGISTRATIONKILLED BEFORE ANY DATA
Two research directions killed at pre-registration — zero data spent, $0
MNQ FREQUENCY AXIS: KILLED 2026-07-29 · THICK-EDGE COMBINE HUNT: KILLED 2026-07-30 AFTER FIVE ADVERSARIAL HUNT ROUNDS
Why this category exists: both specifications were authored blind, hunted adversarially, and killed before a single candidate mechanism was written or a single backtest was run. Strictly cheaper than the FAVBAND / TAKERFOLLOW / crypto15m pattern of dying after tuition has been paid.
Frequency axis: 43 findings, 16 confirmed. Root cause: the target was derived from a tool running a 63-day evaluation budget, while the tool that adjudicates a candidate runs 21. Re-derived honestly at the operative horizon, a thin edge is impossible at every risk level — no daily-loss-feasible frequency clears the bar.
Thick-edge hunt: confirmed critical findings by round ran 26 → 25 → (truncated) → 14 → 15. Repairing all fourteen produced fifteen. The cause of death generalises: a specification that adjudicates a self-produced artefact cannot be repaired by adding assertions to it, because every assertion is satisfiable by construction.
REPLICATIONPASSED, THEN NARROWED
Diversified trend following — replicated, and its limits measured
32 FUTURES MARKETS · 2000–2026 · PRE-REGISTERED BEFORE THE RUN: 2026-07-30 · RESEARCH COST: $0
Result — passed all three pre-registered criteria: net Sharpe +0.47 (t = 2.30 over 26.6 years), positive in four of five non-overlapping five-year sub-periods, and no single sector’s removal flips the sign. A roll-contamination control moved the figure by +0.02, so this is not an artefact of unadjusted contract rolls.
Then narrowed by its own out-of-sample test: the identical rule applied to 33 crypto assets, on a survivorship-bias-free universe that includes coins which went to zero, returned +0.12 with a standard error of 0.30 — t = 0.40, no detectable edge. It therefore stands as a plausibly futures-specific result rather than a durable cross-market phenomenon.
And it is not tradeable on a funded account: maximum drawdown −25.9%, 96% of days spent underwater, longest recovery 1,464 trading days (about 5.8 years). Against a $2,000 trailing drawdown — 4% of a $50,000 account — the most-documented edge in systematic trading breaches the limit routinely and by design.
SHADOWFAILED — GRAVEYARDED
Crypto perpetual funding capture — arbitraged below its own costs
10 MAJORS · 2020–2026 · 68,857 FUNDING OBSERVATIONS · PRE-REGISTERED BEFORE ANY P&L: 2026-07-30 · RESEARCH COST: $0
Result: falsified. The full-period figure looks strong — +9.32% annualised at Sharpe 1.50 — and is carried almost entirely by a 2021 regime that no longer exists. By year: +34.5% (2021) → +10.9% (2024) → +1.5% (2025) → −0.9% (2026). The binding pre-registered condition was placed on the recent window before the run, precisely because that decay was measured first.
Why it dies: roughly 2.4% per year in fees against 1–3% per year of current gross funding. The fee load now exceeds the entire yield. A near-miss on the tail bound (−5.02% against a −5.00% bar) was failed as written rather than rounded through.
SHADOWFAILED — GRAVEYARDED
Signal stacking — two combinations, two distinct reasons
32 FUTURES MARKETS · BOTH PRE-REGISTERED BEFORE ANY SLEEVE P&L: 2026-07-30 · RESEARCH COST: $0
Result: both falsified, and the pair closes the axis from opposite directions. Adding a documented, genuinely negatively-correlated value sleeve (−0.30 and −0.41, as the literature predicts) lowered the combination’s Sharpe, because that sleeve has negative expectancy: negative correlation cannot rescue a losing component. Pairing the two momentum sleeves instead also failed, because at a correlation of +0.58 they are substantially the same bet: positive expectancy cannot rescue an insufficiently decorrelated one.
Method note: the second test was labelled not-blind on its face — it drops a sleeve after seeing that sleeve lose — and its acceptance bar was raised to compensate. Its outcome was predicted in writing beforehand from the measured correlation: +0.35 forecast against +0.36 actual. The mechanism is understood, not merely observed.
SHADOWFAILED — GRAVEYARDED
SPEC_FAVBAND — resting-ask favorites band
MARKET: KALSHI SPORTS · REGISTERED: 2026-07-14 · FAILED: 2026-07-14 · RESEARCH COST: $0 (SHADOW)
Hypothesis: buying resting asks in the 80–90c favorites band nets positive after fees.
Result: n=357 resolved — 77.9% hit at 84.3c average = −6.39c gross / ≈−7.4c net per contract. Both pre-declared kill criteria fired; resolutions spot-checked against the live API (4/4 exact).
Mechanism found: stale resting asks in illiquid series are adverse selection — the research measured fresh taker fills, the implementable signal was stale quotes. At target size this failure would have cost roughly −$780 in six hours. Actual cost: $0 and one day of logging.
SHADOWFAILED — GRAVEYARDED
SPEC_TAKERFOLLOW — fresh-taker-fill following
MARKET: KALSHI SPORTS · REGISTERED: 2026-07-14 · INTERIM GATE MET (WEAKLY): 2026-07-14 AT n=525 · FAILED ON FULLER SAMPLE: 2026-07-15 · RESEARCH COST: $0 (SHADOW)
Hypothesis: following fresh taker fills in the 80–90c band (with a 2.0c freshness guard) captures the edge FAVBAND failed to implement.
Result: the pre-declared gate was met as written at n=525 (+0.39c net, error bars wider than the estimate — logged as such at the time), then failed on the fuller sample. Graveyarded rather than re-tuned: changing the spec resets the evidence to zero.
HISTORICAL — LIVE-DATA AUTOPSYNO-GO
WHALE-FOLLOW — copying large traders
MARKET: KALSHI SPORTS · ANALYZED: 2026-07-12 · SOURCE: THIRD-PARTY BOT'S LIVE PRODUCTION DATABASE (READ-ONLY) · REPLICATED FROM COMMITTED CODE: 2026-07-12
Hypothesis: following large "whale" trades captures edge from informed money.
Result: across n=3,182 resolved sports signals, whales hit 82.0% at ~80.5c average — roughly +1.3 to +1.5c gross, and about breakeven after the measured 1.36c/contract fee. The bot's own 37 settled positions performed exactly at market-implied probability: zero captured edge. Exotics whales were net losers. Findings replicated independently from committed code against a snapshotted dataset.
Verdict: NO-GO as-is. No configuration made it meaningfully profitable.
FUNDED LIVE — SMALLDISABLED
MOMENTUM — third-party bot momentum module
MARKET: KALSHI · OBSERVED: 2026-07-12 · SAMPLE: 3 SETTLED POSITIONS
Result: 0 wins in 3 settled positions, realized −$4.75. Sample far too small for statistical claims — disabled on risk-control grounds rather than kept running to "see".
Method notes
- Pre-registration: specs are frozen in writing before data collection. Pre-registration makes post-hoc changes auditable and forces every changed specification to begin a new evidence record — it reduces the room for curve-fitting; it does not make overfitting impossible.
- Append-only record: entries are never edited after the fact; corrections are new entries. Defects get logged too (the log includes a price-blind-scan defect found and fixed on 2026-07-14, with clocks reset).
- The watchdog: an independent process turns outages into loud evidence. Root cause of early coverage gaps (laptop power management) was found and fixed on 2026-07-15 — those gaps are documented, not smoothed over.
- Fees and costs: measured from real fills (1.36c/contract on the venue analyzed), not assumed away.
Futures and event-contract trading involves substantial risk of loss, including total loss. Nothing on this page is a guarantee of profit or individualized investment advice. Read the full
Trading Risk Disclosure.