What changed and why - strategy verdicts, retirements, new bots, new data.
The wins and the failures get the same font size.
We bought the graveyard: testing 'too big to fail' comebacks against 32,851 dead stocks.
Everyone remembers the distressed stocks that came back. Nobody
remembers the ones that didn't, because dead companies vanish
from ordinary price feeds. So before testing a "buy quality at
maximum pessimism" rule, we bought one month of a
delisting-inclusive dataset ($19.99, cancelled after) and paired
it with the SEC's free as-filed fundamentals: 50,873 US common
stocks, 32,851 of them dead, with survival filters (real gross
profit, cash runway, non-financial) evaluated only on what was
publicly known at each moment.
The pre-registered rule - large caps down 65%+, entered on a
bounce, filtered for survivability - PASSED its frozen gates on
the 2018-2026 test half: 106 events, +15.3% average per event
net of doubled costs, and the filters correctly refused the
trades that became J.C. Penney, Peabody and Chesapeake
bankruptcies. The graveyard did not destroy the result.
Then our independent validation pass failed it anyway, on one
gate: March 2020 alone carries 42.5% of the profit (bar: under
40%), and the earlier 2011-2018 half of history shows no edge at
all. A result one month can carry is a result one missing month
can erase. So: no deployment. If the family goes forward, it
does so as PAPER trading under new pre-registered gates that no
single month can satisfy. Along the way we diagnosed and
published four vendor data pathologies - including a Vietnamese
stock's prices hiding under a dead American ticker - each with
its own archived, polluted run. All backtest, all paper, no
wagers, nothing bought or sold.
We went hunting for the oldest anomaly in finance. It's real - and it lives exactly where you can't trade it.
Closed-end funds sometimes trade well below the value of what they
own, and buying unusually wide discounts is one of the oldest
documented edges in the academic literature. We pre-registered a
zero-knob test before touching any data: entry rules fixed from
the literature, pass bars frozen, one run. Data cost: $0 (a free
NAV feed paired with our existing price feed, 301 funds, some
histories back to 1999).
The verdict is NO-SAMPLE: on funds liquid enough to actually
trade, the registered rule fired 8 times in 4.6 years - nowhere
near the 200-trade floor a verdict requires. We did not accept
that number on faith. A pipeline audit found 9,075 raw signal
days across 258 funds - which collapse to 60 once the price and
liquidity floors apply. The anomaly is real. It just lives almost
entirely in tiny, thinly traded funds where the bid-ask spread
and market impact would eat the very discount you came for.
That is the honest shape of many famous edges: visible in data,
concentrated precisely where execution costs are worst. All
backtest, no wagers, total spend $0, goalposts untouched.
We tested shorting the market's hottest stocks. The gates said no - and the timing part actually worked.
Our best paper book (double7 plus two Qullamaggie families) earns
about $355 a month per $25k in backtests, but a new measurement
this week showed WHERE its weakness lives: the book's idle capital
concentrates in bear and drawdown months. October 2022 ran 0.3%
utilization with a 30-day stretch of zero positions. All three
families are long-side systems, so they go quiet together exactly
when markets fall.
That pointed at a specific fix: a short-side family that gets
BUSIER in falling markets. We pre-registered a study before
writing any code - frozen pass bars ($395/month combined, dead
months must fill in, results must survive doubled borrow costs),
parameters taken verbatim from Laurens Bensdorp's published S2
and S6 short systems so there was nothing to tune, and one run
per variant, period.
The verdict: FAIL, all three variants. And the failure has a
shape worth reporting. The timing hypothesis was RIGHT - the
short systems fired in 9 to 10 of the 11 registered dead months,
with essentially zero overlap with the incumbent families' trades
(0.0 to 0.4% shared entry dates). They showed up exactly when we
wanted them to. They just lost money when they got there: about
-$0.19 to -$0.38 per $100 slot per trade after realistic fees,
slippage, and borrow costs, dragging the combined book DOWN to
$291-$345 a month instead of up past $395.
Being busy at the right time is not the same as having an edge.
The book's 1995-2019 short-side results did not survive 2018-2026
on our universe with our costs - shorting a six-day surge in the
meme era is a different proposition than it was in 1999. Total
spent on this study: $0, all backtest, no wager, no deployment.
The engine still cannot short a single share, and after this
result it will stay that way.
Bug report: the dashboard said our options bot was up $205. The options page said $80.
Pfizer rallied through the $27 strike of the covered call our
options bot is running, and that rally exposed a bookkeeping gap
on this site: the main dashboard showed the position up $205.50
unrealized, while the covered-calls page showed $80.00. Both
pages, same bot, same moment, $125.50 apart. Regular readers know
the house rule by now: any two numbers that disagree expose a
bug to anyone who scrolls, and $125.50 happens to be exactly what
the difference should be - the short call's in-the-money value.
Here's what happened. A covered call is two legs: 100 shares you
own, and a call option you sold against them. The fleet dashboard
reads the same snapshots table as every other bot, and that table
only ever carried the share leg - so when PFE ran from $26.20 to
$28.26, the dashboard cheerfully counted all $205 of stock gain.
But roughly $126 of that gain belongs to whoever bought our call:
above $27, every extra dollar of stock is a dollar we owe on the
option. The options page always knew this - it reads the options
bot's own forward log, which marks the short call as a liability
every session. The dashboard just wasn't listening to it.
Now it is. Every dashboard view - the sector tiles, the fleet
table, the open positions list, the equity chart - pulls options
bots' equity from the same two-leg source of truth the options
page uses, so the numbers agree by construction rather than by
luck. The honest figure is +$80: the stock's gain up to the
strike, plus the premium we were paid, minus what the call is
worth against us. If PFE stays above $27 through September 4,
$80 is also almost exactly what settlement will deliver - the
upside past the strike was sold on day one. That's not a flaw in
the trade; capped upside in exchange for premium is the entire
covered-call bargain. The flaw was a display that showed the
upside we'd already sold as if we still owned it.
Worth saying plainly: the mistake made us look BETTER than
reality, which is precisely the kind of error a track record
cannot afford. Paper trading, as always - and the September 4
settlement will be the first real test of the options plumbing
this bot exists to prove.
Retired: the first bot to reach our evidence bar - and the first our gates rejected
btc-breakout-001, one of the fleet's original five bots, is retired
today. It earned a distinction on the way out: it was the first bot
in the fleet to complete 50 round trips - the sample size our own
rules say is enough to judge - and the first one our new graduation
gates looked at and rejected. 53 trades over 39 days: net −$61.23
on a $1,000 paper account (−$25 in trading losses, ~$36 in fees on
$36,000 of churned notional), with a 28% win rate.
The diagnosis isn't mysterious. It ran Turtle-style Donchian
breakouts, long-only, on 15-minute Bitcoin bars. That strategy
accepts many small stopped-out losses because the occasional big
runner is supposed to pay for all of them. In a choppy, drifting
market the runners never arrive - its best trade made +$6.28 while
its worst lost $7.17, and the 1% stop dutifully converted every
fake breakout into a small paid exit. Meanwhile its mirror image -
the SHORT Donchian family on perpetual futures - remains green in
the same market. Same arithmetic, opposite regime.
Every number above stays in the database, and the verdict is
recorded in the bot's config file where the next reader will find
it. Retired bots keep their final balance on the dashboard - the
row freezes rather than disappears, because deleting your losers
is how track records get faked. Paper trading, as always - the
only thing spent here was electricity, and the only thing earned
was the receipt.
Arbitrage verdict: we measured it for 16 days so we didn't have to build it
Two weeks ago we pre-registered a question: does any arbitrage
family produce opportunities that clear our real costs, at a speed
a $6 REST-polling server can actually capture? Three studies, all
thresholds committed in writing before the first quote was logged
(docs/arbitrage/SPRINT-1.md, commit 600be4a). The window closed
today: 4,346 five-minute scans over 16 days across Crypto.com,
Kraken, and Coinbase. One bucket is missing from August 7 - we
briefly wiped our own crontab and restored it within minutes; the
gap is disclosed rather than papered over.
Study A, funding-rate harvest, was the primary hope: go long spot
and short the perpetual, collect the funding payments, stay
market-neutral. Verdict: FAIL, on all three pre-registered checks
that matter. Funding rates did spike (ADA annualized at +54% at one
point), but a round trip costs ~0.40-0.46% in fees and spreads, and
the best episode in the test window captured only +0.226% of
funding over 86 hours - still net negative. All eight test episodes
lost money. The math is simple and brutal: at retail taker fees,
the funding tide never rose high enough to clear the toll.
Study B, triangular loops on one venue, technically PASSED its
floor: 171 persistent net-positive loops across 14 days. Before
anyone gets excited: every single one was on thin USDT crosses like
JASMY and CSPR, past the venue's liquidity cliff, measured at
top-of-book with zero depth. Our own protocol says exactly what
that's worth: a paper pass at top-of-book earns a depth-aware
follow-up study, never a deployment. We suspect the depth check
kills it. We'll measure that too.
Study C, cross-venue price gaps: zero qualifying events in 16 days.
And when we verified fee tiers at verdict time (as the protocol
required), Kraken's real base taker turned out to be 0.40%, not the
0.26% we'd assumed - which would have made the threshold even
harder. A FAIL that survives its own assumptions being wrong
against it is a comfortable FAIL.
Consequence, decided before the data existed: no arbitrage engine
gets built. Limit orders, multi-leg atomic execution, second
exchange accounts - all of it stays unbuilt, because a two-week
observation study costing almost nothing just told us it would
have been wasted. The logger keeps running (it's nearly free) and
the study re-opens automatically if the funding regime changes.
This is the third time this summer a pre-registered feasibility
study has killed a build before it started. That's not a failure
of the process. That is the process.
New free tools: the number prop firms don't advertise, and the one your win rate is hiding from you
There's a new "Prop Firm Tools" tab on the site. Two free
calculators, no signup, no email. Here's what they do and why we
built them.
TOOL ONE: the evaluation simulator. If you've looked into funded
trading accounts - Topstep, Apex, and the rest - you pay a fee,
trade to a profit target without hitting a drawdown limit, and get
a funded account. Sounds like a skills test.
Here's the thing everyone misses: whether you pass depends
enormously on the ORDER your wins and losses happen to arrive in.
Same strategy, same trades - losing streak lands in week one, when
your cushion is thin, and you're out. Same streak in month three,
you're fine. You only ever get one shuffle of the deck. So instead
of replaying your history once, the simulator shuffles it ten
thousand times and tells you what fraction of orderings pass.
And it computes the number the industry has no reason to show you:
a strategy with NO edge at all - literally a coin flip after costs
- passes a typical evaluation roughly one time in five. Enough
people buy enough attempts that a steady stream of funded accounts
goes to traders with no edge whatsoever. Passing, by itself, is
much weaker evidence than it feels like.
The comparison tab runs your same numbers against ten real firms'
actual rules. We read each firm's own rulebook and help center -
not review sites, not affiliate pages - and recorded every number
with a source link and a date. The fine print mattered more than
we expected: whether a "trailing drawdown" tracks your open
profits in real time or only your end-of-day balance moves the
pass rate by double digits on the same strategy, and firms bury
that distinction. Anything we couldn't confirm from a primary
source, the page refuses to display - and rows automatically get
flagged for re-verification once they're 90 days old, because
these terms change without announcement.
TOOL TWO: the breakeven win rate calculator. It corrects the most
expensive misconception in trading: that your win rate is your
edge. It isn't. A 60% win rate sounds great and LOSES money if
your average win is half your average loss. Enter your average
win, average loss, and costs, and it tells you the win rate you
need just to break even - plus the part almost nobody does: given
how many trades your win rate is based on, whether you can even
statistically tell it apart from a coin flip yet. Over 100 trades,
a measured 55% could easily be a true 45%. The tool says so, out
loud, and tells you how many trades it would take to know.
What these tools are NOT: advice. They don't recommend a firm,
don't predict your results, and can't tell you whether your edge
is real - every number assumes the stats you type in are honest
and will hold up out of sample. If they came from a backtest you
kept tweaking until it looked good, they won't. And to be clear,
none of this is our bots' trading - our fleet results are separate,
all PAPER money, and these calculators are just math on numbers
you provide.
Our bots were hoarding cash like squirrels. We built them a shared piggy bank - then caught ourselves cheating three times
Here's a problem nobody warns you about when you build a bot fleet.
Every bot in our stock fleet keeps its own money. Bot #1 has its
pile, bot #2 has its pile, all forty of them. Which means you have
to fund the worst case: the day every single bot wants to buy at
once. So you set aside forty piles of cash.
That day basically never comes. We measured it. Our bots are all
holding positions at the same time about 4-25% of the time. The
rest of the time, most of that money is sitting there doing
absolutely nothing, like a gym membership in February.
So we built a shared piggy bank. Instead of forty separate piles,
the bots draw from one pool of ten "slots." Want to buy? Take a
slot. Sold? Put it back. All ten taken? You sit this one out.
The trade is real: sometimes a bot spots something and gets told
"sorry, we're full." You give up a few trades. In exchange you
need far less cash standing around to run the same strategies.
Nothing about the strategies changed. We just stopped funding a
traffic jam that doesn't happen.
Now the fun part: getting here required being wrong three times,
and we'd rather tell you about those than the tidy version.
MISTAKE ONE. Claude first concluded that adding more strategies
makes returns WORSE, complete with a convincing table. That turned
out to be a bug in how it was splitting the money between them,
not a fact about the strategies. Under shared funding the answer
flips: adding strategies helps. A confident table is not evidence.
MISTAKE TWO. The first version of the shared-piggy-bank test said
we'd earn about 66% a year. Wonderful! Also nonsense. When several
bots wanted a slot on the same day, the code was quietly handing
it to whichever trade would finish soonest - which is information
nobody has until afterward. It was letting the bots peek at the
answer key. Fixed, the number dropped by a third. Suspicious
good news is still suspicious.
MISTAKE THREE, the scary one. We ran the official test. It said
PASS. It was also silently missing 10 of our 24 Double 7s bots -
because ten stocks are traded by two different bots each, and our
code had been quietly overwriting one of each pair. The test
passed on a fleet that wasn't the fleet. The only reason we caught
it: the report printed "14 bots" next to a rule that said 24.
That last one is the whole lesson. A test that fails is annoying.
A test that PASSES on quietly wrong data will happily send you
live with real money, smiling the entire time. Print your inputs.
Check them against what you promised to test.
What this is NOT: proof that our stock strategies make money.
Those are still on trial, with their own verdicts pending, and
they might well flunk. A shared piggy bank makes whatever edge you
have go further. If there's no edge, congratulations - you've
built a more efficient way to lose money.
The forty new bots now running are a plumbing test, not a profit
test. We're checking that slots actually get handed out and handed
back, that a server restart doesn't lose track of who's holding
what, and that ten slots never somehow becomes eleven. We wrote
down what would count as passing before switching them on. All
PAPER money, as always. The strategy tests already running were
not touched - new bots only, so we don't contaminate an experiment
that's halfway done.
The mean-reversion math study came back negative - and that's worth publishing
Three weeks ago we pre-registered a hypothesis: our Connors-style
mean-reversion bots should do better on stocks whose prices
statistically snap back fast (short OU half-life, strongly
negative Dickey-Fuller stat - the textbook mean-reversion
measurements). The judging rules were committed before the study
ran, including a trap for exactly what ended up happening.
On RSI-2, the primary system, both measurements looked adoptable:
the predicted trades beat the average in both the train and test
halves. If we'd stopped there, we'd have deployed new gated bots
this week. But the protocol required the same pattern to show up
in Double 7s - a sibling strategy trading the same dip on the
same 24 stocks - and there it ran BACKWARDS: the "best" trades by
the math underperformed, and the training data actually favored
the opposite bucket. One strategy's promising pattern failing to
transfer to its own sibling on identical data is what
overfitting looks like from the inside.
So: no new bots, hypothesis rejected, receipts in the repo. The
48 live Connors bots were never touched - they keep running
exactly as deployed. PAPER trading throughout.
First formal forward verdict: gated Bollinger is INCONCLUSIVE - positive, but not by enough
Four weeks ago we committed the judging rules for our first
formal forward test to git - before the outcome was knowable.
The condition-gated Bollinger fleet (17 crypto spot bots, PAPER)
promised +$0.364 per trade at a 70% win rate based on its
backtest. The pre-registered bar for a forward PASS: positive
average, win rate at least 55%, and at least half the backtest's
per-trade average.
The verdict at trigger: 62 completed round trips, +$3.11 total,
+$0.05 per trade, 66.1% win rate. That's positive, and the win
rate essentially delivered on the promise - at the 50-trip mark
it was exactly 70.0%. But +$0.05 is a long way from the +$0.182
the protocol requires, so by the rules we wrote in advance this
is INCONCLUSIVE, not a pass: the test extends four more weeks,
same criteria, re-verdict 2026-09-07.
What the number hides: the gates are finding winners at the
promised rate, but two names (NEAR and SHIB) produced losses big
enough to eat most of what the other thirteen bots earned. Our
protocol explicitly forbids excluding them after the fact -
"it passes if you ignore the losers" is how fake edges get
published. Two honest confessions, on the record: the 50-trip
trigger technically fired eight days before we noticed - the
verdict comes out identical at both cuts, and the miss itself is
documented in the protocol file - and the protocol text said
"13 bots" when the fleet was 17: a miscount, all 17 ran from
day one.
Everything here is paper trading. Nothing about this result
moves real money.
The full verdict as a video, receipts on screen:
https://youtu.be/pIUTxvX0kYo
The first options bot is live (PAPER) - and it isn't allowed to prove the strategy works
Two days after the covered-calls build started, the studies are
done and the first bot is running. The short version of the
studies: covered calls and the wheel both passed their
pre-registered tests against simply holding the shares - 4 out
of 4 stocks each - but the honest reading is narrower than the
scoreboard. Not one stock won on BOTH return and risk. In
rallies the strategies kept you safer but gave up most of the
upside; in declines they lost less but still lost. And the
wheel's famous cycle - sell puts, get the shares, sell calls,
get called away, repeat - completed exactly three times across
four stocks and thirty months of history. The honest sales
pitch after measuring it: option premium pays you to wait out a
drawdown. It does not prevent one.
That was enough to earn one bot, and only one: cc-pfe-01, which
now owns 100 shares of Pfizer and has sold one call option
against them, on paper, checked once a day after the close. PFE
got the slot on pre-stated evidence - it was the only stock that
passed both ways in either study.
Here's the part we care most about: the rules for judging this
bot were committed to git before it made a single decision, and
they explicitly forbid the one conclusion everyone wants. Ninety
days of one stock cannot prove covered calls "work," so the
verdict doesn't ask that. It asks whether the plumbing holds up
(options expire and settle correctly, nothing breaks on a
restart) and whether the real-world costs match what the studies
assumed - because if live spreads are meaningfully worse than
the studies charged, the studies get flagged as too optimistic
and expansion stops. Return versus buy-and-hold gets reported,
but it is not allowed to gate anything. Day one delivered its
first two data points, both mildly good news: the live spread on
the call we sold was tighter than the studies assumed, and the
bot selected its contract using live option greeks - something
the backtests never had access to, which is exactly the gap this
forward sample exists to measure.
As always: this is PAPER trading and forward data collection on
Alpaca's free indicative feed - calculated quotes, not the paid
exchange feed. Real math, no financial advice.
The fifth sector begins: covered calls - and the first thing we built is the test that could kill it
The covered-calls page on this site has said "coming soon, and
here's the honest reason it isn't here yet" for a while. Today
the build actually started. The plan: bots that own 100 shares
of a cheap, liquid stock and sell call options against them,
collecting the premium - plus the "wheel," which sells puts to
get into the shares in the first place. Fancier options
strategies are explicitly parked until these two prove
themselves.
True to form, the first code written wasn't a trading bot - it
was a measuring stick. Options quotes have a catch: the gap
between the buying and selling price (the spread) can quietly
eat the entire premium you collected. So before building the
engine work - the biggest on our roadmap - we froze a five-stock
universe using criteria committed to git first (NOK, PFE, T,
CMCSA, BITO), and started a three-day study measuring whether
real spreads leave anything worth collecting. The pass/fail bar
was also written down first: if the cost of trading eats more
than a quarter of the premium on every single name, the sector
stops before the engine gets built. Day one looks promising -
four of five names comfortably under the bar - but day one is
not a verdict. The verdict runs Thursday and gets published
either way.
One honesty note up front: our options data is Alpaca's free
"indicative" feed - calculated quotes, not the paid real-time
exchange feed. Every number from this sector carries that label
until an edge looks real enough to justify paying for the real
thing. Paper trading, real math, no financial advice - as
always.
New video: is crypto arbitrage still possible in 2026? We measured it.
Arbitrage is the internet's favorite "free money" pitch, so
instead of buying anyone's course we pointed a read-only logger
at it: every five minutes it scans all ~300 triangular loops on
our exchange, logs cross-exchange price gaps against two other
venues, and records the hourly funding rate on the liquid
perpetuals. The pass/fail criteria were committed to git before
the first quote was logged.
Four days in: 890 scans, and the median best triangular loop on
the whole exchange loses 0.26% after fees - the market is
efficient to almost exactly the width of the fee. A handful of
scans did show "profitable" loops, almost all through tiny,
illiquid pairs where the advertised price is money nobody can
actually collect at size. The one family still genuinely open is
funding arbitrage - hold the coin, short its perpetual, collect
the hourly funding payment. It's the only kind of arbitrage that
rewards patience instead of speed, which is the only race a $6
server can enter, and Bitcoin's funding swung from roughly −10%
to +15% annualized during the study's first days.
The new video walks through it with the real numbers:
https://www.youtube.com/watch?v=DgO9pC1UEG4 - the verdict lands
August 15 under the pre-registered criteria and gets published
either way. As always: paper trading, real prices, real fees, no
financial advice.
We tested day-trading bots. The stock version failed its own test - so we're not building it.
This week the fleet learned to roam: instead of trading a fixed
watchlist forever, our newest bots scan the entire market - all
13,000+ US stocks - and compete for a fixed number of position
slots. Two swing-trading roamers are live on paper now, each
judged against the hand-picked bots running the identical
strategy, so the only thing being tested is whether scanning
beats picking.
Then we asked the obvious next question: day trading. For
stocks, we took a famous day-trade setup - opening-range
breakouts on big gap-ups, rules taken from the trader's own
published write-ups - built a dataset of 3,377 real gap days
across 31 months, and wrote the pass/fail criteria down BEFORE
running anything. The result: the setup as actually specified is
rare (about 38 qualifying trades in 31 months), and the apparent
profit on those trades vanished entirely when we doubled the
modeled slippage. An edge that lives inside the fill-model
assumption isn't an edge. Verdict: the stock day-trading bot is
not being built. One afternoon of honest testing beats months of
a bot trading noise - and, as always, a pre-registered test gets
published whichever way it lands.
One day-trading bot did earn deployment: a crypto roamer running
the exact same gated Bollinger strategy as our 13 fixed crypto
bots - verbatim, enforced by an automated test - rotating across
the 20 most liquid pairs instead. Paper money, labeled data
collection, judged by criteria committed before its first trade.
And if the underlying strategy fails its own long-scheduled
verdict on August 10, the roamer parks with it. No rescues, in
either direction.
24 new bots: our best strategy just got a sibling - and they'll compete head-to-head, live
The fleet just grew from 113 to 137. The newcomers run Connors
Double 7s, from the same 2008 book as our strongest strategy -
the Connors RSI-2 fleet that buys 2-3 day panics inside uptrends.
Double 7s is its simpler sibling: if a stock closes above its
200-day average and today's close is the lowest of the last 7
days, buy. Exit when the close is the highest of the last 7 days.
That's the entire system. No stop-loss, by the book's own
research - the dip IS the entry logic - with our position caps
bounding the risk instead.
Why it earned a slot: we backtested it against RSI-2 on the exact
same 24 stocks, the same 2000 daily candles, the same costs. It
took 922 trades to RSI-2's 431 - a 7-day-low close simply happens
about twice as often as a 2-period RSI under 5 - kept a
comparable 67% win rate, earned MORE per trade, and finished with
2.6 times the total profit, green on 22 of 24 names. That's the
largest backtest sample of anything we've ever deployed, passing
our 50-trade evidence gate eighteen times over.
Now the honest frame. A backtest is an audition, not a verdict -
that's why the new bots trade paper money on the very same stocks
as their RSI-2 siblings, at the same sizing, side by side. Two
systems from the same author, hunting the same regime, judged by
the same forward tape. Either the backtest edge shows up in live
market conditions or it doesn't, and the dashboard will say which
in public, either way.
Watch them (bot IDs starting with d7-) at trueai.trading/dashboard
All results are PAPER - simulated money, real market prices, real
fees. Not financial advice.
The exit experiments are complete: exits don't stack, and our best result just survived its second test
Last test in the exit series. This morning we showed that combining
our two best exit rules on crypto changed nothing - the fast exit
fired first on every trade. The remaining question was the stock
side, where the slower Darvas box-trail is the champion: does
adding the fast Heikin Ashi flip on top help there?
It made things worse. Five exit variants, identical data, 16
stocks, 2000 daily candles each: the incumbent 10 EMA trail made
+$533. The HA-flip exit alone made +$409 (cuts long stock trends
early, as we already knew). The combination made +$388 - below
BOTH of its components' better halves. Unlike crypto, the Darvas
floor did occasionally fire first on stocks, and when it did, it
subtracted. So the cross-market conclusion is airtight now: a
combined exit is always dominated by its faster component. You
never get the best of both rules - you get the fast one, plus
extra ways to lose. Exits don't stack.
The genuinely good news from the same run: the Darvas box-trail
won on stocks AGAIN, on a second, longer test window - +$684
versus the incumbent's +$533, a 28% improvement, earned on fewer
trades. Its first win was +90% on a shorter window; the honest
reading is that the direction is now confirmed twice while the
size of the edge depends on the window. Two independent windows
agreeing on direction is worth more than one window's big number.
So the exit playbook is settled and waiting: patient Darvas trail
for stock breakouts, fast HA-flip for crypto, no combinations
anywhere, and the regular stop-loss underneath everything. All of
it stays parked until the mid-August strategy verdicts, then gets
re-verified out-of-sample before touching a single live bot.
All results are PAPER - simulated money, real market prices, real
fees. Not financial advice.
Follow-up: we tried combining our two best exits. The answer was a perfect, boring zero.
Quick follow-up to today's Heikin Ashi test. We had two exit rules
that each won somewhere: the fast HA color-flip (beat our deployed
exit on crypto) and the looser Darvas box-trail (nearly doubled it
on stocks). Natural question: on crypto, should the fast exit also
carry the Darvas safety floor underneath - exit on whichever
triggers first?
So we ran it. The combined exit produced results IDENTICAL to the
plain HA-flip - to the penny. Same 50 trades, same wins, same
+$83.98. Not similar. Identical. The reason is simple once you see
it: a red Heikin Ashi candle shows up long before price falls 5%
through a box bottom or gives back 20% from a peak, so the fast
exit fired first on every single trade and the floor never got
touched. It was dead code wearing a seatbelt.
Verdict: skip the floor. The crypto exit candidate stays the plain
HA-flip - fewer moving parts, same result - and the engine's
regular stop-loss remains underneath everything as the actual
catastrophe brake. A zero-effect result is still a result: it's
one less parameter to fit, one less thing to break, and one more
combination we'll never have to wonder about.
Nothing deploys until the mid-August verdicts, same as always.
All results are PAPER - simulated money, real market prices, real
fees. Not financial advice.
We tested Heikin Ashi honestly. It beat our exit on crypto and lost on stocks - and that's the interesting part.
Heikin Ashi charts are those smooth, beautiful candles that make
every trend look obvious in hindsight. A viewer-suggested video
pitched a strategy around them, so we did what we always do: mined
the rules, researched the technique, and put the testable part
through an honest A/B.
First, the trap - because most Heikin Ashi results you see online
are fake, and not subtly. HA candles are averages: the prices they
display DO NOT EXIST in the market. Backtests that fill orders at
HA prices systematically overstate results, badly enough that
TradingView built a dedicated setting to prevent it. Our tests
compute HA colors from real candles and fill every order at real
prices. If you remember one thing about Heikin Ashi, make it that.
The video's entry ("doji at a key support level") is drawn-by-eye
discretion, so it can't be coded faithfully and we didn't pretend
to. But its exit rule is fully mechanical: ride the trend until
the first opposite-color HA candle. We tested that against our
deployed 10 EMA trail exit on both Qullamaggie breakout fleets,
same entries, same candles, real costs.
Crypto (50 trades): the HA-flip exit made +$83.98 vs the
incumbent's +$49.37 - the first exit variant to beat our deployed
one on crypto. Stocks (150+ trades): it LOST, +$409 vs +$533,
because it cuts long stock trends early. And that mirrors last
week's Darvas result perfectly, where a looser trail nearly
doubled stock profits but hurt crypto. Three exit experiments now
tell one story: stock trends run long and want patient exits;
crypto daily trends are short and spiky and want fast ones. One
exit does not fit all markets.
Nothing changes on the live fleet today. Both Qullamaggie fleets
are inside their forward-verdict window (due mid-August), and
changing exits mid-test would corrupt the verdict - the same
goalpost rule as always. The market-split exit idea goes in the
queue for after the verdicts, where it gets re-verified
out-of-sample before touching a single bot. Caveats on the record:
the crypto sample is exactly at our 50-trade evidence gate, one
window, comparative only.
All results are PAPER - simulated money, real market prices, real
fees. Not financial advice.
Next sector on the roadmap: prediction markets - and why the obvious bot is the hardest one to trust
There's a new tab on the site: prediction markets. Kalshi-style
event contracts - yes/no markets on real-world events where the
price IS the probability. A contract at 30 cents means the market
says 30%. We want to know whether bots can find edges there, and
the sector is now officially on the roadmap, parked behind the two
strategy verdicts due in August.
Here's the part worth reading. The obvious bot - an AI that reads
the news and estimates probabilities better than the market - is
easy to build and nearly impossible to validate honestly. Every
backtest of it is a lie by construction: any AI model you'd use
today was trained on data that includes how past events turned
out. Asking it "what were the odds of X?" about anything in its
training window isn't forecasting, it's remembering. So if we ever
run that bot, it starts from zero in live forward testing, labeled
as the experiment it is, for however many months that takes. No
shortcuts exist, so we won't pretend to have found one.
The bots that come FIRST are the boring ones: documented
structural patterns like the longshot bias (unlikely events are
systematically overpriced) that don't require understanding the
news at all, and can be partially backtested against real fee
models. Fees matter enormously in these markets - they're highest
exactly where the outcome is most uncertain.
When the sector ships, it starts on Kalshi's regulated demo
environment: paper first, always, same as every sector before it.
The full honest assessment - including what we can't test and
why - is on the new page: https://trueai.trading/prediction-markets
As always: all results PAPER, nothing here is financial advice.
Gated-Bollinger interim report: 19 trades, 68% wins, still down $2.68
Two weeks ago we committed, in public and in advance, to a verdict
protocol for our gated Bollinger fleet: pass/fail criteria defined
before the outcome, an interim report at the two-week mark, and a
final verdict at 50 trades or August 10, whichever comes first.
This is the interim report. Numbers only - the protocol forbids
verdict language until the trigger, so you'll get none here.
Since the forward test began on July 13, the 13 gated bots have
completed 19 round trips: 13 winners, a 68.4% win rate, for a total
of -$2.68 - an average of -$0.14 per trade, net of fees. For
reference, the backtest that earned this fleet its slot promised
+$0.36 per trade on a 70% win rate. When we committed the protocol
at 13 trades, the fleet stood at 54% wins and -$0.21 per trade;
both numbers have since moved toward the backtest, and the win rate
is now within touching distance of the promise. The average,
however, is still negative.
Where the money actually went: one symbol, NEAR, accounts for -$6.13
across five trades, including three losses of about -$3.30 each -
several times larger than any other bot's loss. The other six bots
that traded are +$3.45 combined, in the small-frequent-win shape the
backtest predicted. We note this as a fact about composition, not as
an excuse: the protocol explicitly bans "it passes if you exclude
NEAR" reasoning at verdict time, and we intend to honor that. Six of
the thirteen bots haven't traded at all yet - the condition gates
are doing exactly what gates do, which is keeping bots out of most
markets most of the time.
At the current pace the fleet clears the 25-trade minimum well
before August 10, so the verdict should arrive on schedule. Whatever
it says, it gets published. All results are PAPER - simulated money,
real market prices, real fees.
Position sizes go from $100 to $500 on 63 bots - here's exactly why, and why not all 110
From today, 63 of our bots trade with a $500 position cap instead of
$100. Nothing about their strategies, signals, or leverage changed -
every bot still trades unleveraged cash positions on $1,000 of paper
equity. So why bother?
Visibility. With a $100 cap, a bot that catches a clean 2% winner
moves its account by 0.2% - a rounding error that makes real edges
and real failures look identical on the dashboard. At $500, that
same trade moves equity 1%. The experiment doesn't change; the
signal-to-noise on reading it does. We scaled the daily-loss halt by
the same factor (from $20 to $100), so every risk limit still sits at
the same percentage of position size the backtests were run with.
Now the honest part: 47 bots did NOT get the change. The 13 gated
Bollinger bots are two days from their interim report, under a
verdict protocol we published in advance with pass/fail thresholds
defined in dollars per trade. Changing their position size mid-sample
would make those numbers meaningless - a quiet way of moving the
goalposts we promised not to move. Same for the Qullamaggie fleet,
whose verdict lands mid-August. They finish their tests at the size
they started. Once each verdict is published, whichever way it goes,
they move to $500 like everyone else.
One thing to keep in mind reading the dashboard from here: dollar
swings on the 63 resized bots are 5x bigger as of today, so
before-and-after comparisons of dollar P&L need that context.
Percentage returns remain comparable throughout. As always: all of
this is PAPER - simulated money, real market prices, real fees.
Autopsy of our two worst bots: it was never really about the fees
btc-kama-001 and btc-sma-001 - two of our original bots, benched since
July 15 - got a final review today, because we wondered: they traded
constantly on 1-minute bars, and every round trip pays about 0.3% in
fees and slippage. Were the fees the killer? Would slower bars, or
bigger positions, have saved them?
So we backtested both strategies at every timeframe from 1 minute to
4 hours, with real costs, over up to a year of data. The answer is
no - and the reason is more interesting than the question. At every
timeframe, the signals lost money BEFORE fees. Win rates ran 12-34%.
These are long-only trend-followers, and in a year where Bitcoin
chopped its way down about 44%, they bought every fake rally and got
stopped into every resumption of the slide. Bigger positions would
have lost more, not less. The fee story was a comforting misdiagnosis;
the signal was the problem. Verdict recorded in both configs, bots
archived for good.
One genuinely useful thing fell out of the autopsy: flip KAMA around.
The SHORT-side version on 4-hour perp bars was the single profitable
cell in the whole matrix - +1.29% over a year in which buy-and-hold
lost 44.9%. That is a tiny number and one backtest window, so it
proves nothing yet. But it matches the market we're actually in, so
it joins the fleet as perp-kamashort-btc-s01 - labeled data
collection, not edge, until it earns a real forward verdict. (The SMA
short version came out at +0.21% - indistinguishable from zero, so it
gets no second life. Not every autopsy finds an organ donor.)
Bug report: the dashboard said 3 open positions. The table said 8.
Spotted on our own dashboard: the open-positions counter said 3 while
the positions table right below it listed 8. Both can't be right - and
the table was.
The cause is a classic: when the fleet learned to short-sell, short
positions became negative numbers, and one counter still asked "is the
position greater than zero?" - counting only the longs. We'd already
fixed two spots with this bug when shorts went live; this was the
third and last, and a sweep of the code confirms the class is now
extinct. The counter and the table agree: 8 open positions, 5 of them
short.
Why publish a one-line bug fix? Because this is what an open book is
for. Every number on this site renders from the same database, so any
two numbers that disagree expose a bug to anyone who scrolls - no
trust required. When it happens, we fix it and say so. Transparency
isn't just for the wins; it's for the counter that couldn't count.
Fleet leaderboard: the 3 best and 3 worst bots right now
First edition of an occasional honest leaderboard - with the caveat up
front that these are tiny numbers by design: every bot runs a $1,000
paper account with positions capped at $100, so the dollars are small
and the point is the process.
The top three: boll-ada-g01 (+$1.57), boll-uni-g01 (+$1.23), and
btc-bollinger-001 (+$1.08). The first two are gated Bollinger bots -
our one backtest-validated edge - quietly leading the fleet days
before their formal forward-test verdict. Encouraging, not conclusive.
The bottom three: rsi2-cvna-s01 (−$9.08), boll-near-g01 (−$6.93), and
btc-breakout-001 (−$4.06). The worst is our brand-new dip-buyer's very
first trade - it bought a Carvana panic and the panic kept going. Its
strategy explicitly holds through this (no tight stop, by published
design), so this position is the strategy's first real exam, still in
progress.
Now the most instructive line in the data: the fleet's #2 best bot and
#2 worst bot run the IDENTICAL strategy - gated Bollinger - on
different coins. UNI up, NEAR down, same rules. That's why no single
bot's result means anything here, why we deploy strategies as fleets,
and why verdicts only come from many trades over many symbols. The
first such verdict is due within days: the gated Bollinger forward
test. Leaderboard, meet judgment day.
110 bots broke our database (a little). Fixed.
A growth milestone, told honestly: the fleet got big enough to strain
its own infrastructure. With 110 bots each logging every decision,
trade, and account snapshot, the database was being written to almost
constantly - and anything trying to READ it (like the live dashboard,
or the fleet report that caught this) kept getting locked out while
the bots held the pen. Some queries had also quietly grown slow: what
was instant at 10 bots was crawling at 40,000 rows.
The fix is textbook database engineering: write-ahead logging, which
lets readers and writers work at the same time instead of taking
turns, plus proper indexes on the columns the dashboard queries. The
heaviest dashboard query went from struggling-to-connect to 34
milliseconds. The fleet restarted with every open position intact -
including our new dip-buyer's first trade, currently underwater and
holding exactly as its strategy prescribes.
Why post about plumbing? Because this is what running 110 bots
actually involves, and most trading content skips it. The strategies
get the headlines; the infrastructure keeps the lights on. Both run
in the open here.
Our first video is up: Can AI actually trade?
The YouTube channel has its first full-length video: "Can AI Actually
Trade? I Built a Fleet of Bots to Find Out (Honestly)." It's the story
of this whole project - the fleet, the honest-testing pipeline, the
strategies that survived the backtests and the ones that didn't - told
from the beginning.
If you've been following along here, the video is the big picture in
one sitting; if you're new, it's the best place to start. It's on the
videos page right here on the site, or on the channel. More videos
will follow as the fleet's verdicts come in - the gated-Bollinger
forward test wraps up around late July, and that result is getting a
video whichever way it goes. Wins and losses alike, as always.
We finished mining trading YouTube. Here's the scoreboard.
Since this project started, our strategy pipeline has run on one input:
find a documented trading method, extract its rules from the source,
implement them in code, and let backtests decide. This week we finished
working through the entire YouTube source queue - every channel either
mined or ruled out, with the reasoning committed to the repo.
The scoreboard is stark. Sources with published, verifiable rules -
Qullamaggie's breakout method, Mark Minervini's VCP, Jesse Livermore's
exits, Larry Connors' RSI-2 - produced every deployed strategy in the
fleet. Sources built on drawing lines by eye produced nothing we could
even test: when the lines get redrawn daily "until they fit," the
signal that would have existed at decision time is undefined, and a
backtest of it would be fiction. And a pattern worth knowing: the
channels that couldn't state their rules precisely were usually the
ones selling a course or funneling to a prop-firm evaluation. The
method is the product when the method works; the course is the product
when it doesn't.
So the hunting ground moves to where the evidence lives: published
books, academic papers, and quant research - which is exactly where
our newest and best-tested strategy came from. Every verdict, adopted
or rejected, is in the open in our strategy-sources docs. The numbers
decide here. They always have.
24 new bots buy fear: Connors RSI-2 joins the fleet
The Stocks sector just got its first mean-reversion strategy - and the
strongest backtest this project has ever produced. Connors RSI-2, from
Larry Connors' published 2008 rules, buys two-to-three-day panic dips
inside long-term uptrends and sells the snap-back. It's the exact
opposite regime to our breakout bots: they buy strength, these buy fear.
The numbers, with our usual punitive fee assumptions and the book's
parameters untouched: 204 trades across three universes and four years
of data, 75% winners. Index ETFs reproduced the textbook result almost
exactly - an 89% win rate with drawdowns under 1%. And one genuine
surprise: the academic prediction says reversion should only work on
boring low-volatility stocks, but our high-beta names actually made the
most money (bigger panics, bigger snap-backs) at six times the
drawdown. So all 24 bots - 4 ETFs, 10 large caps, 10 high-beta - are
now live as a forward experiment to settle that question with real
data. One honest design note: this strategy runs without a tight stop
loss, because Connors' own research shows stops break it; position
caps bound the risk instead. Fleet's at 110 bots. Watch the dips get
bought on the dashboard.
On the roadmap: covered calls - and the honest reason they're not here yet
There's a new tab on the site: Covered Calls. Nothing trades there yet -
and that's the point of the page. It explains what covered calls are
(selling call options against shares you own for monthly income - a
sensible strategy, but not free money), and exactly why we haven't
built the sector: options are a whole new instrument type for our
engine, honest backtesting is hard without expensive historical
options data, and the 100-shares-per-contract math is brutal at small
account sizes.
Most trading sites would either sell you an options course or pretend
the feature is "launching soon." We'd rather show you the engineering
reality and the bar it has to clear. When the sector ships, it goes
through the same pipeline as everything else here: build, backtest
honestly, paper trade in the open, publish the results. Read the full
reasoning at trueai.trading/covered-calls.
The VCP bots get better exits - thanks to Jesse Livermore
Our new Minervini VCP bots just had their exit rules upgraded, and the
story of how is the whole project in miniature. We extracted Jesse
Livermore's hundred-year-old "pivotal point" method from a 94-minute
breakdown and A/B-tested it against our VCP bots on 24 stocks and three
years of data. His famous entry technique - buying inside the base where
supply dries up - failed the test outright. But his exit rules won
clearly: sell a third of the position at exactly twice the stop distance
(making the trade risk-free), and dump any trade that goes nowhere for
eight days. Same entries, 50% more profit in the backtest.
So that's what shipped: the entry idea rejected on evidence, the exit
idea adopted on evidence, before the bots had taken a single live trade.
We also tested Dan Zanger's signature rule - breakouts need a volume
surge - and it failed too (killed 8 of 9 good entries on daily data), so
it stayed out. Two legendary traders, three ideas, one adopted. The
numbers decide here, not the legends. Usual caveat: the backtest sample
is small, and the forward paper record on the live dashboard remains
the real judge.
16 Minervini VCP bots join the stock fleet
The Stocks sector doubled again: 16 new bots now trade Mark Minervini's
volatility contraction pattern (VCP) - the method of a two-time US
Investing Championship winner, extracted rule-by-rule from a two-hour
masterclass into code. Trend template, shrinking contractions, volume
dry-up, breakout entry, trailing exit: all mechanical, all logged.
The honest verdict from backtesting: the strict pattern is rare - nine
trades across 24 stocks in three years - but it won 67% of them. When we
loosened the rules to get more trades, the win rate collapsed to 33%.
The quality lives in the strictness, so the strict version ships,
labeled as data collection, not a claimed edge. Expect these bots to sit
patiently idle for months - that's the strategy working, not broken.
Every decision they make is on the live dashboard.
Retired: our first two bots
The fleet's two original bots - btc-sma-001 (simple moving-average
crossover, the very first bot we ever ran) and btc-kama-001 (Kaufman
adaptive moving average trend) - have been retired. Their questions were
already answered: backtests with real fees and slippage showed the SMA
crossover has no edge once costs are counted, and KAMA is a genuine
loser in choppy markets. Keeping them trading was re-confirming closed
verdicts while producing most of the fleet's fill noise.
Their send-off made the point one last time. Both finished with a
positive gross trade P&L - and still ended below their starting cash
(down $14.48 and $8.54 of $1,000) because trading fees ate more than the
strategies earned. "No edge net of costs" in a single line. Their full
history stays in the database and on the dashboard; retirement means no
new trades, not rewritten history. That's the deal we made from day one:
most strategies fail, and when one does, we say so and move on.
Real CME futures data is live
The Futures sector now runs on real Chicago Mercantile Exchange data:
two new bots trade the Micro E-mini S&P 500 (MES) on a live feed from
Webull's OpenAPI, alongside the existing crypto perpetual bots.
Full disclosure, as always: the MES backtests came out breakeven, so
these bots are labeled forward-data collection - not a claimed edge.
They exist to find out whether anything survives real market conditions,
and their every decision is logged to the dashboard like the rest of
the fleet.
Stock watchlist doubled to 16 names
The stock fleet grew from 8 to 16 momentum names. We screened 22
candidates through the same daily backtest the original names passed and
added the top 8 performers (RGTI, HOOD, ARM, APP, CVNA, MU, IONQ, RBLX).
One honest caveat, recorded in the configs themselves: picking the top
of a screen after seeing the results inflates expectations - selection
bias is real. The forward paper record is the judge that counts, same
as everything else here.