Fleet news

What changed and why - strategy verdicts, retirements, new bots, new data. The wins and the failures get the same font size.

We bought the graveyard: testing 'too big to fail' comebacks against 32,851 dead stocks.

Everyone remembers the distressed stocks that came back. Nobody remembers the ones that didn't, because dead companies vanish from ordinary price feeds. So before testing a "buy quality at maximum pessimism" rule, we bought one month of a delisting-inclusive dataset ($19.99, cancelled after) and paired it with the SEC's free as-filed fundamentals: 50,873 US common stocks, 32,851 of them dead, with survival filters (real gross profit, cash runway, non-financial) evaluated only on what was publicly known at each moment.

The pre-registered rule - large caps down 65%+, entered on a bounce, filtered for survivability - PASSED its frozen gates on the 2018-2026 test half: 106 events, +15.3% average per event net of doubled costs, and the filters correctly refused the trades that became J.C. Penney, Peabody and Chesapeake bankruptcies. The graveyard did not destroy the result.

Then our independent validation pass failed it anyway, on one gate: March 2020 alone carries 42.5% of the profit (bar: under 40%), and the earlier 2011-2018 half of history shows no edge at all. A result one month can carry is a result one missing month can erase. So: no deployment. If the family goes forward, it does so as PAPER trading under new pre-registered gates that no single month can satisfy. Along the way we diagnosed and published four vendor data pathologies - including a Vietnamese stock's prices hiding under a dead American ticker - each with its own archived, polluted run. All backtest, all paper, no wagers, nothing bought or sold.

We went hunting for the oldest anomaly in finance. It's real - and it lives exactly where you can't trade it.

Closed-end funds sometimes trade well below the value of what they own, and buying unusually wide discounts is one of the oldest documented edges in the academic literature. We pre-registered a zero-knob test before touching any data: entry rules fixed from the literature, pass bars frozen, one run. Data cost: $0 (a free NAV feed paired with our existing price feed, 301 funds, some histories back to 1999).

The verdict is NO-SAMPLE: on funds liquid enough to actually trade, the registered rule fired 8 times in 4.6 years - nowhere near the 200-trade floor a verdict requires. We did not accept that number on faith. A pipeline audit found 9,075 raw signal days across 258 funds - which collapse to 60 once the price and liquidity floors apply. The anomaly is real. It just lives almost entirely in tiny, thinly traded funds where the bid-ask spread and market impact would eat the very discount you came for.

That is the honest shape of many famous edges: visible in data, concentrated precisely where execution costs are worst. All backtest, no wagers, total spend $0, goalposts untouched.

We tested shorting the market's hottest stocks. The gates said no - and the timing part actually worked.

Our best paper book (double7 plus two Qullamaggie families) earns about $355 a month per $25k in backtests, but a new measurement this week showed WHERE its weakness lives: the book's idle capital concentrates in bear and drawdown months. October 2022 ran 0.3% utilization with a 30-day stretch of zero positions. All three families are long-side systems, so they go quiet together exactly when markets fall.

That pointed at a specific fix: a short-side family that gets BUSIER in falling markets. We pre-registered a study before writing any code - frozen pass bars ($395/month combined, dead months must fill in, results must survive doubled borrow costs), parameters taken verbatim from Laurens Bensdorp's published S2 and S6 short systems so there was nothing to tune, and one run per variant, period.

The verdict: FAIL, all three variants. And the failure has a shape worth reporting. The timing hypothesis was RIGHT - the short systems fired in 9 to 10 of the 11 registered dead months, with essentially zero overlap with the incumbent families' trades (0.0 to 0.4% shared entry dates). They showed up exactly when we wanted them to. They just lost money when they got there: about -$0.19 to -$0.38 per $100 slot per trade after realistic fees, slippage, and borrow costs, dragging the combined book DOWN to $291-$345 a month instead of up past $395.

Being busy at the right time is not the same as having an edge. The book's 1995-2019 short-side results did not survive 2018-2026 on our universe with our costs - shorting a six-day surge in the meme era is a different proposition than it was in 1999. Total spent on this study: $0, all backtest, no wager, no deployment. The engine still cannot short a single share, and after this result it will stay that way.

Bug report: the dashboard said our options bot was up $205. The options page said $80.

Pfizer rallied through the $27 strike of the covered call our options bot is running, and that rally exposed a bookkeeping gap on this site: the main dashboard showed the position up $205.50 unrealized, while the covered-calls page showed $80.00. Both pages, same bot, same moment, $125.50 apart. Regular readers know the house rule by now: any two numbers that disagree expose a bug to anyone who scrolls, and $125.50 happens to be exactly what the difference should be - the short call's in-the-money value.

Here's what happened. A covered call is two legs: 100 shares you own, and a call option you sold against them. The fleet dashboard reads the same snapshots table as every other bot, and that table only ever carried the share leg - so when PFE ran from $26.20 to $28.26, the dashboard cheerfully counted all $205 of stock gain. But roughly $126 of that gain belongs to whoever bought our call: above $27, every extra dollar of stock is a dollar we owe on the option. The options page always knew this - it reads the options bot's own forward log, which marks the short call as a liability every session. The dashboard just wasn't listening to it.

Now it is. Every dashboard view - the sector tiles, the fleet table, the open positions list, the equity chart - pulls options bots' equity from the same two-leg source of truth the options page uses, so the numbers agree by construction rather than by luck. The honest figure is +$80: the stock's gain up to the strike, plus the premium we were paid, minus what the call is worth against us. If PFE stays above $27 through September 4, $80 is also almost exactly what settlement will deliver - the upside past the strike was sold on day one. That's not a flaw in the trade; capped upside in exchange for premium is the entire covered-call bargain. The flaw was a display that showed the upside we'd already sold as if we still owned it.

Worth saying plainly: the mistake made us look BETTER than reality, which is precisely the kind of error a track record cannot afford. Paper trading, as always - and the September 4 settlement will be the first real test of the options plumbing this bot exists to prove.

Retired: the first bot to reach our evidence bar - and the first our gates rejected

btc-breakout-001, one of the fleet's original five bots, is retired today. It earned a distinction on the way out: it was the first bot in the fleet to complete 50 round trips - the sample size our own rules say is enough to judge - and the first one our new graduation gates looked at and rejected. 53 trades over 39 days: net −$61.23 on a $1,000 paper account (−$25 in trading losses, ~$36 in fees on $36,000 of churned notional), with a 28% win rate.

The diagnosis isn't mysterious. It ran Turtle-style Donchian breakouts, long-only, on 15-minute Bitcoin bars. That strategy accepts many small stopped-out losses because the occasional big runner is supposed to pay for all of them. In a choppy, drifting market the runners never arrive - its best trade made +$6.28 while its worst lost $7.17, and the 1% stop dutifully converted every fake breakout into a small paid exit. Meanwhile its mirror image - the SHORT Donchian family on perpetual futures - remains green in the same market. Same arithmetic, opposite regime.

Every number above stays in the database, and the verdict is recorded in the bot's config file where the next reader will find it. Retired bots keep their final balance on the dashboard - the row freezes rather than disappears, because deleting your losers is how track records get faked. Paper trading, as always - the only thing spent here was electricity, and the only thing earned was the receipt.

Arbitrage verdict: we measured it for 16 days so we didn't have to build it

Two weeks ago we pre-registered a question: does any arbitrage family produce opportunities that clear our real costs, at a speed a $6 REST-polling server can actually capture? Three studies, all thresholds committed in writing before the first quote was logged (docs/arbitrage/SPRINT-1.md, commit 600be4a). The window closed today: 4,346 five-minute scans over 16 days across Crypto.com, Kraken, and Coinbase. One bucket is missing from August 7 - we briefly wiped our own crontab and restored it within minutes; the gap is disclosed rather than papered over.

Study A, funding-rate harvest, was the primary hope: go long spot and short the perpetual, collect the funding payments, stay market-neutral. Verdict: FAIL, on all three pre-registered checks that matter. Funding rates did spike (ADA annualized at +54% at one point), but a round trip costs ~0.40-0.46% in fees and spreads, and the best episode in the test window captured only +0.226% of funding over 86 hours - still net negative. All eight test episodes lost money. The math is simple and brutal: at retail taker fees, the funding tide never rose high enough to clear the toll.

Study B, triangular loops on one venue, technically PASSED its floor: 171 persistent net-positive loops across 14 days. Before anyone gets excited: every single one was on thin USDT crosses like JASMY and CSPR, past the venue's liquidity cliff, measured at top-of-book with zero depth. Our own protocol says exactly what that's worth: a paper pass at top-of-book earns a depth-aware follow-up study, never a deployment. We suspect the depth check kills it. We'll measure that too.

Study C, cross-venue price gaps: zero qualifying events in 16 days. And when we verified fee tiers at verdict time (as the protocol required), Kraken's real base taker turned out to be 0.40%, not the 0.26% we'd assumed - which would have made the threshold even harder. A FAIL that survives its own assumptions being wrong against it is a comfortable FAIL.

Consequence, decided before the data existed: no arbitrage engine gets built. Limit orders, multi-leg atomic execution, second exchange accounts - all of it stays unbuilt, because a two-week observation study costing almost nothing just told us it would have been wasted. The logger keeps running (it's nearly free) and the study re-opens automatically if the funding regime changes. This is the third time this summer a pre-registered feasibility study has killed a build before it started. That's not a failure of the process. That is the process.

New free tools: the number prop firms don't advertise, and the one your win rate is hiding from you

There's a new "Prop Firm Tools" tab on the site. Two free calculators, no signup, no email. Here's what they do and why we built them.

TOOL ONE: the evaluation simulator. If you've looked into funded trading accounts - Topstep, Apex, and the rest - you pay a fee, trade to a profit target without hitting a drawdown limit, and get a funded account. Sounds like a skills test.

Here's the thing everyone misses: whether you pass depends enormously on the ORDER your wins and losses happen to arrive in. Same strategy, same trades - losing streak lands in week one, when your cushion is thin, and you're out. Same streak in month three, you're fine. You only ever get one shuffle of the deck. So instead of replaying your history once, the simulator shuffles it ten thousand times and tells you what fraction of orderings pass.

And it computes the number the industry has no reason to show you: a strategy with NO edge at all - literally a coin flip after costs - passes a typical evaluation roughly one time in five. Enough people buy enough attempts that a steady stream of funded accounts goes to traders with no edge whatsoever. Passing, by itself, is much weaker evidence than it feels like.

The comparison tab runs your same numbers against ten real firms' actual rules. We read each firm's own rulebook and help center - not review sites, not affiliate pages - and recorded every number with a source link and a date. The fine print mattered more than we expected: whether a "trailing drawdown" tracks your open profits in real time or only your end-of-day balance moves the pass rate by double digits on the same strategy, and firms bury that distinction. Anything we couldn't confirm from a primary source, the page refuses to display - and rows automatically get flagged for re-verification once they're 90 days old, because these terms change without announcement.

TOOL TWO: the breakeven win rate calculator. It corrects the most expensive misconception in trading: that your win rate is your edge. It isn't. A 60% win rate sounds great and LOSES money if your average win is half your average loss. Enter your average win, average loss, and costs, and it tells you the win rate you need just to break even - plus the part almost nobody does: given how many trades your win rate is based on, whether you can even statistically tell it apart from a coin flip yet. Over 100 trades, a measured 55% could easily be a true 45%. The tool says so, out loud, and tells you how many trades it would take to know.

What these tools are NOT: advice. They don't recommend a firm, don't predict your results, and can't tell you whether your edge is real - every number assumes the stats you type in are honest and will hold up out of sample. If they came from a backtest you kept tweaking until it looked good, they won't. And to be clear, none of this is our bots' trading - our fleet results are separate, all PAPER money, and these calculators are just math on numbers you provide.

Our bots were hoarding cash like squirrels. We built them a shared piggy bank - then caught ourselves cheating three times

Here's a problem nobody warns you about when you build a bot fleet.

Every bot in our stock fleet keeps its own money. Bot #1 has its pile, bot #2 has its pile, all forty of them. Which means you have to fund the worst case: the day every single bot wants to buy at once. So you set aside forty piles of cash.

That day basically never comes. We measured it. Our bots are all holding positions at the same time about 4-25% of the time. The rest of the time, most of that money is sitting there doing absolutely nothing, like a gym membership in February.

So we built a shared piggy bank. Instead of forty separate piles, the bots draw from one pool of ten "slots." Want to buy? Take a slot. Sold? Put it back. All ten taken? You sit this one out.

The trade is real: sometimes a bot spots something and gets told "sorry, we're full." You give up a few trades. In exchange you need far less cash standing around to run the same strategies. Nothing about the strategies changed. We just stopped funding a traffic jam that doesn't happen.

Now the fun part: getting here required being wrong three times, and we'd rather tell you about those than the tidy version.

MISTAKE ONE. Claude first concluded that adding more strategies makes returns WORSE, complete with a convincing table. That turned out to be a bug in how it was splitting the money between them, not a fact about the strategies. Under shared funding the answer flips: adding strategies helps. A confident table is not evidence.

MISTAKE TWO. The first version of the shared-piggy-bank test said we'd earn about 66% a year. Wonderful! Also nonsense. When several bots wanted a slot on the same day, the code was quietly handing it to whichever trade would finish soonest - which is information nobody has until afterward. It was letting the bots peek at the answer key. Fixed, the number dropped by a third. Suspicious good news is still suspicious.

MISTAKE THREE, the scary one. We ran the official test. It said PASS. It was also silently missing 10 of our 24 Double 7s bots - because ten stocks are traded by two different bots each, and our code had been quietly overwriting one of each pair. The test passed on a fleet that wasn't the fleet. The only reason we caught it: the report printed "14 bots" next to a rule that said 24.

That last one is the whole lesson. A test that fails is annoying. A test that PASSES on quietly wrong data will happily send you live with real money, smiling the entire time. Print your inputs. Check them against what you promised to test.

What this is NOT: proof that our stock strategies make money. Those are still on trial, with their own verdicts pending, and they might well flunk. A shared piggy bank makes whatever edge you have go further. If there's no edge, congratulations - you've built a more efficient way to lose money.

The forty new bots now running are a plumbing test, not a profit test. We're checking that slots actually get handed out and handed back, that a server restart doesn't lose track of who's holding what, and that ten slots never somehow becomes eleven. We wrote down what would count as passing before switching them on. All PAPER money, as always. The strategy tests already running were not touched - new bots only, so we don't contaminate an experiment that's halfway done.

The mean-reversion math study came back negative - and that's worth publishing

Three weeks ago we pre-registered a hypothesis: our Connors-style mean-reversion bots should do better on stocks whose prices statistically snap back fast (short OU half-life, strongly negative Dickey-Fuller stat - the textbook mean-reversion measurements). The judging rules were committed before the study ran, including a trap for exactly what ended up happening.

On RSI-2, the primary system, both measurements looked adoptable: the predicted trades beat the average in both the train and test halves. If we'd stopped there, we'd have deployed new gated bots this week. But the protocol required the same pattern to show up in Double 7s - a sibling strategy trading the same dip on the same 24 stocks - and there it ran BACKWARDS: the "best" trades by the math underperformed, and the training data actually favored the opposite bucket. One strategy's promising pattern failing to transfer to its own sibling on identical data is what overfitting looks like from the inside.

So: no new bots, hypothesis rejected, receipts in the repo. The 48 live Connors bots were never touched - they keep running exactly as deployed. PAPER trading throughout.

First formal forward verdict: gated Bollinger is INCONCLUSIVE - positive, but not by enough

Four weeks ago we committed the judging rules for our first formal forward test to git - before the outcome was knowable. The condition-gated Bollinger fleet (17 crypto spot bots, PAPER) promised +$0.364 per trade at a 70% win rate based on its backtest. The pre-registered bar for a forward PASS: positive average, win rate at least 55%, and at least half the backtest's per-trade average.

The verdict at trigger: 62 completed round trips, +$3.11 total, +$0.05 per trade, 66.1% win rate. That's positive, and the win rate essentially delivered on the promise - at the 50-trip mark it was exactly 70.0%. But +$0.05 is a long way from the +$0.182 the protocol requires, so by the rules we wrote in advance this is INCONCLUSIVE, not a pass: the test extends four more weeks, same criteria, re-verdict 2026-09-07.

What the number hides: the gates are finding winners at the promised rate, but two names (NEAR and SHIB) produced losses big enough to eat most of what the other thirteen bots earned. Our protocol explicitly forbids excluding them after the fact - "it passes if you ignore the losers" is how fake edges get published. Two honest confessions, on the record: the 50-trip trigger technically fired eight days before we noticed - the verdict comes out identical at both cuts, and the miss itself is documented in the protocol file - and the protocol text said "13 bots" when the fleet was 17: a miscount, all 17 ran from day one.

Everything here is paper trading. Nothing about this result moves real money.

The full verdict as a video, receipts on screen: https://youtu.be/pIUTxvX0kYo

The first options bot is live (PAPER) - and it isn't allowed to prove the strategy works

Two days after the covered-calls build started, the studies are done and the first bot is running. The short version of the studies: covered calls and the wheel both passed their pre-registered tests against simply holding the shares - 4 out of 4 stocks each - but the honest reading is narrower than the scoreboard. Not one stock won on BOTH return and risk. In rallies the strategies kept you safer but gave up most of the upside; in declines they lost less but still lost. And the wheel's famous cycle - sell puts, get the shares, sell calls, get called away, repeat - completed exactly three times across four stocks and thirty months of history. The honest sales pitch after measuring it: option premium pays you to wait out a drawdown. It does not prevent one.

That was enough to earn one bot, and only one: cc-pfe-01, which now owns 100 shares of Pfizer and has sold one call option against them, on paper, checked once a day after the close. PFE got the slot on pre-stated evidence - it was the only stock that passed both ways in either study.

Here's the part we care most about: the rules for judging this bot were committed to git before it made a single decision, and they explicitly forbid the one conclusion everyone wants. Ninety days of one stock cannot prove covered calls "work," so the verdict doesn't ask that. It asks whether the plumbing holds up (options expire and settle correctly, nothing breaks on a restart) and whether the real-world costs match what the studies assumed - because if live spreads are meaningfully worse than the studies charged, the studies get flagged as too optimistic and expansion stops. Return versus buy-and-hold gets reported, but it is not allowed to gate anything. Day one delivered its first two data points, both mildly good news: the live spread on the call we sold was tighter than the studies assumed, and the bot selected its contract using live option greeks - something the backtests never had access to, which is exactly the gap this forward sample exists to measure.

As always: this is PAPER trading and forward data collection on Alpaca's free indicative feed - calculated quotes, not the paid exchange feed. Real math, no financial advice.

The fifth sector begins: covered calls - and the first thing we built is the test that could kill it

The covered-calls page on this site has said "coming soon, and here's the honest reason it isn't here yet" for a while. Today the build actually started. The plan: bots that own 100 shares of a cheap, liquid stock and sell call options against them, collecting the premium - plus the "wheel," which sells puts to get into the shares in the first place. Fancier options strategies are explicitly parked until these two prove themselves.

True to form, the first code written wasn't a trading bot - it was a measuring stick. Options quotes have a catch: the gap between the buying and selling price (the spread) can quietly eat the entire premium you collected. So before building the engine work - the biggest on our roadmap - we froze a five-stock universe using criteria committed to git first (NOK, PFE, T, CMCSA, BITO), and started a three-day study measuring whether real spreads leave anything worth collecting. The pass/fail bar was also written down first: if the cost of trading eats more than a quarter of the premium on every single name, the sector stops before the engine gets built. Day one looks promising - four of five names comfortably under the bar - but day one is not a verdict. The verdict runs Thursday and gets published either way.

One honesty note up front: our options data is Alpaca's free "indicative" feed - calculated quotes, not the paid real-time exchange feed. Every number from this sector carries that label until an edge looks real enough to justify paying for the real thing. Paper trading, real math, no financial advice - as always.

New video: is crypto arbitrage still possible in 2026? We measured it.

Arbitrage is the internet's favorite "free money" pitch, so instead of buying anyone's course we pointed a read-only logger at it: every five minutes it scans all ~300 triangular loops on our exchange, logs cross-exchange price gaps against two other venues, and records the hourly funding rate on the liquid perpetuals. The pass/fail criteria were committed to git before the first quote was logged.

Four days in: 890 scans, and the median best triangular loop on the whole exchange loses 0.26% after fees - the market is efficient to almost exactly the width of the fee. A handful of scans did show "profitable" loops, almost all through tiny, illiquid pairs where the advertised price is money nobody can actually collect at size. The one family still genuinely open is funding arbitrage - hold the coin, short its perpetual, collect the hourly funding payment. It's the only kind of arbitrage that rewards patience instead of speed, which is the only race a $6 server can enter, and Bitcoin's funding swung from roughly −10% to +15% annualized during the study's first days.

The new video walks through it with the real numbers: https://www.youtube.com/watch?v=DgO9pC1UEG4 - the verdict lands August 15 under the pre-registered criteria and gets published either way. As always: paper trading, real prices, real fees, no financial advice.

We tested day-trading bots. The stock version failed its own test - so we're not building it.

This week the fleet learned to roam: instead of trading a fixed watchlist forever, our newest bots scan the entire market - all 13,000+ US stocks - and compete for a fixed number of position slots. Two swing-trading roamers are live on paper now, each judged against the hand-picked bots running the identical strategy, so the only thing being tested is whether scanning beats picking.

Then we asked the obvious next question: day trading. For stocks, we took a famous day-trade setup - opening-range breakouts on big gap-ups, rules taken from the trader's own published write-ups - built a dataset of 3,377 real gap days across 31 months, and wrote the pass/fail criteria down BEFORE running anything. The result: the setup as actually specified is rare (about 38 qualifying trades in 31 months), and the apparent profit on those trades vanished entirely when we doubled the modeled slippage. An edge that lives inside the fill-model assumption isn't an edge. Verdict: the stock day-trading bot is not being built. One afternoon of honest testing beats months of a bot trading noise - and, as always, a pre-registered test gets published whichever way it lands.

One day-trading bot did earn deployment: a crypto roamer running the exact same gated Bollinger strategy as our 13 fixed crypto bots - verbatim, enforced by an automated test - rotating across the 20 most liquid pairs instead. Paper money, labeled data collection, judged by criteria committed before its first trade. And if the underlying strategy fails its own long-scheduled verdict on August 10, the roamer parks with it. No rescues, in either direction.

24 new bots: our best strategy just got a sibling - and they'll compete head-to-head, live

The fleet just grew from 113 to 137. The newcomers run Connors Double 7s, from the same 2008 book as our strongest strategy - the Connors RSI-2 fleet that buys 2-3 day panics inside uptrends. Double 7s is its simpler sibling: if a stock closes above its 200-day average and today's close is the lowest of the last 7 days, buy. Exit when the close is the highest of the last 7 days. That's the entire system. No stop-loss, by the book's own research - the dip IS the entry logic - with our position caps bounding the risk instead.

Why it earned a slot: we backtested it against RSI-2 on the exact same 24 stocks, the same 2000 daily candles, the same costs. It took 922 trades to RSI-2's 431 - a 7-day-low close simply happens about twice as often as a 2-period RSI under 5 - kept a comparable 67% win rate, earned MORE per trade, and finished with 2.6 times the total profit, green on 22 of 24 names. That's the largest backtest sample of anything we've ever deployed, passing our 50-trade evidence gate eighteen times over.

Now the honest frame. A backtest is an audition, not a verdict - that's why the new bots trade paper money on the very same stocks as their RSI-2 siblings, at the same sizing, side by side. Two systems from the same author, hunting the same regime, judged by the same forward tape. Either the backtest edge shows up in live market conditions or it doesn't, and the dashboard will say which in public, either way.

Watch them (bot IDs starting with d7-) at trueai.trading/dashboard

All results are PAPER - simulated money, real market prices, real fees. Not financial advice.

The exit experiments are complete: exits don't stack, and our best result just survived its second test

Last test in the exit series. This morning we showed that combining our two best exit rules on crypto changed nothing - the fast exit fired first on every trade. The remaining question was the stock side, where the slower Darvas box-trail is the champion: does adding the fast Heikin Ashi flip on top help there?

It made things worse. Five exit variants, identical data, 16 stocks, 2000 daily candles each: the incumbent 10 EMA trail made +$533. The HA-flip exit alone made +$409 (cuts long stock trends early, as we already knew). The combination made +$388 - below BOTH of its components' better halves. Unlike crypto, the Darvas floor did occasionally fire first on stocks, and when it did, it subtracted. So the cross-market conclusion is airtight now: a combined exit is always dominated by its faster component. You never get the best of both rules - you get the fast one, plus extra ways to lose. Exits don't stack.

The genuinely good news from the same run: the Darvas box-trail won on stocks AGAIN, on a second, longer test window - +$684 versus the incumbent's +$533, a 28% improvement, earned on fewer trades. Its first win was +90% on a shorter window; the honest reading is that the direction is now confirmed twice while the size of the edge depends on the window. Two independent windows agreeing on direction is worth more than one window's big number.

So the exit playbook is settled and waiting: patient Darvas trail for stock breakouts, fast HA-flip for crypto, no combinations anywhere, and the regular stop-loss underneath everything. All of it stays parked until the mid-August strategy verdicts, then gets re-verified out-of-sample before touching a single live bot.

All results are PAPER - simulated money, real market prices, real fees. Not financial advice.

Follow-up: we tried combining our two best exits. The answer was a perfect, boring zero.

Quick follow-up to today's Heikin Ashi test. We had two exit rules that each won somewhere: the fast HA color-flip (beat our deployed exit on crypto) and the looser Darvas box-trail (nearly doubled it on stocks). Natural question: on crypto, should the fast exit also carry the Darvas safety floor underneath - exit on whichever triggers first?

So we ran it. The combined exit produced results IDENTICAL to the plain HA-flip - to the penny. Same 50 trades, same wins, same +$83.98. Not similar. Identical. The reason is simple once you see it: a red Heikin Ashi candle shows up long before price falls 5% through a box bottom or gives back 20% from a peak, so the fast exit fired first on every single trade and the floor never got touched. It was dead code wearing a seatbelt.

Verdict: skip the floor. The crypto exit candidate stays the plain HA-flip - fewer moving parts, same result - and the engine's regular stop-loss remains underneath everything as the actual catastrophe brake. A zero-effect result is still a result: it's one less parameter to fit, one less thing to break, and one more combination we'll never have to wonder about.

Nothing deploys until the mid-August verdicts, same as always. All results are PAPER - simulated money, real market prices, real fees. Not financial advice.

We tested Heikin Ashi honestly. It beat our exit on crypto and lost on stocks - and that's the interesting part.

Heikin Ashi charts are those smooth, beautiful candles that make every trend look obvious in hindsight. A viewer-suggested video pitched a strategy around them, so we did what we always do: mined the rules, researched the technique, and put the testable part through an honest A/B.

First, the trap - because most Heikin Ashi results you see online are fake, and not subtly. HA candles are averages: the prices they display DO NOT EXIST in the market. Backtests that fill orders at HA prices systematically overstate results, badly enough that TradingView built a dedicated setting to prevent it. Our tests compute HA colors from real candles and fill every order at real prices. If you remember one thing about Heikin Ashi, make it that.

The video's entry ("doji at a key support level") is drawn-by-eye discretion, so it can't be coded faithfully and we didn't pretend to. But its exit rule is fully mechanical: ride the trend until the first opposite-color HA candle. We tested that against our deployed 10 EMA trail exit on both Qullamaggie breakout fleets, same entries, same candles, real costs.

Crypto (50 trades): the HA-flip exit made +$83.98 vs the incumbent's +$49.37 - the first exit variant to beat our deployed one on crypto. Stocks (150+ trades): it LOST, +$409 vs +$533, because it cuts long stock trends early. And that mirrors last week's Darvas result perfectly, where a looser trail nearly doubled stock profits but hurt crypto. Three exit experiments now tell one story: stock trends run long and want patient exits; crypto daily trends are short and spiky and want fast ones. One exit does not fit all markets.

Nothing changes on the live fleet today. Both Qullamaggie fleets are inside their forward-verdict window (due mid-August), and changing exits mid-test would corrupt the verdict - the same goalpost rule as always. The market-split exit idea goes in the queue for after the verdicts, where it gets re-verified out-of-sample before touching a single bot. Caveats on the record: the crypto sample is exactly at our 50-trade evidence gate, one window, comparative only.

All results are PAPER - simulated money, real market prices, real fees. Not financial advice.

Next sector on the roadmap: prediction markets - and why the obvious bot is the hardest one to trust

There's a new tab on the site: prediction markets. Kalshi-style event contracts - yes/no markets on real-world events where the price IS the probability. A contract at 30 cents means the market says 30%. We want to know whether bots can find edges there, and the sector is now officially on the roadmap, parked behind the two strategy verdicts due in August.

Here's the part worth reading. The obvious bot - an AI that reads the news and estimates probabilities better than the market - is easy to build and nearly impossible to validate honestly. Every backtest of it is a lie by construction: any AI model you'd use today was trained on data that includes how past events turned out. Asking it "what were the odds of X?" about anything in its training window isn't forecasting, it's remembering. So if we ever run that bot, it starts from zero in live forward testing, labeled as the experiment it is, for however many months that takes. No shortcuts exist, so we won't pretend to have found one.

The bots that come FIRST are the boring ones: documented structural patterns like the longshot bias (unlikely events are systematically overpriced) that don't require understanding the news at all, and can be partially backtested against real fee models. Fees matter enormously in these markets - they're highest exactly where the outcome is most uncertain.

When the sector ships, it starts on Kalshi's regulated demo environment: paper first, always, same as every sector before it. The full honest assessment - including what we can't test and why - is on the new page: https://trueai.trading/prediction-markets

As always: all results PAPER, nothing here is financial advice.

Gated-Bollinger interim report: 19 trades, 68% wins, still down $2.68

Two weeks ago we committed, in public and in advance, to a verdict protocol for our gated Bollinger fleet: pass/fail criteria defined before the outcome, an interim report at the two-week mark, and a final verdict at 50 trades or August 10, whichever comes first. This is the interim report. Numbers only - the protocol forbids verdict language until the trigger, so you'll get none here.

Since the forward test began on July 13, the 13 gated bots have completed 19 round trips: 13 winners, a 68.4% win rate, for a total of -$2.68 - an average of -$0.14 per trade, net of fees. For reference, the backtest that earned this fleet its slot promised +$0.36 per trade on a 70% win rate. When we committed the protocol at 13 trades, the fleet stood at 54% wins and -$0.21 per trade; both numbers have since moved toward the backtest, and the win rate is now within touching distance of the promise. The average, however, is still negative.

Where the money actually went: one symbol, NEAR, accounts for -$6.13 across five trades, including three losses of about -$3.30 each - several times larger than any other bot's loss. The other six bots that traded are +$3.45 combined, in the small-frequent-win shape the backtest predicted. We note this as a fact about composition, not as an excuse: the protocol explicitly bans "it passes if you exclude NEAR" reasoning at verdict time, and we intend to honor that. Six of the thirteen bots haven't traded at all yet - the condition gates are doing exactly what gates do, which is keeping bots out of most markets most of the time.

At the current pace the fleet clears the 25-trade minimum well before August 10, so the verdict should arrive on schedule. Whatever it says, it gets published. All results are PAPER - simulated money, real market prices, real fees.

Position sizes go from $100 to $500 on 63 bots - here's exactly why, and why not all 110

From today, 63 of our bots trade with a $500 position cap instead of $100. Nothing about their strategies, signals, or leverage changed - every bot still trades unleveraged cash positions on $1,000 of paper equity. So why bother?

Visibility. With a $100 cap, a bot that catches a clean 2% winner moves its account by 0.2% - a rounding error that makes real edges and real failures look identical on the dashboard. At $500, that same trade moves equity 1%. The experiment doesn't change; the signal-to-noise on reading it does. We scaled the daily-loss halt by the same factor (from $20 to $100), so every risk limit still sits at the same percentage of position size the backtests were run with.

Now the honest part: 47 bots did NOT get the change. The 13 gated Bollinger bots are two days from their interim report, under a verdict protocol we published in advance with pass/fail thresholds defined in dollars per trade. Changing their position size mid-sample would make those numbers meaningless - a quiet way of moving the goalposts we promised not to move. Same for the Qullamaggie fleet, whose verdict lands mid-August. They finish their tests at the size they started. Once each verdict is published, whichever way it goes, they move to $500 like everyone else.

One thing to keep in mind reading the dashboard from here: dollar swings on the 63 resized bots are 5x bigger as of today, so before-and-after comparisons of dollar P&L need that context. Percentage returns remain comparable throughout. As always: all of this is PAPER - simulated money, real market prices, real fees.

Autopsy of our two worst bots: it was never really about the fees

btc-kama-001 and btc-sma-001 - two of our original bots, benched since July 15 - got a final review today, because we wondered: they traded constantly on 1-minute bars, and every round trip pays about 0.3% in fees and slippage. Were the fees the killer? Would slower bars, or bigger positions, have saved them?

So we backtested both strategies at every timeframe from 1 minute to 4 hours, with real costs, over up to a year of data. The answer is no - and the reason is more interesting than the question. At every timeframe, the signals lost money BEFORE fees. Win rates ran 12-34%. These are long-only trend-followers, and in a year where Bitcoin chopped its way down about 44%, they bought every fake rally and got stopped into every resumption of the slide. Bigger positions would have lost more, not less. The fee story was a comforting misdiagnosis; the signal was the problem. Verdict recorded in both configs, bots archived for good.

One genuinely useful thing fell out of the autopsy: flip KAMA around. The SHORT-side version on 4-hour perp bars was the single profitable cell in the whole matrix - +1.29% over a year in which buy-and-hold lost 44.9%. That is a tiny number and one backtest window, so it proves nothing yet. But it matches the market we're actually in, so it joins the fleet as perp-kamashort-btc-s01 - labeled data collection, not edge, until it earns a real forward verdict. (The SMA short version came out at +0.21% - indistinguishable from zero, so it gets no second life. Not every autopsy finds an organ donor.)

Bug report: the dashboard said 3 open positions. The table said 8.

Spotted on our own dashboard: the open-positions counter said 3 while the positions table right below it listed 8. Both can't be right - and the table was.

The cause is a classic: when the fleet learned to short-sell, short positions became negative numbers, and one counter still asked "is the position greater than zero?" - counting only the longs. We'd already fixed two spots with this bug when shorts went live; this was the third and last, and a sweep of the code confirms the class is now extinct. The counter and the table agree: 8 open positions, 5 of them short.

Why publish a one-line bug fix? Because this is what an open book is for. Every number on this site renders from the same database, so any two numbers that disagree expose a bug to anyone who scrolls - no trust required. When it happens, we fix it and say so. Transparency isn't just for the wins; it's for the counter that couldn't count.

Fleet leaderboard: the 3 best and 3 worst bots right now

First edition of an occasional honest leaderboard - with the caveat up front that these are tiny numbers by design: every bot runs a $1,000 paper account with positions capped at $100, so the dollars are small and the point is the process.

The top three: boll-ada-g01 (+$1.57), boll-uni-g01 (+$1.23), and btc-bollinger-001 (+$1.08). The first two are gated Bollinger bots - our one backtest-validated edge - quietly leading the fleet days before their formal forward-test verdict. Encouraging, not conclusive.

The bottom three: rsi2-cvna-s01 (−$9.08), boll-near-g01 (−$6.93), and btc-breakout-001 (−$4.06). The worst is our brand-new dip-buyer's very first trade - it bought a Carvana panic and the panic kept going. Its strategy explicitly holds through this (no tight stop, by published design), so this position is the strategy's first real exam, still in progress.

Now the most instructive line in the data: the fleet's #2 best bot and #2 worst bot run the IDENTICAL strategy - gated Bollinger - on different coins. UNI up, NEAR down, same rules. That's why no single bot's result means anything here, why we deploy strategies as fleets, and why verdicts only come from many trades over many symbols. The first such verdict is due within days: the gated Bollinger forward test. Leaderboard, meet judgment day.

110 bots broke our database (a little). Fixed.

A growth milestone, told honestly: the fleet got big enough to strain its own infrastructure. With 110 bots each logging every decision, trade, and account snapshot, the database was being written to almost constantly - and anything trying to READ it (like the live dashboard, or the fleet report that caught this) kept getting locked out while the bots held the pen. Some queries had also quietly grown slow: what was instant at 10 bots was crawling at 40,000 rows.

The fix is textbook database engineering: write-ahead logging, which lets readers and writers work at the same time instead of taking turns, plus proper indexes on the columns the dashboard queries. The heaviest dashboard query went from struggling-to-connect to 34 milliseconds. The fleet restarted with every open position intact - including our new dip-buyer's first trade, currently underwater and holding exactly as its strategy prescribes.

Why post about plumbing? Because this is what running 110 bots actually involves, and most trading content skips it. The strategies get the headlines; the infrastructure keeps the lights on. Both run in the open here.

Our first video is up: Can AI actually trade?

The YouTube channel has its first full-length video: "Can AI Actually Trade? I Built a Fleet of Bots to Find Out (Honestly)." It's the story of this whole project - the fleet, the honest-testing pipeline, the strategies that survived the backtests and the ones that didn't - told from the beginning.

If you've been following along here, the video is the big picture in one sitting; if you're new, it's the best place to start. It's on the videos page right here on the site, or on the channel. More videos will follow as the fleet's verdicts come in - the gated-Bollinger forward test wraps up around late July, and that result is getting a video whichever way it goes. Wins and losses alike, as always.

We finished mining trading YouTube. Here's the scoreboard.

Since this project started, our strategy pipeline has run on one input: find a documented trading method, extract its rules from the source, implement them in code, and let backtests decide. This week we finished working through the entire YouTube source queue - every channel either mined or ruled out, with the reasoning committed to the repo.

The scoreboard is stark. Sources with published, verifiable rules - Qullamaggie's breakout method, Mark Minervini's VCP, Jesse Livermore's exits, Larry Connors' RSI-2 - produced every deployed strategy in the fleet. Sources built on drawing lines by eye produced nothing we could even test: when the lines get redrawn daily "until they fit," the signal that would have existed at decision time is undefined, and a backtest of it would be fiction. And a pattern worth knowing: the channels that couldn't state their rules precisely were usually the ones selling a course or funneling to a prop-firm evaluation. The method is the product when the method works; the course is the product when it doesn't.

So the hunting ground moves to where the evidence lives: published books, academic papers, and quant research - which is exactly where our newest and best-tested strategy came from. Every verdict, adopted or rejected, is in the open in our strategy-sources docs. The numbers decide here. They always have.

24 new bots buy fear: Connors RSI-2 joins the fleet

The Stocks sector just got its first mean-reversion strategy - and the strongest backtest this project has ever produced. Connors RSI-2, from Larry Connors' published 2008 rules, buys two-to-three-day panic dips inside long-term uptrends and sells the snap-back. It's the exact opposite regime to our breakout bots: they buy strength, these buy fear.

The numbers, with our usual punitive fee assumptions and the book's parameters untouched: 204 trades across three universes and four years of data, 75% winners. Index ETFs reproduced the textbook result almost exactly - an 89% win rate with drawdowns under 1%. And one genuine surprise: the academic prediction says reversion should only work on boring low-volatility stocks, but our high-beta names actually made the most money (bigger panics, bigger snap-backs) at six times the drawdown. So all 24 bots - 4 ETFs, 10 large caps, 10 high-beta - are now live as a forward experiment to settle that question with real data. One honest design note: this strategy runs without a tight stop loss, because Connors' own research shows stops break it; position caps bound the risk instead. Fleet's at 110 bots. Watch the dips get bought on the dashboard.

On the roadmap: covered calls - and the honest reason they're not here yet

There's a new tab on the site: Covered Calls. Nothing trades there yet - and that's the point of the page. It explains what covered calls are (selling call options against shares you own for monthly income - a sensible strategy, but not free money), and exactly why we haven't built the sector: options are a whole new instrument type for our engine, honest backtesting is hard without expensive historical options data, and the 100-shares-per-contract math is brutal at small account sizes.

Most trading sites would either sell you an options course or pretend the feature is "launching soon." We'd rather show you the engineering reality and the bar it has to clear. When the sector ships, it goes through the same pipeline as everything else here: build, backtest honestly, paper trade in the open, publish the results. Read the full reasoning at trueai.trading/covered-calls.

The VCP bots get better exits - thanks to Jesse Livermore

Our new Minervini VCP bots just had their exit rules upgraded, and the story of how is the whole project in miniature. We extracted Jesse Livermore's hundred-year-old "pivotal point" method from a 94-minute breakdown and A/B-tested it against our VCP bots on 24 stocks and three years of data. His famous entry technique - buying inside the base where supply dries up - failed the test outright. But his exit rules won clearly: sell a third of the position at exactly twice the stop distance (making the trade risk-free), and dump any trade that goes nowhere for eight days. Same entries, 50% more profit in the backtest.

So that's what shipped: the entry idea rejected on evidence, the exit idea adopted on evidence, before the bots had taken a single live trade. We also tested Dan Zanger's signature rule - breakouts need a volume surge - and it failed too (killed 8 of 9 good entries on daily data), so it stayed out. Two legendary traders, three ideas, one adopted. The numbers decide here, not the legends. Usual caveat: the backtest sample is small, and the forward paper record on the live dashboard remains the real judge.

16 Minervini VCP bots join the stock fleet

The Stocks sector doubled again: 16 new bots now trade Mark Minervini's volatility contraction pattern (VCP) - the method of a two-time US Investing Championship winner, extracted rule-by-rule from a two-hour masterclass into code. Trend template, shrinking contractions, volume dry-up, breakout entry, trailing exit: all mechanical, all logged.

The honest verdict from backtesting: the strict pattern is rare - nine trades across 24 stocks in three years - but it won 67% of them. When we loosened the rules to get more trades, the win rate collapsed to 33%. The quality lives in the strictness, so the strict version ships, labeled as data collection, not a claimed edge. Expect these bots to sit patiently idle for months - that's the strategy working, not broken. Every decision they make is on the live dashboard.

Retired: our first two bots

The fleet's two original bots - btc-sma-001 (simple moving-average crossover, the very first bot we ever ran) and btc-kama-001 (Kaufman adaptive moving average trend) - have been retired. Their questions were already answered: backtests with real fees and slippage showed the SMA crossover has no edge once costs are counted, and KAMA is a genuine loser in choppy markets. Keeping them trading was re-confirming closed verdicts while producing most of the fleet's fill noise.

Their send-off made the point one last time. Both finished with a positive gross trade P&L - and still ended below their starting cash (down $14.48 and $8.54 of $1,000) because trading fees ate more than the strategies earned. "No edge net of costs" in a single line. Their full history stays in the database and on the dashboard; retirement means no new trades, not rewritten history. That's the deal we made from day one: most strategies fail, and when one does, we say so and move on.

Real CME futures data is live

The Futures sector now runs on real Chicago Mercantile Exchange data: two new bots trade the Micro E-mini S&P 500 (MES) on a live feed from Webull's OpenAPI, alongside the existing crypto perpetual bots.

Full disclosure, as always: the MES backtests came out breakeven, so these bots are labeled forward-data collection - not a claimed edge. They exist to find out whether anything survives real market conditions, and their every decision is logged to the dashboard like the rest of the fleet.

Stock watchlist doubled to 16 names

The stock fleet grew from 8 to 16 momentum names. We screened 22 candidates through the same daily backtest the original names passed and added the top 8 performers (RGTI, HOOD, ARM, APP, CVNA, MU, IONQ, RBLX).

One honest caveat, recorded in the configs themselves: picking the top of a screen after seeing the results inflates expectations - selection bias is real. The forward paper record is the judge that counts, same as everything else here.