RaceChat AI/ Blog (mirror)
This is a static, read-only mirror kept for continuity. For the live site, comments, and full app, visit racechatai.com.

We Tried Every Way to Filter Our AI Down to Its "Best" Picks. Then We Discovered What Actually Works.

6 September 2026·16 min read
We Tried Every Way to Filter Our AI Down to Its "Best" Picks. Then We Discovered What Actually Works.

On 11 March 2026, at Cheltenham Festival, our Elite v4 model produced this day at the office:

Race Elite's pick Result Profit on £10
13:20 King Rasko Grey (11/1) WON +£110
14:00 Kitzbuhel (11/1) WON +£110
14:40 Give It To Me Oj (50/1) 18th −£10
15:20 Final Orders (7/1) WON +£70
16:00 Quilixios (14/1) 5th −£10
16:40 Martator (66/1) WON +£660
17:20 Chicker (80/1) 18th −£10

4 winners from 7 picks. £920 profit on a £70 outlay. A 1,314% return, on Cheltenham Festival day.

If you saw that, you'd have every reason to think our AI understands Cheltenham. Every reason to want to filter its picks down to a short list of "trusted" courses like this one.

We tested that instinct honestly. It's a trap.

Here's what actually holds up when you test it properly — and what your AI is really doing on every card.


The obvious idea — and why it failed

Our Elite model covers 79 UK, Ireland and France courses. The natural instinct is: filter to the ones where the model performs best. Ignore the tracks where it's losing. Cheltenham looks great — surely we should just back its Cheltenham picks and skip everywhere else?

We tested it. Method:

  1. Training window — use the last 60 days of Elite v4 predictions to identify the top-performing courses by ROI.
  2. Test window — measure how those courses actually performed over the next 30 days.
  3. Rebalance every 30 days, the same way an index fund refreshes its constituents.
  4. Walk forward across the full model history so we get multiple independent tests.

This is the only honest way to evaluate a filter strategy. You're not allowed to peek at the future when picking your list.

The top-20 course filter, walk-forward tested:

Test window Bets ROI
Apr 26 → May 25 54 +21.46%
May 26 → Jun 24 82 +78.38%
Jun 25 → Jul 24 305 −39.56%
Jul 25 → Jul 28 28 +39.73%
AGGREGATE 469 −7.18%

Two profitable windows. One catastrophic month (−39.56% on 305 bets — the dominant sample). Net loss.

The tighter refinement — top 10 courses

If top-20 diluted the signal with weaker picks, surely top-10 concentrates it? We ran the same walk-forward test with a top-10 filter.

The top-10 course filter, walk-forward tested:

Test window Bets ROI
Apr 26 → May 25 15 −3.33%
May 26 → Jun 24 39 +17.16%
Jun 25 → Jul 24 165 −46.94%
Jul 25 → Jul 28 15 +0.83%
AGGREGATE 234 −30.40%

Worse. The tighter filter didn't rescue the strategy — it amplified the drawdown. −30% ROI over 234 bets is what backing a losing tipster looks like.

At this point, most content-hungry AI racing sites would quietly file this away and publish a different article about how great their "top 10 courses" system is. We're publishing it because it matters.

The counter-intuitive finding

Here's what actually delivered profit over the same period — backing every single rank-1 pick, no filter at all:

Test window Bets Strike Elite ROI SP Favourite ROI Edge
Feb 25 → Mar 26 882 21.4% +36.77% −8.60% +45.4pp
Mar 27 → Apr 25 672 27.8% +53.31% −8.79% +62.1pp
Apr 26 → May 25 403 21.8% −4.12% −9.95% +5.8pp
May 26 → Jun 24 415 21.0% +29.44% −11.41% +40.9pp
Jun 25 → Jul 24 953 15.6% −20.32% −12.50% −7.8pp
Jul 25 → Jul 28 105 13.3% −10.15% −13.96% +3.8pp
AGGREGATE 3,430 20.8% +17.02% −10.51% +27.5pp

No filter. Just back every rank-1 pick. In this 155-day walk-forward test:

What that looks like on the equity curve

£-7,000 £-6,000 £-5,000 £-4,000 £-3,000 £-2,000 £-1,000 +£0 +£1,000 +£2,000 +£3,000 +£4,000 +£5,000 +£6,000 +£7,000 +£8,000 FebMarAprMayJunJul Cumulative Profit — Elite v4 rank-1 vs SP Favourite (£10 level stakes) Elite v4 (final: +£5838) SP Favourite (final: -£6395)

The dashed red line is what a punter blindly backing the SP favourite would have done: a slow bleed to −£6,395. The green line is Elite v4's rank-1 picks over the same days — trending up despite one obvious drawdown in late June/early July. A £12,232 swing between the two strategies over 141 trading days.

The one bad month (Jun 25 → Jul 24) lost 20% ROI. It happens. Every real signal has drawdowns. What matters is whether the strategy recovers — and it did.

An important caveat before you get excited

3,430 bets over 155 days is a much more reliable sample than the course-level slices we looked at earlier, but it isn't proof of a permanent edge. Markets evolve, models decay, and future performance may differ from what we've measured here. What this test shows is that the signal survived a large out-of-sample walk-forward evaluation and substantially outperformed a simple market benchmark over the period tested. That's the honest claim.

The edges most likely to survive going forward are structural ones — features of the model's underlying inputs (form, class movement, going, connections) that reflect genuine patterns in the sport — rather than recent performance streaks on specific tracks. That's what course-filtering was: a bet on recent performance streaks. That's why it failed.

Why filtering broke what worked

Small samples are mostly noise. A course with 15–30 bets can look like a "strong edge" or a "heavy loss" depending on which 5 races land differently. When we filter to "top 10 courses," we're picking the 10 courses that got lucky in the training window, not the 10 where the model has genuine skill. And courses that got lucky tend to regress.

The full model is where the skill lives. Elite v4's edge comes from its underlying probability estimates — how it weighs form, class movement, going preferences, trainer form, jockey booking, pedigree signals. Those signals apply across every course, in aggregate. Slicing the model into per-course subsets doesn't make it smarter. It just makes it noisier.

The parallel with investing is imperfect but instructive: a diversified index tends to beat a "top 10 stocks by last quarter's return" portfolio because recent winners regress and the underlying market average reasserts itself.


What this means for how you use RaceChatAI

Course-based cherry-picking underperformed. The temptation to say "our AI is best at Cheltenham" or "we should skip Wolverhampton" is real — but the walk-forward evidence points the other way.

The historical evidence suggests the model performs best when evaluated across all rank-1 selections, rather than after applying course-based filters. When RaceChatAI surfaces a rank-1 pick, back it or don't based on the price, not on which course it's at.

Respect variance. +17% ROI over 155 days doesn't mean +17% every week. Some weeks the picks fire; some weeks they don't. The edge is aggregate, over many bets, across a full season. Level-stake or fractional-Kelly it; don't chase or fade it based on the last three results. Past performance is not a promise of future performance.


Try it in the chat

Ask RaceChatAI:

"What's Elite's rank-1 pick in every race today?"

You'll get the full list — every course, every race, every rank-1 the model rates. That's the strategy the historical evidence supports. Not a curated subset. The whole card.


The honest bottom line

Three strategies tested. Two failed:

One retained a positive edge in this walk-forward evaluation:

That's the process. Most AI racing sites won't publish the failures. That's how their "systems" tend to be Cheltenham at +1,314% ROI in bold letters — until you check the base rate and realise you're being sold noise.

Our promise isn't that we'll pick every winner. It's that when we say something works, we've tested it in a way that would embarrass us if it didn't.

In our walk-forward testing, the full unfiltered rank-1 strategy was the only approach that consistently retained its edge out of sample.


Methodology — how the numbers were produced

For readers who want to reproduce or stress-test the claim:

If any of the above raises a question, get in touch — we'd rather answer a sceptic than have them silently distrust the numbers.


RaceChatAI — the AI horse racing analyst that shows its working. Chat live at racechatai.com.

This article describes historical model performance in a walk-forward evaluation. Racing is variance-heavy and no strategy guarantees future returns. Only bet what you can afford to lose.

backtestwalk-forwardtransparencymethodologyedgeelite-v4