The Tested Desk
Would this strategy have passed a prop-firm evaluation? I run the numbers.

Eval verdicts: 1 scored · 1 rejected · 0 passed · Death reports: 1
REJECT

The doji tells you nothing: a death report

2026-09-19·PS-DOJI·holdout sealed, never opened
Backtested · negative result

Simulated, backtested, and not a funded-account result. No account was traded. This is a negative result — a documented failure, published because the failure is the finding. There is no strategy here to adopt and no P&L series to show, because this line died before it ever had one.

The question

A doji is a candle that opens and closes at nearly the same price. The folklore says it marks indecision, and that indecision resolves into something you can trade. The question put to it was narrow and answerable: does a doji carry information about the next move — its direction, its magnitude, or its dispersion?

Not "is the doji real". Whether the thing it is supposed to tell you is large enough to be worth acting on after costs.

What was tested

Bars 5-minute, resampled from 1-minute
Universe SPY, QQQ, IWM, AAPL, AMZN, META, MSFT, NVDA, TSLA
Discovery window 2003-09-10 → 2018-12-12
Sample 11,881,916 bars across 3,842 day-blocks
Holdout sealed 2018-12-13, never opened

The holdout matters more than it looks. It was sealed before the discovery work began and it is still sealed, because nothing in discovery earned the right to spend it. A shortlist of candidates would have opened it. The shortlist came back empty, so the holdout is untouched and remains available to a future question that deserves it.

The doji definition was banded rather than fixed, so the result could not hinge on one arbitrary threshold: body-to-range ratios of 0.02, 0.05 and 0.10, with wick balance bands at 0.70 and 0.30. Those are specification, not discovery — any reader can reimplement them exactly.

The decision rule, registered before the run

Results: zero separations

Four of the twenty cleared BY at q = 0.10. All four landed below the cost floor, so under the registered rule none was eligible for the shortlist, and the shortlist closed empty.

contrast effect × cost floor 95% CI p state
dispersion vs controls +2.9% 0.29× [+2.5%, +3.2%] 0.0001 sub-threshold
|forward move| vs controls +2.2% 0.22× [+1.1%, +3.3%] 0.0001 sub-threshold
opex week vs not −1.55 bps 0.16× [−2.62, −0.48] 0.0040 withdrawn — artifact
above vs below session VWAP +1.39 bps 0.14× [+0.93, +1.86] 0.0001 withdrawn — artifact

The other sixteen showed no separation at all.

The two withdrawals

Two of those four were withdrawn after they had already cleared BY. Not because the numbers were wrong — because a noise control reproduced them.

The VWAP contrast reported that dojis above the session VWAP outperform dojis below it: +1.39 bps at p = 0.0001. Then the same contrast was run on populations that were not dojis at all:

population effect (bps) 95% CI p
real dojis +1.394 [+0.93, +1.86] 0.0001
noise control 5 +1.259 [+0.76, +1.74] 0.0001
noise control 1 +1.244 [+0.76, +1.72] 0.0001
noise control 3 +1.121 [+0.64, +1.60] 0.0001
noise control 4 +1.083 [+0.60, +1.57] 0.0001
noise control 6 +0.760 [+0.35, +1.18] 0.0004
noise control 2 +0.655 [+0.17, +1.13] 0.0110

All six reproduced it. Four at the same p = 0.0001, at 78–90% of the real magnitude, with intervals that overlap the real one almost entirely. The contrast measured nothing about dojis. What it measured is that bars above session VWAP have higher forward returns than bars below it, for any base-rate-matched selection of bars — intraday drift, and a property of the conditioning variable. The opex-week contrast fell the same way, at −1.55 bps real against controls at −1.18 and −1.16, all six sharing direction.

The two that survived are not survivors

The dispersion result, +2.9%, is implementation-unstable. Four defensible implementations of the same pre-registered contrast returned a range from −15.9% to +21.1%. A finding that changes sign depending on which reasonable estimator you pick is not a finding; it is a description of the estimator.

The absolute-move result, +2.2%, is era-unstable: +1.6% in the early era against +4.7% in the late one. One regime wearing two faces.

Three implementation errors were caught by cross-implementation checks during the block and corrected before anything was reported. Two of them moved results enough to change conclusions. That is the argument for computing a contrast more than one way even when the first way looks fine.

The two findings worth your time

1. The confirmation condition is not a condition

A common rule says: take the doji, but only if the next k bars confirm it. Tested at k = 3 on 5-minute bars, 151,957 of 157,033 dojis are "confirmed" — 96.8%.

The filter excludes about three percent of the population. It is not selecting a better subset; it is passing almost everything and charging you three bars of delay for the privilege. Any rule of the form pattern plus confirmation should be asked, first, what fraction it actually rejects. If the answer is a few percent, the confirmation is decoration.

2. A noise control must replicate the contrast structure, not just the pattern

This is the expensive lesson, and it cost two results that had already survived multiple-comparison correction.

The original noise controls were built correctly as far as they went: take a non-doji population matched on the same base rates, and check that the pattern effect disappears. It did. Every control returned null. Their nullity was reassuring and, it turned out, meaningless.

The controls varied the pattern. The contrast was conditioned on a variable — position relative to VWAP. A control that varies only the pattern tests whether the pattern matters and is silent on whether the conditioning does. The fix is a rule that is easy to state and easy to forget:

For any contrast conditioned on a variable V, the noise control is that same contrast on V, applied to a non-pattern population.

Splitting the controls by the same conditioning variable is what exposed both artifacts. Neither would have been caught by a control that only asked whether dojis were special.

What this does and does not say

It does say that across twenty pre-registered contrasts on nine liquid names over fifteen years of 5-minute bars, no doji contrast produced a separation large enough to cover a 10 bp cost floor — and that the two largest apparent effects were properties of the conditioning variables rather than of the pattern.

It does not say that dojis are meaningless everywhere. Specifically:

Desk verdict card

○ REJECTFailed the pre-registered gate. Recorded, not deleted — a rejected run is evidence.

Run PS-DOJI
Logged 2026-08-08
Data window 2003-09-10 → 2018-12-12

Headline metrics

metric value
pre-registered contrasts 20
separations found 0
cleared BY at q = 0.10 4
best of those, vs cost floor 0.29×
withdrawn as artifacts 2
bars / day-blocks 11,881,916 / 3,842
shortlist empty

Flags


Want to see what happens when a strategy does have a P&L series? → Read post 1.