20,934 Chart Patterns, Zero Passed
A point-in-time reconstruction of S&P 500 membership from 2000 to 2017, 20,934 detected chart patterns, three layers of benchmark, and a multiple-testing correction. Of 245 test cells with adequate sample size, zero survived. The scope is chart patterns only — RSI, MACD, stochastics and the rest were never tested. This article is about that zero, about an intermediate conclusion I had to retract, and about a question the study cannot answer: even if you do beat the index, is the excess worth what it costs you.
By Elnath Finance Academy
The chart pattern chapters of technical analysis textbooks make a set of unusually concrete claims: a head and shoulders top precedes a decline, a converging triangle precedes a breakout, a double bottom precedes a reversal. Those claims have one useful property — they can be tested.
This study reconstructed S&P 500 membership month by month across 2000 to 2017, detected 20,934 instances of nine classic pattern types across daily, weekly, and monthly bars for 674 stocks, split them into 354 test cells (one pattern type × one timeframe × one holding period × one benchmark), and tested whether any of them carries statistically detectable information about subsequent returns.
The result: of the 245 cells with adequate sample size, zero passed. The median hit rate — the share of events where the excess return was positive — was 50.45%.
But the zero is not what this article is really about. The zero is unsurprising; it is consistent with the weak-form efficient market hypothesis. What is worth writing down is the process. This study once produced the opposite conclusion: 16 significant cells that held up on data the analysis had never seen. It looked exactly like a real finding, and it was wrong. The error was not in the data, not in the pattern definitions, and not look-ahead bias. It was somewhere almost nobody thinks to check.
One thing to establish up front, so the next six sections are not misread: this is not an article arguing that technical analysis does not work. I use it myself, and there is a documented record of people profiting from it over long periods. Section 7 deals with that directly — why those cases and this study do not contradict each other, and what each can actually establish.
What this study tested, and what it did not
The scope needs stating up front, because "technical analysis" covers a great deal more than this study did:
| Coverage | |
|---|---|
| ✅ Tested | Nine geometric pattern types × three timeframes (daily, weekly, monthly) × five holding periods × three benchmarks = 354 test cells |
| ⚠️ Computed, but never a test subject | RSI, MACD, stochastics, Bollinger Bands, moving averages, ATR, ADX — all 24 indicator columns were computed and stored. In the pipeline they serve only as inputs to the market-regime labels and the ZigZag threshold. They were never defined as events and tested |
| ⚠️ Computed, but never a conditioning variable | Market regime (trend / range / transition, plus a volatility band). It was used to build the matched control group, but it is not a dimension of the test cells — so this article cannot answer "does the same head and shoulders behave differently in a bull market than a bear market" |
| ❌ Not attempted at all | Fibonacci retracements, Elliott wave, volume patterns, cross-asset and intermarket signals |
The conclusions here therefore cover chart patterns only. Reading them as "technical analysis does not work," or as "RSI and MACD are useless," goes well past what the data supports — those things were never tested.
Why not the second and third rows? The second is a design choice: the project's architecture document defines a test cell as pattern × timeframe × horizon, so indicators were positioned as downstream inputs from the start, not as subjects. The third is an oversight: the original research question explicitly read "after a pattern and market regime appear," and the test-cell definition dropped the regime dimension. The first was a decision; the second was a miss. Both are on the table here.
The missing pieces have since been filled
(Updated 25 August 2026.) The ⚠️ and ❌ rows above were a to-do list at the time. All three are done, written up as a follow-up: 49 Technical Signals, Zero Passed.
| Listed then as | Result in the follow-up |
|---|---|
| Market regime as a conditioning variable | Multi-timeframe trend expanded into 4 states, three combinations at 2,940 cells each — zero passed |
| Indicator signals (RSI / MACD / stochastics / Bollinger / moving averages) | 30 signal types plus 13 candlestick patterns, 2,205 cells — zero passed |
| Fibonacci retracements | 6 types, anchored on the already-frozen ZigZag pivots — zero passed |
That is 11,025 test cells over 1,779,160 events, none of which cleared the pre-registered threshold.
This section originally closed with:
I do not know what those results will be, and I am not going to guess. … Whether the answer comes back "yes" or "no," it gets written up the same way.
That promise was kept. And the most important part of the follow-up is not the zero — it is the 5 false discoveries I nearly published: they passed every check already in place, and they appeared because I brought an implementation into line with its specification. What caught them was the shuffled control, which produced 5 as well.
If you only quote one sentence
Quote this one, and do not shorten it further:
These nine chart pattern types, under this one frozen set of definitions, on S&P 500 large caps between 2000 and 2017, carry no detectable information once market drift, stock-specific characteristics, and same-day same-sector moves are controlled for.
Every shorter version claims more than the evidence. Two in particular are worth avoiding:
- "Chart patterns don't work." What was tested is one specific definition of nine specific patterns, not the family of chart patterns. A different definition, a different parameter set, or small caps could all give different answers. That is what Section 8, point one is about.
- "Chart patterns aren't useful." This study did not measure usefulness. It measured information content. There is no transaction cost model, no execution model, no position sizing, no risk management (limitation 5). "No detectable information" and "loses money when traded" are different statements, and this study only addresses the first — while the overnight gap finding in Section 5 suggests that even a pattern carrying information might not be reachable.
Two layers, at different strengths
- The finding is narrow. Nine geometric pattern types, after controlling for the market, for stock-specific characteristics, and for what the same sector did on the same day, carry no detectable information. That is all it says.
- The methodological lesson is general. What Sections 2 and 6 describe — how to set a benchmark, how to account for multiple testing, how a null distribution gets built wrong — applies to anyone trying to validate a method against historical data, whether that method is patterns, indicators, or something else entirely.
So how hard is it to beat the index with technical analysis? This study cannot answer that question — it tested one slice of it. What it can answer is a more basic one: it is very difficult to know whether you have beaten it. And that problem is not specific to patterns or indicators. Every method has to face it.
This is a research and methodology article. It describes how a test was designed, what it found, and where the boundaries of that finding lie. It is not a recommendation regarding any security or any method of analysis, and it is not investment advice.
1. What exactly was tested
Stated as a falsifiable proposition:
After a given chart pattern appears and is confirmed, does the distribution of subsequent returns differ detectably from the distribution when no pattern is present?
The project's original research question was broader — it read "after a pattern and market regime appear." Regime ended up being used only in the matched control group and never entered the test cells, so what this article actually answers is the narrower version above. That is the gap described in the third row of the scope table.
Nine pattern types, all defined by conventional charting rules. The program first reduces the price series to a sequence of turning points using ZigZag (a pivot-detection method whose threshold scales automatically with volatility via ATR, so small wiggles are not counted as turns), then matches geometry against the relative positions of those pivots. Counts detected between 2000 and 2017:
| Pattern | Count |
|---|---|
| Rising wedge | 5,999 |
| Double top | 3,520 |
| Double bottom | 3,415 |
| Falling wedge | 2,864 |
| Symmetric triangle | 2,297 |
| Head and shoulders top | 947 |
| Inverse head and shoulders | 852 |
| Ascending triangle | 739 |
| Descending triangle | 301 |
| Total | 20,934 |
Patterns were not picked by eye. They were marked by a program following fixed rules. Three head-and-shoulders instances:
The AVY chart is worth a second look. Textbook head and shoulders is a bearish pattern, and after confirmation the price went straight up. That is not an exception. That is the everyday content of the table in Section 3.
The data was cut into three parts, and only the first was used
This affects every section that follows, so it belongs here:
| Period | Purpose | Used by this study? |
|---|---|---|
| 2000–2017 | Development tier. All analysis, every attempt, happens here | ✅ Every result in this article comes from this segment |
| 2018–2021 | Validation tier. Used only to check whether something found in development holds up | ⚠️ Used once — see Section 6 |
| 2022 onward | Final holdout. Locked throughout; the project may open it exactly once | ❌ Never opened |
The split exists because searching for an answer and verifying an answer cannot happen on the same data. While searching, you inevitably try many combinations. To verify, you need data you genuinely did not touch while searching — otherwise you are only measuring your own footprints.
Three things frozen before any statistics were run
These are the most common failure points in this kind of research.
First, the usable timestamp is the bar on which the close broke the neckline — not the bar on which the shape completed. The right shoulder of a head and shoulders does not count until the neckline breaks; before that you are telling a story about a chart in hindsight. The confirmation timestamp also may not precede the moment every pivot was itself confirmed — deciding whether a high is a turning point requires later prices, so treating pivots as "known at the time" is the classic source of look-ahead bias in this field.
Second, every parameter was frozen before the statistics ran. Thresholds, tolerances, and window lengths were set by charting convention, written into a config file, and not touched again. "This pattern is detecting too few instances, let me loosen the tolerance and re-run" looks harmless and manufactures significant results directly.
This constraint also creates the study's deepest limitation, which I will flag here and develop in Section 7: freezing the definition is what makes a pattern testable, and the moment it is frozen it stops being the thing a human actually does. A hundred people read the same chart a hundred ways. A program reads it one way.
Third, what is measured is information content (does the pattern carry information), not tradability (can you make money on it). These are very different questions. A pattern could be statistically significant and still be unusable because of transaction costs, slippage, or the overnight gap problem in Section 5. The reverse also holds: "not tradable" does not imply "no information." This study answers only the first question.
2. Three things that make backtests look like they work
When a pattern backtest produces an attractive result, check these three things before believing it.
2.1 Survivorship bias: the sample contains only the winners that survived
Taking today's 503 constituents and running them back to 2000 is the most common error in this field. By definition, the companies on that list are the ones that lasted 26 years. Everything that went bankrupt, got acquired, or was dropped from the index is missing. Any pattern tested on that sample is standing on a stage that only goes up.
This study instead uses point-in-time membership — the list of companies actually in the index at each moment — reconstructed from the historical revisions of a Wikipedia article (one snapshot per month) combined with the index change log. Across the full history that is 985 distinct tickers, not 503.
That number deserves a pause on its own: the companies that entered and left the S&P 500 over 26 years come to nearly twice the number in the index at any single moment.
2.2 Benchmark: the market goes up anyway
Raw returns after a pattern appears are meaningless on their own, and it is worth showing exactly how meaningless. Within this sample — 2000 to 2017, S&P 500 constituents, total return, excluding days with obvious data errors — buying a randomly chosen constituent on a randomly chosen trading day and holding 20 trading days returns +1.08% on average, and is positive 57.4% of the time (1,358,968 observations).
None of that comes from patterns. That is the drift of the market over the period. Any claim of the form "this pattern is up 57% of the time" says nothing at all unless that layer has been subtracted first. (For completeness: the +1.08% itself carries the survivorship bias discussed in Section 8, so true drift was somewhat lower.)
So the study uses three benchmark layers, each stricter than the last:
| Benchmark | Calculation | What it removes |
|---|---|---|
| Market-adjusted | Raw return − SPY return over the same window (SPY tracks the S&P 500) | Market drift |
| Unconditional | Raw return − that stock's own full-history mean | Stock-specific characteristics |
| Matched control | Raw return − mean of a control group on the same date, in the same market regime, in the same sector | Date + regime + sector rotation |
The matched control is the layer that matters. If a pattern merely happens to show up more often in technology stocks during bull years, this layer removes that effect entirely — it compares against stocks that were in the same sector, on the same day, in the same market regime, and did not produce a pattern.
2.3 Multiple testing: 354 tests were run
Pattern × timeframe × horizon × benchmark comes to 354 test cells. A cell is one specific question — for example: "After a double top confirms on daily bars, held for 20 trading days, measured against the matched control, is the excess return different from zero?" Of those, 245 have adequate sample size (at least 30 events).
In a world with no effect at all, running 245 tests at a 5% significance level produces roughly 12 "significant" cells by luck alone.
This is the most persuasive set of numbers in the study. Same data, three levels of rigor:
| How the question is asked | "Significant" cells |
|---|---|
| 95% confidence interval on mean excess return excludes 0 | 32 / 245 |
| Tested against the correct null distribution, uncorrected p < 0.05 | 13 / 245 (12.2 expected by luck) |
| After Benjamini–Hochberg FDR correction (q = 0.10) | 0 / 245 |
The first row is what the large majority of pattern backtests actually do: subtract the market from subsequent returns and test whether the remainder differs from zero. Doing that yields 32 "findings."
The second row fixes a hidden error in the first — after subtracting the market, the no-effect case does not sit at zero either. Patterns do not appear randomly on the calendar; they cluster in particular market conditions. Draw a set of dates with the same time distribution at random and the expected excess return is not zero. So the study does not test against 0. It constructs the null distribution (what the number looks like when patterns carry no information at all) using a randomization test: shift every pattern date by the same offset, recompute thousands of times, and see where the real value falls among those recomputations. Fixing this drops 32 cells to 13.
The third row applies the multiple-testing correction — Benjamini–Hochberg, which controls the FDR (false discovery rate: among the cells you declare to be findings, what share are actually false), at a threshold of q = 0.10. Thirteen is almost exactly the 12.2 expected from luck alone, and after correction nothing is left. The smallest q-value is 0.780, nowhere near the 0.10 threshold.
There is a version of this that traders already say, and Section 7 returns to it: making money once with a hundred different setups, and making money a hundred times with one setup, are not the same thing statistically.
3. Results
Using the strictest benchmark (matched control), daily bars, 20-trading-day holding period (about one month). Effect size is the excess return left after subtracting both the benchmark and the null level; hit rate is the share of events where that excess was positive; a q-value near 1 means the cell is very unlikely to be a real finding (the threshold is 0.10):
| Pattern | Events | Effect size | Hit rate | 95% CI | q |
|---|---|---|---|---|---|
| Head and shoulders top | 532 | +0.01% | 48.5% | −0.51% to +0.46% | 0.99 |
| Inverse head and shoulders | 497 | −0.33% | 50.9% | −0.85% to +0.15% | 0.96 |
| Double top | 1,929 | −0.17% | 48.6% | −0.49% to +0.12% | 0.96 |
| Double bottom | 1,916 | −0.13% | 48.9% | −0.38% to +0.17% | 0.96 |
| Rising wedge | 3,150 | −0.02% | 50.9% | −0.22% to +0.33% | 0.98 |
| Falling wedge | 1,340 | +0.51% | 46.7% | −0.55% to +1.25% | 0.93 |
| Ascending triangle | 443 | +0.16% | 54.2% | −0.26% to +0.67% | 0.96 |
| Descending triangle | 173 | +0.93% | 53.2% | −0.35% to +1.94% | 0.91 |
| Symmetric triangle | 1,055 | +0.37% | 49.9% | −0.13% to +0.90% | 0.96 |
All nine confidence intervals straddle zero — meaning the true effect in every row could be zero, could be positive, could be negative, and the data cannot tell them apart.
Across all 245 reliable cells:
| Metric | Value |
|---|---|
| Median |effect size| | 0.26% |
| 90th percentile |effect size| | 1.83% |
| Median hit rate | 50.45% |
| Cells with hit rate between 45% and 55% | 198 / 245 (81%) |
| Smallest uncorrected p-value | 0.0055 |
| Smallest q-value | 0.780 |
A median hit rate of 50.45%. What matters is not that it is close to half, but that it is close to half to an almost embarrassing degree: across 245 cells, four fifths of the hit rates fall within five percentage points of a coin flip.
The directions are not tidy either. Head and shoulders — conventionally a bearish pattern — shows an effect size of +0.01% here, while double bottom, conventionally bullish, shows −0.13%. This is not evidence that patterns work in reverse. Both numbers are inside the noise. It simply illustrates that when effect sizes are this small, the sign itself carries no meaning.
4. Statistical power: the zero is not blindness
A test that cannot detect anything and a world with nothing to detect produce identical reports. The two have to be separated.
The way to separate them is a power analysis — the probability that this test detects an effect that genuinely exists. Take the real event structure (double top, n = 3,147), inject an effect of known size, and see whether the test finds it:
| Injected effect | Detection rate |
|---|---|
| 0.5% | 83% |
| 1.0% | 100% |
| 2.0% | 100% |
Under the null (no effect present), the false positive rate is 0% against a nominal 5% — so the test is, if anything, conservative.
A 1% effect, if it existed, would be detected 100% of the time. The effect sizes actually observed have a median of 0.26% and a 90th percentile of 1.83%.
Put differently: the telescope is not too dim. There is no star at that position — at least none as large as 1%.
5. Patterns mark news days; they do not predict news
This is the most concrete mechanical finding in the study, and in my view the most useful section.
Split each day's price move into two parts:
The overnight gap is open ÷ prior close − 1, covering every piece of news between the close and the next open. The intraday move is close ÷ open − 1.
The first result: the overnight segment accounts for 60.6% of total daily variance. More than half of all price movement happens while the market is shut.
The second result matters more. Pattern confirmation days are disproportionately days with large gaps:
| Metric | Confirmation days | All days | Ratio |
|---|---|---|---|
| Mean absolute gap | 1.29% | 0.69% | 1.87× |
| Share with gap above 2% | 15.57% | 5.91% | 2.63× |
| Share with gap above 5% | 5.53% | 0.96% | 5.76× |
| Share falling on market-wide gap days | 1.09% | 0.40% | 2.70× |
And for 22.3% of confirmations, the overnight segment alone accounts for more than half of that day's total move.
"Market-wide gap days" are identified by cross-sectional co-movement: if 20% or more of constituents gap in the same direction on the same day, it is classified as a macro or systemic event. The 27 days identified this way all correspond to known events — the March 2020 COVID crash and rebound, the October 2008 financial crisis, the August 2015 renminbi devaluation flash crash, the 9 November 2020 Pfizer vaccine announcement. Not one is noise, which means the classification method validates itself.
Put together, the picture is clear:
Patterns did not predict that news. The price moves caused by that news triggered the pattern confirmations. The neckline breaks, very often, because an earnings report or macro release pushed the price straight through it before the open — there was no intraday "break" at all. This also explains why the matched control removes the effect completely: once you subtract what every stock in the same sector, on the same day, in the same regime did, the contribution attributable to the pattern itself is gone.
Incidentally, this matters more for tradability than for information content. A substantial share of confirmations occur inside gaps, which means that even a statistically significant pattern would leave nothing to react to during the session — the price is already there before the market opens.
6. The intermediate conclusion I had to retract
This section is about the place where the study got something wrong. It is worth more than every result above it combined — because the error was not carelessness. It was a structural trap: it left every number on the report looking normal, and the check specifically designed to catch errors passed.
Here is what the first version of the analysis produced: 16 significant cells in the development tier (2000–2017); carried into the validation tier (2018–2021, data untouched during the search) the direction agreed on 15 of 16, and 5 cells replicated.
That looks exactly like a real finding. Modest, undramatic, directionally stable, and it survived out of sample. If I had to design a report capable of fooling me, it would look like that.
All of it was an artifact.
The problem was in the construction of the null distribution from Section 2.3.
Pattern confirmations are heavily concentrated on particular days — statistically, they cluster. In the development tier, 18,441 events fall across 3,836 trading days: a median of 3 per day, a maximum of 63, with the busiest 1% of days accounting for 7.6% of all events. The reason is not mysterious. On a day the whole market drops hard, hundreds of stocks break their necklines at once. Events sharing a day share a market shock, and are therefore highly correlated with each other.
The original randomization drew an independent new date for each event. That destroys the clustering: in the resulting null world, events are spread evenly across the calendar and independent of one another, so the averaged noise is far smaller than in the real world.
Measured directly (20-day horizon, identical event set):
| Date reassignment method | Null distribution SD |
|---|---|
| Independent redraw per event (old) | 0.000667 |
| Uniform shift of all events (new) | 0.001754 |
The null distribution was 2.63× too narrow. The consequence is direct: an effect that is genuinely 1σ — one standard deviation, entirely within noise — gets scored as 2.6σ, which looks like a discovery. And because the bias is systematic rather than random, the wrong results came out looking tidy and convincing.
The fix is a uniform shift: every event moves by the same offset along its own stock's bar sequence. This preserves the event count, the spacing between events within a stock, and the same-day clustering across stocks. The only thing destroyed is the alignment between events and real market dates — which is precisely what the null hypothesis is supposed to destroy.
After the fix, all three benchmarks went to zero.
The part worth remembering
Throughout, the study ran a null-hypothesis smoke test: scramble the event labels and confirm the test no longer reports significance. If it still finds something after scrambling, the pipeline is broken.
That smoke test passed under both versions.
The reason: the scrambled data and the null distribution it was compared against were built by the same logic. Both were wrong in the same way, so the two errors cancelled. A smoke test catches "does the pipeline mistake pure noise for a finding" — but it cannot catch an error in the construction of the null distribution itself, because the flaw is present in the ruler and the measured object simultaneously.
What finally caught it was not a test at all. It was a blunter move: measure the width of the null distribution directly, and compare it against a version that preserves the clustering.
If one sentence survives from this article, I would like it to be this one: all the lights being green does not mean nothing is wrong. It can also mean every light is wired to the same circuit.
7. So why do people still make money with technical analysis
If the six sections above read as "technical analysis does not work," I wrote them badly. I use technical analysis myself. And there is a group of people who have profited from it over long periods — this is not in dispute, and there is a public record.
Several cases, at differing levels of verifiability:
| Case | What it was | Verifiability |
|---|---|---|
| The Turtle experiment (1983–84) | Richard Dennis and William Eckhardt bet on whether trading could be taught, recruited novices, and taught them a fully mechanical breakout system | High: the rules were later published by participants and documented in several books. Individual results vary by source |
| U.S. Investing Championship | A real-money trading competition running since the 1980s; winners include Marty Schwartz, David Ryan, and Mark Minervini, all openly technical traders | Medium-high: the competition and placings are public; return figures come mostly from the organizer and press |
| Paul Tudor Jones and 1987 | Heavy use of chart analysis; positioned correctly ahead of the 1987 crash | Medium: primarily documented in Market Wizards and the 1987 documentary Trader |
| Systematic trend-following funds (Man AHL, Winton, Millburn and others) | Rule-based strategies driven by price data, with decades of regulated performance records | Highest: regulated, audited, regularly reported |
The table describes what each case is and how far it can be verified. It deliberately quotes no return figures — most are not independently audited, and every retelling adds a layer of distortion. The primary sources are below; go read them yourself.
- The Turtle experiment — the full rules were released free by participant Curtis Faith with Richard Dennis's permission; the official site is originalturtles.org, and the stated reason for giving them away is written into the document: to undercut the people selling trading systems and seminars. See also Curtis Faith, Way of the Turtle (McGraw-Hill, 2007) and Michael Covel, The Complete TurtleTrader (HarperBusiness, 2007).
- U.S. Investing Championship — organizer's site: financial-competitions.com, run by Norman Zadeh; the contest began in 1983, lapsed, and was restarted in 2019. Entrants link a checkable brokerage account and are ranked on percentage return — that is verification by the organizer, not an independent audit, and the difference matters. Press release for Minervini's 2021 win: PR Newswire.
- Paul Tudor Jones — Jack D. Schwager, Market Wizards (New York Institute of Finance, 1989; lendable copy at the Internet Archive); the documentary Trader (PBS, 1987, directed by Michael Glyn, IMDb). Jones later had the film pulled from circulation, so the versions circulating online are unofficial transfers — worth knowing when citing it.
- Systematic trend-following funds — Man AHL (founded 1987) and Winton (founded 1997 by David Harding). Both are institutional systematic managers with decades of public record, which is why the table rates them the most verifiable of the four.
Setting these cases beside this study yields four conclusions, none of which conflict.
One: what they do is not what this study measured
Look at the rightmost column. The most verifiable category is rule-based trend following, not chart patterns. The Turtle rules were "enter on a breakout of the N-day high, size the position by volatility, exit on a reverse break." There is no head and shoulders in there, no neckline, no wedge.
And the most-discussed individual traders do not rely on a single pattern either. Their approach binds growth criteria, relative strength, base structure, and strict risk control together; the pattern is the thing that pulls the trigger, not the source of the edge.
This is precisely where the study's scope limit bites. I measured whether the geometric shape carries information on its own. The answer is that it does not. That is a different question from whether a pattern is useful as a trigger inside a complete process — and I did not test the second one.
Two: the names you know were selected out of a very large pool
I want to state this precisely, because it is easy to turn into "those people just got lucky," which is not what I mean.
Return to Section 2.3: across 245 cells, luck alone produces about 12 that look significant. Competitions and hall-of-fame lists are the same structure in the real world. Each competition has hundreds of entrants; what gets reported is the top line of the results table. Nobody tabulates the distribution across all those entrants, and few follow up on what happened to last year's winner.
This does not mean the winners lack skill. It means that looking only at the ones selected out, you cannot tell how much is skill and how much is the tail of a large distribution. Both readings fit the data, and the truth is probably some of each.
Separating them requires exactly what this study required: rules fixed in advance, a complete record of every attempt (not just the successful ones), and a control group. Which is why a published trading journal is more persuasive than any competition win — the first records everything, the second records only the tail.
Three: a hundred people read the same chart a hundred ways
What makes technical analysis genuinely interesting is that it is not a fixed rule set. Two people can look at the same daily chart and one sees a base about to complete while the other sees a downtrend with further to fall. Switch to weekly and it is a different story again; monthly, another. Each person's version is something they ground out over years, and it does not necessarily transfer.
This is fatal for the study, and I should say so plainly. To make patterns testable at all, I had to freeze the definition into a fixed set of parameters (Section 1). Freezing makes it measurable — and at the moment of freezing, it stops being the thing a live human does.
So the honest statement is: this study can confirm or refute a fixed, writable pattern rule. It can neither confirm nor refute the accumulated, adaptive judgment of a person with years of screen time. The latter is not something I measured with the wrong method; it is not measurable this way in principle, because it is never exactly the same twice.
That is not a defense of technical analysis. It is a double-edged sword: untestable also means unverifiable. You cannot demonstrate to anyone else that it works, and no one can demonstrate to you that it doesn't — yourself included.
Four: one setup a hundred times beats a hundred setups once each
There is a line that circulates among traders: making money a hundred times with one setup beats making money a hundred times with a hundred setups.
It usually gets passed along as a piece of discipline. But it is exactly the statistical problem from Section 2.3, phrased differently:
- A hundred setups, one win each = a hundred tests. Some of them are pure luck, and afterwards you cannot tell which.
- One setup, a hundred wins = one test with n = 100. That is statistically identifiable — and more importantly, you yourself can find out whether it works.
FDR correction is the accounting for the first case. And the real reason "trade one setup" holds up in practice may not be that the setup is unusually strong. It is that it gives you enough repetitions to discover that the setup does not work.
Someone running a hundred different setups never finds that out. They only remember the ones that worked.
8. What this does and does not establish
Three levels, in decreasing strength.
One: this is not proof that technical analysis does not work — it does not even establish that chart patterns as a family do not work. What was measured is nine specific geometric patterns, one frozen parameter set, on S&P 500 large caps, across three timeframes. Different parameters, a different pattern definition, small caps, or another asset class could all give different answers. Indicator-based methods (RSI, MACD, stochastics, Bollinger Bands, moving averages) and Fibonacci were never tested at all, and the regime dimension never entered the test cells. Generalizing this result to technical analysis as a whole goes well past what the data supports — the scope table above and all of Section 7 are about that point.
Two: it is consistent with the weak-form efficient market hypothesis. If these patterns carried detectable information, then with 20,934 instances and 100% power against a 1% effect, it should have been visible. It was not, and the simplest explanation is that the information is already in the price.
Three: the most concrete and best-supported result is the mechanism in Section 5. Confirmation days are 5.76× more likely to carry a large gap and 2.70× more likely to land on a macro event day than an ordinary trading day. Patterns mark news; they do not predict it.
Problems with the study itself
Listed honestly, each with the direction of its bias — whether it makes the result look better or worse than reality:
-
Scope: only chart patterns were tested, and the regime dimension is missing. This is more fundamental than any of the data problems below. Indicator-based methods were not tested at all, and market regime was computed but never entered a test cell. This is not a bias, it is a gap — it does not push the numbers up or down, but it makes the conclusions apply to a far narrower domain than a reader might assume at first glance.
-
Membership before 2007 is incomplete. Only 447 names were reconstructed for 2000 against roughly 500 in reality. What is missing are companies that were in the index, left before 2007, and whose departure was not logged — mostly losers. This direction works against the study: it makes early-sample patterns look more effective than they were.
-
Price data is missing for delisted stocks. Only 674 of the 985 tickers (68%) have usable price data. So this version of the result still carries survivorship bias, again in the direction of making patterns look better.
-
Pattern definitions are subjective. The geometric rules are one reasonable definition among many. Parameters were frozen before analysis, but different parameters would give different results (see Section 7, part three).
-
No transaction cost model. What was measured is information content, not tradability.
-
S&P 500 large caps only. Not extendable to small caps, non-US equities, or other asset classes.
-
Vendors restate historical data. The 2010 data downloaded today is not the 2010 data that was visible in 2010. This layer cannot be eliminated.
Points 2 and 3 deserve a note of their own: the two largest data biases both push in the direction of making patterns look more effective. Under those conditions, the result was still zero.
The final holdout remains unopened
The third segment in Section 1 — data from 2022 onward — is locked for the duration, openable exactly once, enforced in code.
Zero significant cells in development means there is no candidate to carry forward to a final test. So that single opportunity was not spent, and the data is still locked. This is part of the design: if "we found nothing, so let's run the whole set against the holdout and see" were permitted, the holdout would not exist in any meaningful sense.
There is a detail here that is easy to misread, so it is worth stating precisely. Pattern detection itself ran across the entire history, including 2022 onward — there are pattern records from the holdout period sitting on disk. That does not contradict the claim that the holdout was never opened; the distinction is which layer:
- Pattern detection is a deterministic function of price and never touches forward returns. Drawing a head and shoulders on 2023 prices reveals nothing about whether it went up or down afterwards.
- What is locked is the statistics layer. Returns, benchmarks, and tests only ever see the development tier. Holdout-period results were never computed at all.
That is not the real protection, though. Even with the returns uncomputed, a human who has looked at holdout-period charts can let that feed back into parameter choices — the hardest kind of leakage to prevent, and the hardest to notice in yourself. What actually blocks it is the second point in Section 1: every parameter in the pattern definitions was frozen before any statistics ran. Without that rule, a holdout preserves numbers but not judgment.
9. The hard part is not the market. It is the measurement — and whether measuring is worth it
Back to the opening question: how hard is it to beat the index with technical analysis?
This study cannot answer the whole question — it tested nine geometric pattern types, with indicators and market regime outside its scope (see the scope table in the introduction). It cannot put a number on "how hard." What it does offer is two more practical answers, neither of which is specific to patterns.
One: your judgment about whether you are winning is worse than you think
Same data, same patterns, same person. The first run produced 16 significant cells that held direction 15 of 16 on data it had never seen. After fixing one error buried in the null distribution, it produced zero. The difference was not data, not honesty, and not effort. It was a technical detail invisible on the report and undetectable by the smoke test built to catch exactly that class of problem.
And that happened under the most favorable conditions available: complete data, parameters frozen in advance, a full inventory of all 245 cells, and a machine that could re-run the whole thing thousands of times.
The real world is harder to measure, not easier. There is no complete inventory to apply an FDR correction to — you only remember which of the ideas you tried seemed to work. There are no frozen parameters, only "let me adjust this a little." There is no control group, only "I made money that quarter." Every trade pays transaction costs. And Section 5 still applies: more than half of all price movement happens when you cannot act.
Two: even if you measure correctly and genuinely win, there is a question this study cannot answer
The ongoing effort cost of an index position is close to zero. That is true by definition, not an empirical finding. So the bar any active approach has to clear is not just the index return — it is the index return plus the time, energy, and attention the approach consumes.
The real question, then, is not whether excess return is achievable. It is, and there is a public record (Section 7). The real question is:
After subtracting the time, the energy, and the opportunity cost you paid for it, how much of that excess return is left? And is what remains worth what you paid?
I will not answer that for anyone, and this article does not attempt to. The answer depends on how much time you have, what that time is worth spent elsewhere, and how much you enjoy the activity itself — three variables only you can fill in. Some people run the numbers and find it clearly worth it; others find it clearly is not. Both can be right, because they substituted different values.
There is only one thing I can contribute: making the question "have I actually beaten it" somewhat more measurable than it was.
I wrote this study up in full, including the intermediate conclusion I had to retract, because in this field the only thing more valuable than finding a method that works is a process that can tell you when you haven't.
The first may not exist. The second you can build yourself.
Data sources
The data this study itself relied on, and what is wrong with each:
| Data | Source | Caveat |
|---|---|---|
| S&P 500 membership | The historical revisions of the Wikipedia article List of S&P 500 companies — one snapshot per month from July 2007 | Before 2007 membership can only be inferred backwards from the change log, and coverage is incomplete — the gap described in Section 8, point 2 |
| Index change log | Wikipedia, Historical components of the S&P 500 | Community-maintained and unofficial; early entries have gaps |
| Daily OHLCV | Yahoo Finance, via yfinance | Excludes delisted stocks, which is the main source of survivorship bias in this study (Section 8, point 3) |
One thing should be stated plainly: the official membership history is proprietary data owned by S&P Dow Jones Indices and requires a license. This study reconstructs it from public sources, and the gaps listed in Section 8 are the price of doing so. That is not an oversight — it is a known trade-off at this budget, and the direction of each gap has been quantified and written into the limitations.
Study window 1 January 2000 to 31 December 2017. Universe is point-in-time S&P 500 membership; daily, weekly, and monthly bars. Every figure is recomputed by the analysis pipeline from raw result files rather than transcribed by hand. Pattern charts show instances from within the study window.
Disclaimer
Content on this site is produced by Elnath Finance Academy for general informational and educational purposes only. It is not investment advice and is not a personalized recommendation for any individual reader. Elnath Finance Academy is not a registered investment adviser (RIA) and does not provide regulated advisory services. Data and analysis may be delayed or contain errors; past performance does not guarantee future results. Investing involves risk, including possible loss of principal. Make your own decisions and consult a qualified professional.