Research · Empirical finance · Reproducible
The Ghost of an Anomaly
Turn-of-the-Month Anomaly Decay · Reproducible Research
An anomaly can still look compelling in a long backtest even when much of its premium belongs to a market that no longer exists.
I started with a simple question: if the turn-of-the-month effect has been documented for decades, does it still belong to today's market?
The interesting part was not confirming that it once existed. That was already known. The interesting part was that becoming public does not coincide with an immediate collapse, while the large deterioration appears later.
The question
How much of this result still belongs to today's market?
The turn-of-the-month effect is one of the most documented calendar anomalies in U.S. equities (Lakonishok and Smidt, 1988; McConnell and Xu, 2008).
I was not interested in proving again that it worked historically. I wanted to know something more useful: how much of that historical result still belongs to the market that exists today.
I split the sample through time, compared canonical TOM days with every other trading day, and repeated the analysis using a second data source and a different market universe.
What I found
Four readings of the same arc
- Pre-1987 TOM premium in the S&P 500.
12.74bps/day
Pre-1987 TOM premium in the S&P 500.
- Between the publication era and the pre-decimalization period.
11.66bps/day
Between the publication era and the pre-decimalization period.
- After 2001 and before T+2.
2.72bps/day
After 2001 and before T+2.
- Modern 10-year rolling HAC 95% intervals include zero.
≈ 0
Modern 10-year rolling HAC 95% intervals include zero.
The anomaly survived becoming well known. Its magnitude did not.
The independent Kenneth French U.S. market replication shows the same central pattern.
First twist
Becoming public did not appear to kill it
If the story were simply “the anomaly was published and arbitrage removed it,” I should see a clear break around 1987 (Ariel, 1987).
I do not.
In Yahoo/S&P 500 data, the premium moves from 12.74 to 11.66 bps per day. In the matched 1950+ French replication, it moves from 12.58 to 14.43 bps.
Direct tests between those regimes do not detect a significant difference.
Per-source comparison of the TOM premium before 1987 and between 1987 and 2001, with the p-value of the direct change between the two regimes.
No immediate collapse detected
Research note
Failing to detect a break does not prove that publication or arbitrage had no effect. It only means that the simple story of an immediate disappearance does not fit these data particularly well.
Main evidence
The decay came later
After 2001, the premium falls to 2.72 bps/day in the S&P 500 and 2.97 bps/day in the matched French replication.
In later regimes the estimate continues to compress and is no longer statistically distinguishable from the rest of the market.
Several microstructure changes overlap — settlement, decimalization, electronic trading, costs, and liquidity. I do not identify any single one as the cause.
What the data do show is a large deterioration during that broader transformation of the market.
Rolling 10-year TOM premium
Fixed 10-year trailing window, with HAC 95% interval
The rolling 10-year TOM premium starts above 15 bps/day in the mid-twentieth century, stays wide through the publication era, and compresses toward zero after 2001. In recent windows the HAC 95% interval includes zero in both sources.
The vertical marks flag changes in market structure. They are time references, not identified causes.
Independent replication
Same pattern, independent data
The first sample uses the S&P 500 from Yahoo Finance.
The second reconstructs the U.S. market return from Kenneth French's Mkt-RF + RF, using a CRSP-based value-weighted market universe.
They are not the same series, and I do not treat them as if they were.
That is exactly why the comparison is useful: change the provider and the market universe, and the central decay pattern remains.
Yahoo S&P 500
- Source
- Yahoo Finance via yfinance
- Universe
- S&P 500 index (^GSPC)
- Return
- Provider adjusted-close percentage change
Kenneth French US Market
- Source
- Kenneth French Data Library
- Universe
- CRSP-based value-weighted U.S. market
- Return
- (Mkt-RF + RF) / 100
TOM premium by regime and source
Daily TOM premium by publication regime in both sources. The first two regimes are wide and similar; later regimes fall to a few bps per day.
T+1 is a short sample. I do not draw a strong modern negative-alpha conclusion from it.
| Regime | Yahoo S&P 500 | French US Market |
|---|---|---|
| Pre-publication | 12.74p (HAC) < 0.0001 | 12.58p (HAC) < 0.0001 |
| Published / pre-decimalization | 11.66p (HAC) 0.0082 | 14.43p (HAC) 0.0010 |
| Post-decimalization / pre-T+2 | 2.72p (HAC) 0.5296 | 2.97p (HAC) 0.4986 |
| T+2 | 2.22p (HAC) 0.7633 | 1.84p (HAC) 0.8103 |
| T+1Short sample | 1.21p (HAC) 0.9027 | -4.19p (HAC) 0.6924 |
A different provider, a different market universe, the same central decay.
Ghost alpha
What a long backtest can hide
A longer sample often feels safer.
But length and stability are not the same thing.
If I combine decades in which the premium was large with decades in which the estimate fluctuates around zero, the historical average still carries part of the old signal.
The backtest is not necessarily wrong. It may simply be answering a question that is no longer the one I care about.
The question I do care about is:
How much of the historical result still exists inside the regime I would actually trade?
Layers of evidence
1950–1987
A wide and statistically clear premium.
1987–2001
The anomaly is public and the premium is still wide.
2001–2017
The estimate compresses to a few bps per day.
Recent windows
The 95% intervals include zero.
A historical average can survive much longer than the phenomenon that created it.
Negative results
I tested the mechanism. Not everything survived.
I also tested several pieces that could help explain the pattern.
Result
Prior pressure → reversal
The S&P 500 result is consistent with the proposed mechanism. In the matched 1950+ French sample, the difference is only marginally significant.
Verdict
Suggestive, not robust enough to headline
Result
Quarter-end / semi-year concentration
The differences do not survive robustly in the matched-sample comparison.
- French US Market
- All pairwise p-values > 0.20
Verdict
Not robust
Result
Automatic breakpoint
The exploratory algorithm selects 1964 in Yahoo and 1991 in French. The disagreement is informative: no single robust date is selected across universes.
- Yahoo S&P 500
- Selected year 1964
- French US Market
- Selected year 1991
Verdict
Exploratory; no single robust date
I did not remove these results to make the story cleaner. They stay because they are part of the research.
Claim boundary
What the evidence supports — and what it doesn't
The evidence supports
- a large historical TOM premium;
- persistence around the publication era;
- substantial later decay;
- modern rolling estimates statistically indistinguishable from zero;
- independent replication of the central decay pattern.
The evidence does not establish
- decimalization as the cause;
- an exact date of disappearance;
- enough T+1 history for a strong conclusion;
- equally strong replication of the pressure/reversal mechanism;
- a current trading recommendation.
Why it matters
The part I actually care about
This research ended up being less about a calendar anomaly and more about a recurring weakness in backtesting.
A 75-year sample can be statistically correct and economically misleading if it combines regimes in which the mechanism was strong with regimes in which it became indistinguishable from noise.
For me, robustness should not only ask whether a result survives costs, parameters, or a bootstrap.
Does the relationship I am trying to exploit still exist?
That is why I separate temporal stability, replication, and falsification from the full-sample result.
I am not presenting this as a new trading strategy. I am presenting it as an example of why a historical result can remain alive inside a backtest long after it has lost relevance in the current market.
Verification
Reproduce the research
The page publishes the same verification package used for the final validation.
To run the reproduction you need the .do script and qtomdecay v0.3.1.
The two output bundles contain the frozen results used on this page and let you verify that your reproduction matches the published analysis.
Reproduction: 1 + 2 · Full verification: 1 + 2 + 3 + 4
Reproduction · Step 1
reproduce_tom_decay.do
Public script that reproduces both samples using relative output paths.
Type: Stata .doDownloadReproduction · Step 2
qtomdecay
v0.3.1The frozen research tool used in the final validation.
Type: ZIPDownloadVerification · Yahoo
Yahoo S&P 500 1950+ outputs
Regimes, adjacent tests, breaks, rolling windows and report for the Yahoo sample.
Type: CSV · JSONDownloadVerification · French
French US Market 1950+ matched outputs
The same outputs for the independent, horizon-matched replication.
Type: CSV · JSONDownload
Every number published on this page can be traced back to an output from the research bundle.
Validation environment
- StataNow 18.5 MP
- Python 3.11.16
- pandas 2.0.3
- HAC/Newey-West inference
- Yahoo Finance
- Kenneth French Data Library
- qtomdecay 0.3.1
Hashes and manifests
SHA-256 for every published file. The original manifests ship alongside the outputs.
Kenneth French source download (SHA-256)
39f9ae1d0e9f575024bc23145980ac270cea508fb67e592578b3f4d65f36d006Yahoo S&P 500
- publication_regime_summary.csv
443d967696d0f2cb34b49efca4711430aa6eee692bba6da637f7c30f70b5083a - publication_adjacent_tests.csv
5e0917ef46b56865e543df9f12e563de59a8eb7852e0b812fdc4b0ed67eeb58c - break_tests.csv
1c6bddc54107c0bccf62ee47f736ce7dd5c63bc72f093444cfa6fee050214f95 - rolling_premium.csv
3ab6b8a2f206edbedc6d6065eeb7e0da2af91ef22cfaafac364e1776a2b054bd - calendar_pairwise_tests.csv
349748075b44613714be9e9f310059719bb289498949c5b20fe5b41c7f4f8c07 - research_report.json
c7ed6a902c46939c6053637fb365e594964f270a67bcebda5557fd362c0acdff
French US Market
- publication_regime_summary.csv
ee2ce0b4f5b08f2f583c19dd75346807125cf77958bde34e566ba89cf0fadf01 - publication_adjacent_tests.csv
667d8d46fcd69377d3ac7f921bd166391242a1b5404d71d047232c5e40a1e0b7 - break_tests.csv
ffc03e41864e2ab8d9825f12c25ae74cfede056a86fb17855d32f8632fe34e71 - rolling_premium.csv
43ed1384ba7a9277b88c437a10ea240b6a036b2a1df4e5aa1613f21974b4d331 - calendar_pairwise_tests.csv
59ed20895e1ee64bf2196aa5244a7854b10f214bec627145759c9382be29462b - research_report.json
16ca1e0db85c927499e90e2ddd8ed14d3efdf8486ac6161f0f5db1539ab21d50
Download
- reproduce_tom_decay.do
04234568cdceda44760bca58bf49d8976a00e1d4c642f70aaf00de8c02b95edd - qtomdecay_v0_3_1_statanow185.zip
95392006fb854083619577911ed0d9810e6f70ecbfd872ee2fec4c8e6ebd0f3f
Methods and sources
Methodological detail
Opened section by section so it does not interrupt the main read.
Event definition
T = 0 is the last trading day of the month. The canonical TOM window is T, T+1, T+2 and T+3.
The benchmark is every other trading day. The window is never optimized at any point in the study.
Data provenance
Yahoo Finance via yfinance, symbol ^GSPC, with the return defined as the provider's adjusted-close percentage change.
Kenneth French Data Library, Fama/French 3 Factors [Daily], with the market return reconstructed as (Mkt-RF + RF) / 100 over a CRSP-based value-weighted universe.
The two series differ in both provider and market universe. The comparison is deliberately an independent replication, not a duplication.
Statistical inference
All tests use HAC/Newey-West standard errors with lags selected by sample length (Newey and West, 1987).
Published p-values come from the frozen bundle outputs and are not recomputed in the browser.
Historical cutoffs
Cutoffs are external and fixed ahead of the analysis: publication era (1987), T+5 → T+3 (1995), decimalization complete (2001), T+3 → T+2 (2017) and T+2 → T+1 (2024).
They are time boundaries, not causal identification. Several structural changes overlap inside each span.
Rolling estimation
A fixed 10-year trailing window with a year-end endpoint, a HAC 95% interval, and the premium expressed in bps per day.
Rolling windows are descriptive. The final window in each series ends after the data end and is marked as incomplete.
Matched-sample replication
The main public comparison uses the matched horizon from 1950 onward in both sources.
The full French sample from 1926 exists as supporting evidence, but it does not replace the matched comparison.
Exploratory analyses
The automatic breakpoint search is explicitly exploratory and its p-value is not adjusted for multiplicity.
It must not be read as a confirmatory test or as the detection of a disappearance date.
Limitations
This is not a causal-identification study.
The T+1 sample is short and exploratory.
The pressure/reversal mechanism does not replicate with the same strength in the matched sample.
Calendar concentration does not survive robustly.
The study does not evaluate transaction costs, capacity or implementation, and is not an investment recommendation.
Reference
Glossary
Concise definitions of the statistical terms used throughout the study.
- BPSBasis points
BPS means basis points. 1 bp equals 0.01%, and 100 bps equal 1%.
They express small return differences clearly without confusing percentages with percentage points.
- HACRobust standard errors
HAC means Heteroskedasticity and Autocorrelation Consistent. It is a standard-error adjustment robust to heteroskedasticity and serial correlation.
It helps prevent statistical uncertainty from looking artificially small when variability changes or nearby observations are related.
- p / p-valueEvidence under the null
A p-value is the probability, under the null hypothesis and the test model, of observing a result at least as extreme as the one obtained.
A smaller value indicates greater incompatibility with the null; it does not measure the probability that the null is true or the result's economic importance.
Scientific literature
References
Publications cited for the study's historical and methodological context.
Robert A. Ariel (1987).
A Monthly Effect in Stock ReturnsJournal of Financial Economics, 18(1), 161–174.
DOI: 10.1016/0304-405X(87)90066-3
Josef Lakonishok and Seymour Smidt (1988).
Are Seasonal Anomalies Real? A Ninety-Year PerspectiveThe Review of Financial Studies, 1(4), 403–425.
DOI: 10.1093/rfs/1.4.403
John J. McConnell and Wei Xu (2008).
Equity Returns at the Turn of the MonthFinancial Analysts Journal, 64(2), 49–64.
DOI: 10.2469/faj.v64.n2.11
Whitney K. Newey and Kenneth D. West (1987).
A Simple, Positive Semi-Definite, Heteroskedasticity and Autocorrelation Consistent Covariance MatrixEconometrica, 55(3), 703–708.
DOI: 10.2307/1913610