Research · Empirical finance · Reproducible

The Ghost of an Anomaly

Turn-of-the-Month Anomaly Decay · Reproducible Research

An anomaly can still look compelling in a long backtest even when much of its premium belongs to a market that no longer exists.

I started with a simple question: if the turn-of-the-month effect has been documented for decades, does it still belong to today's market?

The interesting part was not confirming that it once existed. That was already known. The interesting part was that becoming public does not coincide with an immediate collapse, while the large deterioration appears later.

1987Publication era1995T+32001Decimalization2017T+22024T+1Wide historical premiumIndistinguishable from zero
Visual abstractAn editorial illustration of the study's arc, not a measured series. The exact data appear further down.

The question

How much of this result still belongs to today's market?

The turn-of-the-month effect is one of the most documented calendar anomalies in U.S. equities (Lakonishok and Smidt, 1988; McConnell and Xu, 2008).

I was not interested in proving again that it worked historically. I wanted to know something more useful: how much of that historical result still belongs to the market that exists today.

I split the sample through time, compared canonical TOM days with every other trading day, and repeated the analysis using a second data source and a different market universe.

What I found

Four readings of the same arc

Pre-1987 TOM premium in the S&P 500.

12.74bps/day

Pre-1987 TOM premium in the S&P 500.

Between the publication era and the pre-decimalization period.

11.66bps/day

Between the publication era and the pre-decimalization period.

After 2001 and before T+2.

2.72bps/day

After 2001 and before T+2.

Modern 10-year rolling HAC 95% intervals include zero.

≈ 0

Modern 10-year rolling HAC 95% intervals include zero.

The anomaly survived becoming well known. Its magnitude did not.

The independent Kenneth French U.S. market replication shows the same central pattern.

First twist

Becoming public did not appear to kill it

If the story were simply “the anomaly was published and arbitrage removed it,” I should see a clear break around 1987 (Ariel, 1987).

I do not.

In Yahoo/S&P 500 data, the premium moves from 12.74 to 11.66 bps per day. In the matched 1950+ French replication, it moves from 12.58 to 14.43 bps.

Direct tests between those regimes do not detect a significant difference.

Yahoo S&P 500bps/day
Yahoo S&P 500: 12.74 → 11.66 bps/day12.74Pre-198711.661987–2001
Change
-1.09 bps/day
Change p
0.8293
French US Marketbps/day
French US Market: 12.58 → 14.43 bps/day12.58Pre-198714.431987–2001
Change
+1.85 bps/day
Change p
0.7106

Per-source comparison of the TOM premium before 1987 and between 1987 and 2001, with the p-value of the direct change between the two regimes.

No immediate collapse detected

Research note

Failing to detect a break does not prove that publication or arbitrage had no effect. It only means that the simple story of an immediate disappearance does not fit these data particularly well.

Main evidence

The decay came later

After 2001, the premium falls to 2.72 bps/day in the S&P 500 and 2.97 bps/day in the matched French replication.

In later regimes the estimate continues to compress and is no longer statistically distinguishable from the rest of the market.

Several microstructure changes overlap — settlement, decimalization, electronic trading, costs, and liquidity. I do not identify any single one as the cause.

What the data do show is a large deterioration during that broader transformation of the market.

Rolling 10-year TOM premium

Fixed 10-year trailing window, with HAC 95% interval

Data source
Rolling 10-year TOM premium-10.000.0010.0020.0030.001987 · Publication era1995 · T+32001 · Decimalization complete2017 · T+22024 · T+11960197019801990200020102020Rolling window endbps/day

The rolling 10-year TOM premium starts above 15 bps/day in the mid-twentieth century, stays wide through the publication era, and compresses toward zero after 2001. In recent windows the HAC 95% interval includes zero in both sources.

The rolling 10-year TOM premium starts above 15 bps/day in the mid-twentieth century, stays wide through the publication era, and compresses toward zero after 2001. In recent windows the HAC 95% interval includes zero in both sources.

Yahoo S&P 500French US MarketHAC 95% intervalMarket changes

The vertical marks flag changes in market structure. They are time references, not identified causes.

Independent replication

Same pattern, independent data

The first sample uses the S&P 500 from Yahoo Finance.

The second reconstructs the U.S. market return from Kenneth French's Mkt-RF + RF, using a CRSP-based value-weighted market universe.

They are not the same series, and I do not treat them as if they were.

That is exactly why the comparison is useful: change the provider and the market universe, and the central decay pattern remains.

Yahoo S&P 500

Source
Yahoo Finance via yfinance
Universe
S&P 500 index (^GSPC)
Return
Provider adjusted-close percentage change

Kenneth French US Market

Source
Kenneth French Data Library
Universe
CRSP-based value-weighted U.S. market
Return
(Mkt-RF + RF) / 100

TOM premium by regime and source

TOM premium by regime and source0.005.0010.0015.0012.7412.58Pre-198711.6614.431987–20012.722.972001–20172.221.84T+21.21-4.19T+1Short samplebps/day

Daily TOM premium by publication regime in both sources. The first two regimes are wide and similar; later regimes fall to a few bps per day.

T+1 is a short sample. I do not draw a strong modern negative-alpha conclusion from it.

Yahoo S&P 500French US Market
Daily TOM premium by regime and source, in bps/day.
RegimeYahoo S&P 500French US Market
Pre-publication12.74p (HAC) < 0.000112.58p (HAC) < 0.0001
Published / pre-decimalization11.66p (HAC) 0.008214.43p (HAC) 0.0010
Post-decimalization / pre-T+22.72p (HAC) 0.52962.97p (HAC) 0.4986
T+22.22p (HAC) 0.76331.84p (HAC) 0.8103
T+1Short sample1.21p (HAC) 0.9027-4.19p (HAC) 0.6924

A different provider, a different market universe, the same central decay.

Ghost alpha

What a long backtest can hide

A longer sample often feels safer.

But length and stability are not the same thing.

If I combine decades in which the premium was large with decades in which the estimate fluctuates around zero, the historical average still carries part of the old signal.

The backtest is not necessarily wrong. It may simply be answering a question that is no longer the one I care about.

The question I do care about is:

How much of the historical result still exists inside the regime I would actually trade?

Layers of evidence

  1. 1950–1987

    A wide and statistically clear premium.

  2. 1987–2001

    The anomaly is public and the premium is still wide.

  3. 2001–2017

    The estimate compresses to a few bps per day.

  4. Recent windows

    The 95% intervals include zero.

A historical average can survive much longer than the phenomenon that created it.

Negative results

I tested the mechanism. Not everything survived.

I also tested several pieces that could help explain the pattern.

Result

Prior pressure → reversal

The S&P 500 result is consistent with the proposed mechanism. In the matched 1950+ French sample, the difference is only marginally significant.

Yahoo S&P 500
Difference 30.04 bps · p 0.0148
French US Market
Difference 23.44 bps · p 0.0585

Verdict

Suggestive, not robust enough to headline

Result

Quarter-end / semi-year concentration

The differences do not survive robustly in the matched-sample comparison.

French US Market
All pairwise p-values > 0.20

Verdict

Not robust

Result

Automatic breakpoint

The exploratory algorithm selects 1964 in Yahoo and 1991 in French. The disagreement is informative: no single robust date is selected across universes.

Yahoo S&P 500
Selected year 1964
French US Market
Selected year 1991

Verdict

Exploratory; no single robust date

I did not remove these results to make the story cleaner. They stay because they are part of the research.

Claim boundary

What the evidence supports — and what it doesn't

The evidence supports

  • a large historical TOM premium;
  • persistence around the publication era;
  • substantial later decay;
  • modern rolling estimates statistically indistinguishable from zero;
  • independent replication of the central decay pattern.

The evidence does not establish

  • decimalization as the cause;
  • an exact date of disappearance;
  • enough T+1 history for a strong conclusion;
  • equally strong replication of the pressure/reversal mechanism;
  • a current trading recommendation.

Why it matters

The part I actually care about

This research ended up being less about a calendar anomaly and more about a recurring weakness in backtesting.

A 75-year sample can be statistically correct and economically misleading if it combines regimes in which the mechanism was strong with regimes in which it became indistinguishable from noise.

For me, robustness should not only ask whether a result survives costs, parameters, or a bootstrap.

Does the relationship I am trying to exploit still exist?

That is why I separate temporal stability, replication, and falsification from the full-sample result.

I am not presenting this as a new trading strategy. I am presenting it as an example of why a historical result can remain alive inside a backtest long after it has lost relevance in the current market.

Verification

Reproduce the research

The page publishes the same verification package used for the final validation.

To run the reproduction you need the .do script and qtomdecay v0.3.1.

The two output bundles contain the frozen results used on this page and let you verify that your reproduction matches the published analysis.

Reproduction: 1 + 2 · Full verification: 1 + 2 + 3 + 4

  • Reproduction · Step 1

    reproduce_tom_decay.do

    Public script that reproduces both samples using relative output paths.

    Type: Stata .doDownload
  • Reproduction · Step 2

    qtomdecay

    v0.3.1

    The frozen research tool used in the final validation.

    Type: ZIPDownload
  • Verification · Yahoo

    Yahoo S&P 500 1950+ outputs

    Regimes, adjacent tests, breaks, rolling windows and report for the Yahoo sample.

    Type: CSV · JSONDownload
  • Verification · French

    French US Market 1950+ matched outputs

    The same outputs for the independent, horizon-matched replication.

    Type: CSV · JSONDownload

Every number published on this page can be traced back to an output from the research bundle.

Validation environment

  • StataNow 18.5 MP
  • Python 3.11.16
  • pandas 2.0.3
  • HAC/Newey-West inference
  • Yahoo Finance
  • Kenneth French Data Library
  • qtomdecay 0.3.1
Hashes and manifests

SHA-256 for every published file. The original manifests ship alongside the outputs.

Kenneth French source download (SHA-256)

39f9ae1d0e9f575024bc23145980ac270cea508fb67e592578b3f4d65f36d006

Yahoo S&P 500

  • publication_regime_summary.csv
    443d967696d0f2cb34b49efca4711430aa6eee692bba6da637f7c30f70b5083a
  • publication_adjacent_tests.csv
    5e0917ef46b56865e543df9f12e563de59a8eb7852e0b812fdc4b0ed67eeb58c
  • break_tests.csv
    1c6bddc54107c0bccf62ee47f736ce7dd5c63bc72f093444cfa6fee050214f95
  • rolling_premium.csv
    3ab6b8a2f206edbedc6d6065eeb7e0da2af91ef22cfaafac364e1776a2b054bd
  • calendar_pairwise_tests.csv
    349748075b44613714be9e9f310059719bb289498949c5b20fe5b41c7f4f8c07
  • research_report.json
    c7ed6a902c46939c6053637fb365e594964f270a67bcebda5557fd362c0acdff

French US Market

  • publication_regime_summary.csv
    ee2ce0b4f5b08f2f583c19dd75346807125cf77958bde34e566ba89cf0fadf01
  • publication_adjacent_tests.csv
    667d8d46fcd69377d3ac7f921bd166391242a1b5404d71d047232c5e40a1e0b7
  • break_tests.csv
    ffc03e41864e2ab8d9825f12c25ae74cfede056a86fb17855d32f8632fe34e71
  • rolling_premium.csv
    43ed1384ba7a9277b88c437a10ea240b6a036b2a1df4e5aa1613f21974b4d331
  • calendar_pairwise_tests.csv
    59ed20895e1ee64bf2196aa5244a7854b10f214bec627145759c9382be29462b
  • research_report.json
    16ca1e0db85c927499e90e2ddd8ed14d3efdf8486ac6161f0f5db1539ab21d50

Download

  • reproduce_tom_decay.do
    04234568cdceda44760bca58bf49d8976a00e1d4c642f70aaf00de8c02b95edd
  • qtomdecay_v0_3_1_statanow185.zip
    95392006fb854083619577911ed0d9810e6f70ecbfd872ee2fec4c8e6ebd0f3f

Methods and sources

Methodological detail

Opened section by section so it does not interrupt the main read.

Event definition

T = 0 is the last trading day of the month. The canonical TOM window is T, T+1, T+2 and T+3.

The benchmark is every other trading day. The window is never optimized at any point in the study.

Data provenance

Yahoo Finance via yfinance, symbol ^GSPC, with the return defined as the provider's adjusted-close percentage change.

Kenneth French Data Library, Fama/French 3 Factors [Daily], with the market return reconstructed as (Mkt-RF + RF) / 100 over a CRSP-based value-weighted universe.

The two series differ in both provider and market universe. The comparison is deliberately an independent replication, not a duplication.

Statistical inference

All tests use HAC/Newey-West standard errors with lags selected by sample length (Newey and West, 1987).

Published p-values come from the frozen bundle outputs and are not recomputed in the browser.

Historical cutoffs

Cutoffs are external and fixed ahead of the analysis: publication era (1987), T+5 → T+3 (1995), decimalization complete (2001), T+3 → T+2 (2017) and T+2 → T+1 (2024).

They are time boundaries, not causal identification. Several structural changes overlap inside each span.

Rolling estimation

A fixed 10-year trailing window with a year-end endpoint, a HAC 95% interval, and the premium expressed in bps per day.

Rolling windows are descriptive. The final window in each series ends after the data end and is marked as incomplete.

Matched-sample replication

The main public comparison uses the matched horizon from 1950 onward in both sources.

The full French sample from 1926 exists as supporting evidence, but it does not replace the matched comparison.

Exploratory analyses

The automatic breakpoint search is explicitly exploratory and its p-value is not adjusted for multiplicity.

It must not be read as a confirmatory test or as the detection of a disappearance date.

Limitations

This is not a causal-identification study.

The T+1 sample is short and exploratory.

The pressure/reversal mechanism does not replicate with the same strength in the matched sample.

Calendar concentration does not survive robustly.

The study does not evaluate transaction costs, capacity or implementation, and is not an investment recommendation.

Reference

Glossary

Concise definitions of the statistical terms used throughout the study.

BPSBasis points

BPS means basis points. 1 bp equals 0.01%, and 100 bps equal 1%.

They express small return differences clearly without confusing percentages with percentage points.

HACRobust standard errors

HAC means Heteroskedasticity and Autocorrelation Consistent. It is a standard-error adjustment robust to heteroskedasticity and serial correlation.

It helps prevent statistical uncertainty from looking artificially small when variability changes or nearby observations are related.

Newey and West (1987)

p / p-valueEvidence under the null

A p-value is the probability, under the null hypothesis and the test model, of observing a result at least as extreme as the one obtained.

A smaller value indicates greater incompatibility with the null; it does not measure the probability that the null is true or the result's economic importance.

Scientific literature

References

Publications cited for the study's historical and methodological context.

  1. Robert A. Ariel (1987).

    A Monthly Effect in Stock Returns

    Journal of Financial Economics, 18(1), 161–174.

    DOI: 10.1016/0304-405X(87)90066-3

  2. Josef Lakonishok and Seymour Smidt (1988).

    Are Seasonal Anomalies Real? A Ninety-Year Perspective

    The Review of Financial Studies, 1(4), 403–425.

    DOI: 10.1093/rfs/1.4.403

  3. John J. McConnell and Wei Xu (2008).

    Equity Returns at the Turn of the Month

    Financial Analysts Journal, 64(2), 49–64.

    DOI: 10.2469/faj.v64.n2.11

  4. Whitney K. Newey and Kenneth D. West (1987).

    A Simple, Positive Semi-Definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix

    Econometrica, 55(3), 703–708.

    DOI: 10.2307/1913610