EOFY Tax-Loss/Gain Correlation
Does a stock's own Jul-Mar return predict its own Apr-Jun return, year after year?
The Australian financial year (FY) runs 1 July → 30 June. 30 June is also the deadline for realising capital losses (or gains) for tax purposes, which concentrates a wave of selling — and sometimes buying — into the last quarter of the FY.
This page tests a simple, stock-specific question: for a given symbol, does its return over the first nine months of the FY (Q1-3, Jul-Mar) correlate with its return over the final three months (Q4, Apr-Jun)? Unlike the lead-lag correlations page, this is not cross-sectional — each stock is tested only against its own history. A stock with 20 years of listed history contributes up to 20 (Q1-3, Q4) pairs to its own r.
A negative r means a strong Jul-Mar run tends to be followed by a weak Apr-Jun quarter (and vice versa) — consistent with profit-taking or tax-loss-selling reversal. A positive r means the trend tends to continue into Q4 — momentum rather than reversal.
⚠ This is a historical pattern in one stock's own returns, not a cross-sectional edge shared by many stocks at once. A high |r| with few years of history (n_years close to the floor of 5) can look dramatic and still be noise — check n_years and fdr_p together, not r alone.
For FY "N" (e.g. FY2025-26 = 1 Jul 2025 → 30 Jun 2026): Q1-3 runs 1 Jul (N-1) → 31 Mar (N), and Q4 runs 1 Apr (N) → 30 Jun (N). Each window's return is computed from the last available close at or before the start date to the last available close at or before the end date. Consecutive FYs share the same 30 June boundary close, so it's computed once and reused as both the exit of one FY's Q4 and the entry of the next FY's Q1-3.
If the closest available trading day to a boundary is more than 10 calendar days away (a halted or highly illiquid stock), that FY is skipped rather than using a stale price.
Raw end-of-day price history is not always split-adjusted at the source. Two guards protect against this showing up as a fake extreme return:
- Any FY where a known split or consolidation date (from corporate_events) falls inside its Q1-3 or Q4 window is excluded entirely.
- Any single-quarter return greater than ±300% is excluded as a likely un-adjusted-split artifact, even if no corporate event is on record for it.
Excluded FYs still show up in the detail-drawer chart (as grey ✕ markers) so you can see what was left out and why, but they are not used in the correlation, the mean/std figures, or the regression line.
A stock needs at least 5 included FY pairs before it gets a row at all. n_years in the table is the count of FYs actually used after the guards above — not the stock's total years of listed history. Statistically, n=5 needs |r| ≈ 0.88 to reach p < 0.05, so short histories only surface at all when the pattern is very strong; use the min-years filter to require more before trusting a row.
Every eligible stock (currently listed, ≥5 included FYs — around 1,500 symbols) gets one correlation test. At a raw p < 0.05 threshold that alone would produce roughly 75 false positives by chance. Benjamini-Hochberg FDR correction is applied across the full set of p-values in one pass, producing the fdr_p column — this is the number to check before treating a result as real.
| Control | What it does |
|---|---|
| Industry | Restrict to one GICS industry. Leave blank to see all industries. |
| Market cap min/max | Filter by current market cap, in $M. Market cap is computed live (shares outstanding × latest close), not a stale snapshot. |
| Min years | Minimum number of included FY pairs (default 8). Raise this to focus on stocks with a longer, more trustworthy history; lower it to see newer listings. |
| Min |r| | Minimum absolute correlation strength (default 0.20). Raise it to see only the strongest self-relationships. |
| Direction | Positive = trend tends to continue into Q4 (momentum). Negative = strong Q1-3 tends to reverse in Q4 (consistent with tax-loss/gain selling). Both shows everything. |
Click any column header to sort (sorting happens on the server, over the full filtered set — not just the rows on screen). Click a row to open the detail drawer on the right. A ★ badge next to the symbol marks fdr_p < 0.05.
| Column | Description |
|---|---|
| Years | Number of FY pairs actually used (after split/outlier exclusions). |
| r | Pearson correlation between the stock's own Q1-3 and Q4 returns across those years. |
| FDR p | Benjamini-Hochberg corrected p-value. Below 0.05 is the bar for "probably not chance." |
| Dir | positive (momentum) or negative (reversal) — the sign of r in plain words. |
| Mkt Cap | Current market capitalisation. |
Clicking a row opens a side panel with:
- A scatter chart of every FY's (Q1-3 return, Q4 return) pair — blue circles are included years, grey ✕ markers are years excluded by the split/outlier guards (hover for the reason).
- A dashed regression line through the included points, using the same slope/intercept the backend computed for r.
- Key statistics: r, p-value, fdr_p, n_years, direction, FY range, and mean ± std for both quarters.
💡 If most of the scatter points cluster near the regression line except one or two outliers dragging it, treat the fit with more caution than fdr_p alone suggests — a handful of years can dominate a Pearson r when n_years is small.
The three tabs above the table (Full Quarter, Late May, Rest of Q4) ask a finer question: is the Q1-3-vs-Q4 pattern spread evenly across the whole quarter, or concentrated in a specific part of it? Q4's 91 days were numbered day 1 (1 Apr) through day 91 (30 Jun), then split into Late May = day 57-70 (roughly 27 May-10 Jun) and Rest of Q4 = day 71-91 (10-30 Jun). Everything else about the methodology — split/outlier guards, FDR correction, min-years floor — is identical to the Full Quarter tab; only the end date of the return window changes.
The Late/Rest split wasn't arbitrary: an earlier investigative pass pooled the 50 stocks with the strongest Full Quarter |r| and broke their Q4 into 13 individual weeks. Two adjacent weeks — day 57-63 and day 64-70 — stood out sharply (r ≈ 0.22 and 0.26, both p < 1e-5) against a scatter of weak, mostly non-significant weeks either side. Late May is that combined two-week window; Rest of Q4 is everything after it, through 30 June.
⚠ That 50-stock pooled test is not on this page and should not be treated as confirmation by itself: it pre-selected stocks that were already known to correlate on the Full Quarter test, so a strong pooled r within that group is partly circular — of course a hand-picked correlated subset also correlates on a sub-window of the same data. The Late May / Rest of Q4 tabs you see here are the independent check: every eligible current stock (~1,500, not a pre-selected 50) gets its own r for each window, exactly like the Full Quarter tab.
Result of that independent, full-universe check: Late May clears FDR p < 0.05 for about 1.3% of eligible stocks, versus about 0.6% for Rest of Q4 — roughly double the hit rate. That's a real, reproducible difference between the two windows, but keep the base rate in view: even Late May's stronger showing means the overwhelming majority of stocks show no significant pattern in either window. Use these tabs to find the specific stocks where the pattern holds, not as evidence of a market-wide late-May effect.
Negative (reversal) is the pattern the tax-loss-selling theory predicts: a stock that has run up over Jul-Mar attracts profit-taking into 30 June, or a stock that has fallen attracts bargain-hunting after loss-sellers are done. In practice this is the more common direction among the significant results.
Positive (momentum) means Q4 tends to extend whatever happened in Q1-3 rather than reverse it — the opposite of the tax-effect story, and worth checking for an alternative explanation (e.g. a stock whose Q1-3 news flow keeps compounding into Q4 earnings season).
- fdr_p < 0.05 — the ★ badge. This is the single most important filter; without it, |r| alone is not evidence.
- n_years ≥ 10-15 — comfortably above the n=5 floor, where a given |r| is much less likely to be a small-sample fluke.
- n_outliers_excluded is small relative to n_years — if a third of a stock's history was excluded, the remaining sample is thinner than n_years alone suggests.
- mean Q1-3 and mean Q4 both make sense together — for a reversal pattern, mean Q4 should sit notably closer to zero (or negative) than mean Q1-3; that's the tax-effect signature, not just a negative r driven by noise.
- Thin samples near n_years = 5 — need |r| ≈ 0.88 just to reach p < 0.05; a dramatic-looking r with few years is the least trustworthy result on the page, not the most.
- Pre-2000s history on very old listings — split-adjustment for splits/consolidations more than a couple of decades ago is harder to independently verify; the outlier guard (±300%) catches the worst cases but isn't a complete guarantee for the very oldest stocks.
- Survivorship — only currently-listed stocks get a row. Stocks that were delisted, taken over, or went to zero partway through their history are not included, which biases the visible universe toward survivors.
- One test per stock, once — there is no train/backtest split here (unlike the lead-lag correlations page). A pattern holding across n_years is the only out-of-sample evidence available; there's no separate held-out period to confirm it in.
Suppose you see this row:
| Symbol | Industry | Years | r | FDR p | Dir |
|---|---|---|---|---|---|
| WBC | Banks | 37 | -0.68 | 0.0014 | negative |
Reading this: across 37 FYs of Westpac's trading history (1988-89 to 2025-26), the average Q1-3 return was +17.6% and the average Q4 return was -1.0% — Q4 gives back almost none of a typical Jul-Mar gain on average, but the correlation of -0.68 says the relationship is stronger than the averages alone suggest: FYs with the biggest Jul-Mar rallies tend to have the weakest (or most negative) Apr-Jun quarters, and FYs with a soft Q1-3 tend to see a comparatively better Q4.
With fdr_p = 0.0014 (well under 0.05) and n_years = 37 (well above the floor of 5), this is a well-supported result, not a small-sample artifact. The direction is consistent with the tax-loss/gain-selling theory: a large, widely-held stock like WBC attracts profit-taking and portfolio rebalancing into the 30 June deadline after a strong run.
This does not mean Q4 will always be weak after a strong Q1-3 — it means that, historically, it has tended to be. The scatter chart in the detail drawer is worth checking before relying on this: it shows whether the -0.68 comes from a broadly consistent pattern across most of the 37 years, or from a few extreme years pulling the line.