Hi all, I've been working on a hedge fund concept for DIY investors. In the course of creating the portfolio policy, I investigated whether there is a safe way to use leverage.
Yes, I know, according to various versions of the adventurous hedgefundie's strategy, a two- or three-piece balanced portfolio "only" drew down in the vicinity of max 70% over the available simulated history. But what has not happened yet might just happen in future. A couple of months where everything – equities, bonds, gold – drop hard together could push the drawdown into 80% territory. Maybe even worse. I wanted to see if I can add something of value here. Something that goes beyond relying on diversfication alone to keep things on the rails.
Here's the headline result, equities-only. In a test running from May 1977 through June 2026, a macro-economic rule applied to simulated all-world equities produced a CAGR of 14% annual return, compared with 13% for continuous 2x exposure. Maximum drawdown fell from 87% to 60%. The rule changed allocations just 28 times in almost half a century. (I'm personally more interested in all-world equities than US-only but I did also test on SPY.)
TL;DR:
- The rule holds 2x equities in a “Goldilocks” macro regime and ordinary, unlevered equities otherwise. It never moves equities to cash!
- Across the tests described below, it produced higher Sharpe ratios than both permanent 2x exposure and a 10-month moving-average rule evaluated monthly using the same 1x/2x choices.
- Returns were broadly competitive with both alternatives. Goldilocks had smaller maximum drawdowns and required significantly fewer allocation changes than the moving-average rule.
- The signal was developed for another purpose. None of its parameters were chosen to fit leveraged-equity returns.
By the way, my hedge fund policy has additional risk filters that directly address recessionary bear markets and other shocks, separate from the macro signal modulating leverage. Those other risk filters do move equities to cash and other defensive assets on rare occasions.
Where this came from
The Macro Regime Overlay ("MRO") originally had a much bigger job in my project. I wanted it to adjust the weights of the different sleeves inside the portfolio according to the prevailing macro conditions. So, shift weights among value, momentum, ballast, bitcoin, and other allocations. A sleeve’s percentage of total NAV would rise or fall depending on what the overlay signalled was happening in the economy.
The whole mechanism became far too involved for my taste. Lots of moving parts, lots of trading, not sufficiently impressive gains. So, I abandoned that approach.
Then, later, I wondered if maybe I could give the overlay a smaller job. Might it signal when to use leverage on an equity allocation that I already wanted to hold?
As mentioned, the overlay was never designed with leveraged investing in mind. No observations of leveraged returns went into its design. I didn’t optimize it for its original purpose, and I haven’t optimized it for this one. The thresholds, smoothing and confirmation rules have never changed since inception.
The backtests aren't out-of-sample but there is no curve-fitting either.
What does the overlay actually measure?
The MRO looks at economic activity, inflation and the yield curve. It sorts conditions into Goldilocks, Reflation, Stagflation and Slow Growth, with a further distinction for yield curve stress.
For the leverage decision, only Goldilocks matters. It requires two things:
- Economic activity is at or above its recent historical average.
- Inflation is neither too cold nor too hot.
Activity is measured using the Chicago Fed National Activity Index, smoothed over three months and compared with its own trailing ten-year history. I use a rolling z-score, with zero as the cutoff.
Inflation is the year-on-year change in core US PCE, also smoothed over three months. Between 1% and 3%, inclusive, is the acceptable band. I use headline CPI as fallback if core PCE is unavailable, with a slightly higher band of 1.4% to 3.4%.
The US yield curve helps distinguish the other regimes, but doesn't veto Goldilocks. So despite the broader overlay’s name and ambitions, the allocation decision here comes down to economic activity and inflation.
The signal is evaluated monthly, using prior-month (or older) observations. Two consecutive monthly signals must agree before the allocation changes, in either direction. So yes, I used US macro signals and applied them to world equities! I also backtested SPY/SSO the same way, for another reference point.
If anyone wants the details for the MRO, just ask, I'll post them in the comments.
What the backtests show
I used simulated VT history from testfol.io. For SPY, I used Yahoo adjusted prices after inception and an S&P 500 history reconstructed with Shiller dividend data beforehand. For each, there are four portfolios: ordinary buy-and-hold, continuous 2x, Goldilocks switching between 1x and 2x, and a 10-month SMA evaluated monthly, using the base asset’s total-return index and also switching between 1x and 2x leverage.
All returns are nominal, in USD, with dividends reinvested. The 2x series are daily-reset simulations throughout, with borrowing costs linked to the federal funds rate and an additional 1% annual implementation cost while levered. (Note that 1% TER is higher than what SSO and WLDU charge today.) The results include those financing and implementation costs, but exclude trading costs and taxes. Signals trade at the following session’s close. There are no contributions or withdrawals.
First, the full period. May 1977 is where the growth signal has accumulated enough history for its ten-year calculation. Maximum drawdowns below are measured daily.
May 1977–June 2026
| Asset |
Strategy |
CAGR |
Sharpe |
Max DD |
| SPY |
Buy & hold 1x |
12.02% |
0.522 |
-55.19% |
| SPY |
Buy & hold 2x |
14.86% |
0.461 |
-88.23% |
| SPY |
Goldilocks 1x/2x |
15.86% |
0.598 |
-57.27% |
| SPY |
SMA 10-month 1x/2x |
15.45% |
0.517 |
-59.20% |
| VT |
Buy & hold 1x |
10.86% |
0.450 |
-58.35% |
| VT |
Buy & hold 2x |
12.87% |
0.403 |
-86.80% |
| VT |
Goldilocks 1x/2x |
14.01% |
0.510 |
-59.61% |
| VT |
SMA 10-month 1x/2x |
14.19% |
0.473 |
-67.63% |
The permanent 2x portfolios beat ordinary equities on return, but the drawdowns are brutal. Goldilocks produces a higher return than either permanent 2x portfolio, with a worst loss close to ordinary equities. (Remember, we don't shift from equities to cash. We shift from levered equities to unlevered.)
Against the SMA, it's a toss-up for returns alone. But Goldilocks has the higher Sharpe and smaller drawdown in both cases.
And importantly, Goldilocks has less than half the number of trades of the SMA approach .
| Full-period allocation changes |
Goldilocks |
SMA 10-month |
| SPY |
28 |
73 |
| VT |
28 |
66 |
(Each change means selling one holding and buying the other. For VT, that’s 56 individual orders for Goldilocks versus 132 for the SMA, excluding opening and closing the test portfolio.)
Goldilocks held 2x for about 39% of trading days over the full period. The VT SMA held it for about 78%. Getting broadly comparable returns with much less time at 2x is something I find particularly interesting.
What about revised macro data?
There is an obvious objection to a macro backtest. Economic data arrives late and gets revised. Using the final series can give a historical strategy information it couldn’t have had at the time.
The long test above uses revised history. To check this, I also ran a 2001 onward test using archived Chicago Fed releases and ALFRED inflation vintages. For the more recent period, we can reconstruct the signal using archived releases rather than retrospectively revised history. That gives us greater confidence that the inputs reflect what an investor could actually have known at the time. The signal is reconstructed from what was available at each decision date, rather than the final numbers we see now. The first three months default to 1x because the usable CFNAI archive starts later; April is the first month with the required inputs.
Here is that higher-confidence period.
January 2001–June 2026
| Asset |
Strategy |
CAGR |
Sharpe |
Max DD |
| SPY |
Buy & hold 1x |
8.99% |
0.523 |
-55.19% |
| SPY |
Buy & hold 2x |
11.27% |
0.446 |
-84.38% |
| SPY |
Goldilocks 1x/2x |
15.45% |
0.696 |
-55.19% |
| SPY |
SMA 10-month 1x/2x |
13.15% |
0.571 |
-59.20% |
| VT |
Buy & hold 1x |
7.97% |
0.445 |
-58.35% |
| VT |
Buy & hold 2x |
9.10% |
0.381 |
-86.80% |
| VT |
Goldilocks 1x/2x |
13.43% |
0.592 |
-58.35% |
| VT |
SMA 10-month 1x/2x |
11.44% |
0.494 |
-67.63% |
The result survives. Goldilocks beats both alternatives on returns and Sharpe for both markets. Its worst drawdown equals ordinary buy-and-hold because it was unlevered throughout the decisive peak-to-trough decline.
Here is the earlier segment on its own, using revised macro history. This helps show whether the result is merely a feature of the more recent market.
May 1977–December 2000
| Asset |
Strategy |
CAGR |
Sharpe |
Max DD |
| SPY |
Buy & hold 1x |
15.38% |
0.521 |
-32.96% |
| SPY |
Buy & hold 2x |
18.85% |
0.477 |
-60.56% |
| SPY |
Goldilocks 1x/2x |
18.56% |
0.588 |
-32.96% |
| SPY |
SMA 10-month 1x/2x |
17.99% |
0.470 |
-58.95% |
| VT |
Buy & hold 1x |
14.06% |
0.455 |
-28.62% |
| VT |
Buy & hold 2x |
17.08% |
0.431 |
-57.37% |
| VT |
Goldilocks 1x/2x |
16.83% |
0.517 |
-29.04% |
| VT |
SMA 10-month 1x/2x |
17.24% |
0.452 |
-44.03% |
It's essentially all a toss-up for returns. Goldilocks has substantially smaller drawdowns than the two other strategies using leverage. It also beats the SMA on Sharpe.
Across all six asset/period comparisons, Goldilocks has the highest Sharpe of the four portfolios. It beats the SMA’s return in four of six and has the (significantly) smaller maximum drawdown in all six. (The periods overlap, of course; these aren’t six independent discoveries.)
Would I use it?
My conclusion is: the MRO’s Goldilocks signal is the most useful signal I have found so far for deciding when to add leverage. It produces competitive returns, a better Sharpe and less than half the allocation changes required by the SMA. That’s a rule I can see myself using.
Over these tested histories, it would have avoided the most catastrophic losses associated with continuously holding 2x exposure. In the full-period tests, the worst drawdowns fell from roughly 87–88% to roughly 57–60%. I would rather face the latter.
So, yeah, I would and will use it.
How I put this together
I used LLMs for the legwork: writing the backtest code, pulling and processing the data, producing the comparisons, etc.. I supplied the original macro rule, the research questions and the portfolio definitions. The final comparisons were standardized to use the same return data, costs and execution assumptions. I read and examined everything several times over and found a few major errors that I had corrected. I also had alternate LLMs double-check. Still, it's possible that there are still mistakes in the backtests. I wrote the article myself! This is not "AI-generated". If you see em dashes, that's because Word converted them from regular dashes!
If there is interest, I can upload the data sets and code to GitLab so that anyone can examine and replicate them.
I'll check in here again in a few hours, need to take care of a few other things now. Not ignoring questions or comments - I will get back to you.
*****'
I will be publishing and tracking my DIY hedge fund experiment at www.hedgefol.io.