Combining Weak Edges Under Risk and Turnover Constraints

Wait 5 sec.

Combining Weak Edges Under Risk and Turnover ConstraintsE-mini S&P 500 FuturesCME_MINI_DL:ES1!EdgeToolsPart 4 of 5: The Signal Book Five indicators can sit on a SPY chart forever and cost you nothing. Parts 1 to 3 treated them that way. A mean-reversion score has a small but real link with the next five-day return, an information coefficient of about 0.076: on days when the score is more bullish than usual, the next five days tend to be a bit more bullish than usual, which is not yet a trading rule. Four of five such scores are largely one idea wearing different costumes, so counting five indicators is counting about 1.9 independent ideas. And that 0.076 is an average across calm stretches and chaotic ones, not a description of either. All of that was measurement. This part is the first time measurement becomes a position, and a position, unlike a score, has to be bought and sold. The question is therefore narrower. Given five weak, overlapping, regime-dependent forecasts, how do you turn them into one holding, and how much of the combined edge is still there after the broker has been paid? From a shelf of scores to one book The unit of work from here on is not the indicator. It is the signal book. Plain version: a fixed recipe that takes today's readings from all your forecasts and returns one number, the position you should be holding. Precise version: a weighted combination of standardised forecasts, mapped through a sizing rule into a target exposure. Tiny example: two scores on the chart today, +1.0 and −0.4, with 70% of the vote given to the first and 30% to the second, produce a combined score of 0.7(1.0) + 0.3(−0.4) = 0.58. A rule that goes fully invested at two standard deviations turns that into 0.58 / 2 = 29% of capital, long. The 70/30 split in that sum is the forecast weight, and choosing those numbers is the entire subject of this part. Everything else in the chain is bookkeeping, though the bookkeeping has to be honest, which is reason enough to see the whole pipeline before arguing about any single link in it. Figure 1: From five scores to one position. The weights are re-estimated as history accumulates, and only from data that had already happened. That last point is not decoration. Part 3 showed that a regime model fitted on the whole sample quietly reads the future and disagreed with a real-time version on roughly 9% of days. Weights are worse. A weight fitted on the full history already knows which of your five signals later worked. Every number below therefore comes from weights that were re-estimated twice a year, after a three-year warm-up, using only days whose five-day outcome had already been observed. Sixty-two such updates produce a book that runs from January 1997 to September 2026, 7,450 daily observations of profit and loss. The instinct to average, and why it is hard to beat Ask a trader with five indicators what to do and the answer is usually to average them, or to demand that a majority agree. This is worth taking seriously rather than dismissing, because equal weighting has a strong record against cleverer schemes in exactly the situation we are in, namely one where the inputs must be estimated from noisy data (DeMiguel, Garlappi and Uppal, 2009). It also has one obvious defect that Part 2 already exposed: it hands a full share to a score that is a near-copy of another, and none of the arithmetic notices. So begin with the upgrade most people reach for next, which is to give the strongest forecast the largest share. The second number in that comparison needs a name first. An information ratio here is the book's annualised daily profit divided by how much that profit jumps around. Plain version: return per unit of wobble, before anyone takes a cut. Precise version: the annualised mean of daily book profit and loss, divided by its standard deviation. It is not a Sharpe ratio against cash and it is not a comparison with holding SPY. Tiny example: a book that makes 6.8% a year at 11.6% volatility has an information ratio of 6.8 / 11.6 = 0.59. Run both books on SPY and the upgrade is not encouraging. Equal weight across all five scores: combined IC 0.085, information ratio before costs 0.55 Weighted by each score's own IC: combined IC 0.087, information ratio before costs 0.51 The quality-weighted book has a slightly better combined forecast and a distinctly worse result. That is not a paradox once you remember what the rule was allowed to see. It looked only at how good each forecast is on its own, and the four price-based scores all look good on their own for the same underlying reason, so the rule loaded up on four descriptions of one idea and let the tilt drift toward whichever description happened to look best. Quality without uniqueness is not new information, and a weighting rule that cannot tell the difference will keep buying the same bet under different names. Before a better rule can be written, the thing it needs to know has to be measured. Mistakes that move together Part 2 measured overlap between the scores themselves, finding correlations from 0.65 to 0.93 within the price family. What matters for sizing is a related but distinct quantity: how often the forecasts are wrong on the same day. Plain version: if two forecasts tend to fail on the same days, holding both of them is not two bets, it is one bet held twice. Precise version: the forecast errors move together, so the combined score jumps around more than independence would imply. Tiny example: two forecasts each contributing half the book, with a correlation of 0.9 between them, produce a combined variance of 0.25(1 + 1 + 2 x 0.9) = 0.95 rather than the 0.50 that independence would give, so the position is carrying about sqrt(0.95 / 0.50) = 1.38 times the risk its author thinks it is. Both halves of that are measurable here. Call a miss a day on which a score pointed one way and the realised five-day return went the other. The −z(20) and RSI(14) reversion scores miss together on 47.5% of days where independence would predict 26.9%, a ratio of 1.77, and the correlation between their miss indicators is 0.83. Across the six pairs inside the price family the ratio runs from 1.46 to 1.77. The control matters more than the headline. If this clustering were merely an artefact of every score being graded against the same return series, then the turn-of-month calendar signal, which Part 2 kept precisely because it is unrelated to price, would cluster just as much. It does not: its four pairings with the price scores give ratios of 1.03 to 1.08, with miss correlations between 0.03 and 0.10. The coincidence in the price family is therefore a property of the scores, not of the grading. The consequence for position size follows from the arithmetic. Weighting the four price scores equally, the variance of the combined score is 0.870 where genuine independence across four inputs would give 0.250, so an author who counted four forecasts and sized as if they were four separate bets is holding 1.87 times the risk that count implies. Including the calendar signal softens it to 1.72 times, which is the whole contribution of adding a genuinely unrelated forecast: it does not raise the edge, since its own IC is near zero, but it does stop the book from being quite so concentrated in one bet. Optional detail: where the factor comes from. Let the standardised scores have correlation matrix Sigma and let the book use weights w. The variance of the combined score is w' Sigma w. Under equal weights w = 1/k this becomes the average of all entries of Sigma, whereas independence would give 1/k. The ratio of standard deviations, risk factor = sqrt( (w' Sigma w) / (1/k) ) is the multiple by which the book is larger than a naive count of k independent forecasts suggests. For the four price scores, sqrt(0.870 / 0.250) = 1.87. This is a statement about the size of the raw weighted sum, not about its expected return, and not about the books below, which re-standardise that sum before sizing. A risk budget for a book is simply a decision, made in advance, about how much bounce the combined position is allowed to have. Someone who gives each of five forecasts a fifth of the book believes they have set a budget in units of risk when they have only set a budget in units of signals, and the gap between the two is 1.72 times on this evidence. That factor describes a book which treats the raw weighted sum as the position. The books whose results follow do not: they re-scale the combined score before turning it into an exposure, so their traded size is not 1.72 times a naive count. What overlap still costs them is forecast quality and, later, turnover, not a silently oversized bet. A weighting rule that knows about overlap The old answer to combining correlated forecasts comes from forecasting rather than from finance: look at how the errors move together, not only at how good each forecast is on its own (Bates and Granger, 1969). The version used here does the same job with the scores themselves and their ICs (Grinold and Kahn, 1999). Each forecast is credited for its quality and debited for its overlap with everything else in the book. The signed misses counted above do not go into that calculation; they are why a calculation of that kind is needed at all. Optional detail: the equation. The weight vector solves Sigma w = ic, where Sigma is the Pearson correlation of the scores and ic is the vector of their Spearman associations with the five-day return. That is not the Bates-Granger estimator, which inverts the covariance of the forecast errors. Absolute weights are then scaled to sum to one. Figure 2: The same five forecasts under four sizing rules. The bars above are the IC and the average overlap from the most recent honest estimate, using only data that had already happened. Notice what happens to −z(20): it has a perfectly respectable IC of 0.074 on its own, yet the overlap-aware rule gives it a weight of −0.35, because whatever it contributes is already supplied by −z(10) and −z(50). That negative weight is the pivot of this section. It is not a coding error and it is not nonsense: given −z(10) and −z(50), the −z(20) score carries almost no new information, and the arithmetic finds it slightly useful as a correction term rather than as a forecast. Part 2 had already shown the same thing from the other direction, with −z(50)'s IC falling from 0.078 to 0.026 once −z(20) was removed from it. Still, a rule that turns a positive forecast into a short position on the strength of a correlation estimate is fragile by construction. Inverting a cluster of highly correlated inputs amplifies whatever error is in those inputs, which is why this kind of optimiser has been described as an error maximiser rather than an optimiser (Michaud, 1989). The usual repair is shrinkage. Plain version: do not take the estimated overlaps at face value; pull them part of the way toward "the scores are independent" before you invert them. Precise version: replace the correlation matrix Sigma with a blend of Sigma and the identity matrix, so the inverse is less jumpy. Tiny example: pulling 30% of the way, a fixed amount rather than the data-driven one Ledoit and Wolf (2004) estimate, moves the −z(20) weight from −0.35 to −0.01, and gives the calendar signal 0.12 instead of 0.04. Both overlap-aware books beat the two naive ones on forecast quality and on the result before costs. Equal weight: combined IC 0.085, information ratio before costs 0.55, turnover 50 times capital a year IC-weighted: combined IC 0.087, information ratio before costs 0.51, turnover 51 times Overlap-aware: combined IC 0.100, information ratio before costs 0.61, turnover 70 times Overlap-aware with shrinkage: combined IC 0.090, information ratio before costs 0.59, turnover 66 times Read that table honestly and it says two things, only one of which is comfortable. Accounting for overlap does raise the combined forecast, from 0.085 to 0.100. Shrinkage, on this sample, gives a little of that back rather than adding to it, taking the IC from 0.100 to 0.090 and the information ratio from 0.61 to 0.59, so the case for it cannot rest on a better number here. It rests instead on the instability of a −0.35 weight on a positively forecasting signal, which is an argument about what happens next year rather than about what happened in this sample, and no experiment run in this part settles it. The uncomfortable column is the last one. The cleverer books turn over 66 to 70 times a year against the plain book's 50, and nothing so far has charged them for it. Turnover is a tax Plain version: turnover is how much buying and selling the book forces you to do, and every unit of it costs the spread. Precise version: the average absolute change in target position per day, multiplied by 252. Tiny example: the shrunk book moves its position by 0.26 of capital on an average day, so across 252 trading days it turns over about 66 times capital a year. A basis point is 0.01 of one percent; at 2 basis points per unit traded that is 66 x 0.0002 = 1.3% of capital surrendered annually. Set that 1.3% against what the book earns and the scale of the problem is plain. The shrunk book made 6.8% a year at 11.6% volatility before costs, which is where the information ratio of 0.59 comes from. A 1.3% charge takes the return to 5.5% and the ratio to 0.48. At 10 basis points the same arithmetic removes 6.6% from a 6.8% return and leaves an information ratio of 0.02, which is to say the entire measured edge belonged to the broker. This is the mechanism behind the finding that anomaly returns survive or fail largely according to how much trading they demand (Novy-Marx and Velikov, 2016). There are two dials available. The first is to trade less by smoothing the position: instead of jumping to today's target, take a short average of the recent targets, and accept a staler forecast in exchange for a smaller bill. Figure 3: Information ratio for the shrunk book before and after costs, as smoothing is relaxed from 55 days down to none, which drives turnover from under 4 to nearly 66 times capital a year. The ring marks the best result after a 2 basis point charge. The grey line before costs keeps rising all the way to the fastest book, while every line after costs turns over and comes back down, and the more expensive the trading the earlier it turns. The grey line rising all the way is the Part 1 half-life result reappearing as a cost problem. Smoothing over 55 days drags the position's own IC from 0.093 down to 0.058, because a five-day forecast averaged over eleven weeks is mostly a record of what used to be true. Freshness has value, and the curve prices it. What the lines after costs add is that the value is finite: at 2 basis points the best book smooths over two days and nets 0.48, at 5 basis points the best mix has moved to three days, and at 10 basis points it has moved to five days and nets only 0.22. The second dial is the weighting rule itself, and this is where the earlier table has to be re-read. Taking the four unsmoothed books and charging each of them at every cost level gives the following. At no cost: overlap-aware 0.61, shrunk 0.59, equal 0.55, IC-weighted 0.51 At 2 basis points: overlap-aware 0.48, shrunk 0.48, equal 0.46, IC-weighted 0.42 At 5 basis points: equal 0.33, shrunk 0.31, overlap-aware 0.28, IC-weighted 0.28 At 10 basis points: equal 0.10, IC-weighted 0.06, shrunk 0.02, overlap-aware −0.05 The ranking inverts completely. The sophisticated book wins on every measure before costs and is the only one of the four that goes negative once trading is dear, while the crude equal-weight book, which nobody would defend on theoretical grounds, finishes first at 5 and at 10 basis points because it trades 50 times a year instead of 70. Whether it was worth modelling the overlap at all is not a question about the signals. It is a question about what you pay to trade. Shrinkage earns its place in that column rather than in the one before costs. Having given up 0.02 of information ratio, the shrunk book is ahead of the unshrunk one at 5 basis points, 0.31 against 0.28, and still positive at 10 where the unshrunk book is losing money. The mechanism is not the estimation stability that motivated shrinkage in the first place, though; it is simply that a book with a −0.01 weight where another has −0.35 has less reason to trade, turning over 66 times a year instead of 70. Two arguments for the same repair, arriving from different directions, and only one of them is visible in this data. Figure 4: Information ratio of the shrunk book after costs, across every combination of smoothing span and trading cost. The yellow ridge is the fastest book still worth running at each cost level; it slides toward heavier smoothing as costs rise, and the whole right-hand side of the surface sags toward zero. The shape of that surface is an instruction. There is no such thing as the correct rebalancing frequency for a signal, only a correct frequency given a cost, and the two decisions cannot be made in separate meetings. It also says something less obvious about weak edges in general: once trading is expensive enough, the surface is nearly flat in the smoothing direction, so tuning the speed of the book stops mattering because there is very little left to tune. Where the Fundamental Law fits There is a back-of-the-envelope that says how good a book can be if bets are independent and trading is free: the information ratio is roughly the information coefficient multiplied by the square root of breadth, the number of independent bets (Grinold, 1989). It belongs here as a ceiling, not as a plan, and one confusion has to be cleared first. Breadth in that relation counts independent bets per year, not the number of signals on the shelf. Part 2's effective breadth of 1.88 answers a different question, namely how many distinct ideas the five scores amount to, and substituting it here would be a category error. This book trades one asset on a five-day horizon, so its breadth is roughly 252 / 5, about 50 bets a year. With a combined IC of 0.090 the relation predicts an information ratio of 0.090 x sqrt(50.4) = 0.64. The book actually delivered 0.59 before costs and 0.48 after a 2 basis point charge. The ceiling is a ceiling, then: it was not breached. Of the gap from 0.64 to 0.48, the first slice, down to 0.59, is the independence the relation assumes and a daily-marked five-day forecast does not have; the rest is the 2 basis point charge. What the relation cannot do is tell you how to build anything, since it assumes bets are independent, which a daily-marked five-day forecast plainly violates, and it assumes trading is free, which Figure 4 exists to contradict. Treat it as an arithmetic check on ambition. A retail book claiming an information ratio of 3 from one asset and a five-day signal is claiming an IC of 0.42, and no such thing has been measured anywhere in this series. What a book gives you, and what it does not The fourth research question of this series was how to turn several weak forecasts into one position without the combination destroying the advantage it was built on. On SPY the answer is that a rule which notices overlap raises the combined forecast from an IC of 0.085 to 0.100, and that whether this survives contact with a broker depends entirely on a number the signals know nothing about. It does not tell you: that these weights are right in any absolute sense; five signals, one horizon, one shrinkage setting and one grid of smoothing spans were all chosen by a person who had seen the whole history, and no expanding-window estimate closes that particular door that 7,450 daily observations are 7,450 independent bets, for the same overlapping-window reason Part 1 raised and Part 2 repeated what your costs actually are; a flat charge per unit traded ignores that spreads widen exactly when volatility rises, which is when Part 3's turbulent regime arrives how large the book can be before its own trading moves the price against it, which is the next problem Weights are where measurement becomes commitment. Sizing five forecasts by their individual quality is worse than sizing them equally, because four of them are the same idea and quality alone cannot see that; sizing them by quality net of overlap raises the combined forecast from 0.085 to 0.100, and produces a negative weight on a signal that forecasts perfectly well alone. Then the bill arrives, the sophisticated book turns out to trade 70 times a year against the crude book's 50, and at ten basis points the crude one wins while the sophisticated one loses money. The edge you can measure and the edge you can keep are separated by your execution, not by your statistics. You should now be able to explain A signal book is one position made from several forecasts, not five strategies running side by side. Weighting by quality alone double-counts near-copies, which is why the IC-weighted book here did worse than simple averaging. When two forecasts fail on the same days, holding both is one bet held twice; the price family missed together 1.46 to 1.77 times as often as independence allows, while the calendar signal stayed between 1.03 and 1.08. Turnover is a tax with a rate: 66 turns a year at 2 basis points costs 1.3% of capital, and at 10 basis points it removed essentially the whole edge. Smoothing buys a lower bill with a staler forecast, and the best trade-off moves toward slower trading as costs rise. IC times the square root of bets per year is a frictionless ceiling, and this book came in below it. Next: Capacity, Impact, and Alpha Decay Every cost charged in this part was charged as if the book were a spectator, paying a fixed toll for the privilege of trading against a market that never noticed it was there. That assumption holds while the position is small and fails when it is not, because a large enough order moves the price it is trying to capture, and a well-known enough edge attracts the competition that erases it. Part 5 turns to the two limits that no amount of careful weighting can design around: how much capital a book of this kind can carry before it spoils its own prices, and what happens to a measured edge once other people have measured it too. References Bates, J.M. and Granger, C.W.J. (1969) 'The combination of forecasts', Operational Research Quarterly DeMiguel, V., Garlappi, L. and Uppal, R. (2009) 'Optimal versus naive diversification: how inefficient is the 1/N portfolio strategy?', Review of Financial Studies Grinold, R.C. (1989) 'The fundamental law of active management', The Journal of Portfolio Management Grinold, R.C. and Kahn, R.N. (1999) Active Portfolio Management. 2nd edn. New York: McGraw-Hill. Ledoit, O. and Wolf, M. (2004) 'A well-conditioned estimator for large-dimensional covariance matrices', Journal of Multivariate Analysis Michaud, R.O. (1989) 'The Markowitz optimization enigma: is "optimized" optimal?', Financial Analysts Journal Novy-Marx, R. and Velikov, M. (2016) 'A taxonomy of anomalies and their trading costs', Review of Financial Studies