Key takeaways
- In Gneezy and Potters' 1997 experiment, subjects who saw results every round bet 50.5% of their endowment; those who saw results every third round bet 67.4%.
- The odds were identical in both groups. Only the feedback schedule changed, and the difference was significant at p = 0.002 over all nine rounds.
- Benartzi and Thaler found the historical equity premium implies an evaluation period of 9 to 13 months, not the decades most investors say they hold for.
- Of 66,465 households studied by Barber and Odean, the most active traders earned 11.4% a year against 17.9% for the market over 1991 to 1996.
- The effect has failed to replicate in several settings, including a sample of 755 adolescents, so treat it as a tendency rather than a law.
Two people buy the same tracker on the same Monday. One of them opens the app before work, every morning. The other checks once a quarter and does nothing in between. Twelve months on they own exactly the same fund. Will they make the same decisions about it?
Almost certainly not. And the reason isn't willpower. It comes from an experiment that had nothing to do with portfolios at all.
Two groups played the same gamble, with the same odds, for the same money. One group saw the result after every round. The other saw results only after every third round. The second group put 33% more of their money at risk.
Nothing about the bet changed. Only how often they looked.
The myopic loss aversion experiment
Uri Gneezy and Jan Potters ran the study and published it in the Quarterly Journal of Economics in 1997. Subjects played nine rounds. In each round they chose how much of a 200-cent endowment to stake on a lottery. It paid 2.5 times the stake with probability one third, and nothing with probability two thirds.
Picture those two hundred cents sitting on the table in front of you. The bet is worth taking. Stake one unit and you gain 2.5 units a third of the time and lose one unit two thirds of the time, so the expected value is (1/3 × 2.5) − (2/3 × 1), a positive sixth of a unit per round. How much you stake should depend on your risk aversion. Why would it depend on how often somebody shows you a scoreboard?
The paper spells out why the schedule matters. Weighing losses more heavily than gains at a rate of λ, a single lottery is attractive only if λ is below 1.25. Evaluate three of them together and it stays attractive up to λ of 1.56, because the probability of seeing a loss at all falls from 0.67 on one round to a much lower figure across three.
So the schedule ought to move behaviour. It did. Averaged across all nine rounds, the high-frequency group staked 50.5% of their endowment. The low-frequency group staked 67.4%. And the gap held in every block of three rounds, so it wasn't one stretch of luck doing the work.
| Share of endowment staked | Feedback every round | Feedback every third round |
|---|---|---|
| Rounds one to three | 52.0 | 66.7 |
| Rounds four to six | 44.8 | 63.7 |
| Rounds seven to nine | 54.7 | 71.9 |
A Mann-Whitney test put the one-tailed significance at p = 0.002 across all rounds. In a second part of the experiment the pattern repeated, with 39.0% against 48.9%.
Why checking your portfolio more often makes losses louder
The mechanism has two parts, and neither is exotic.
The first is loss aversion. Losses register more heavily than equivalent gains. The second is the arithmetic of observation windows. A volatile asset with a positive drift shows a loss surprisingly often over short windows and rarely over long ones. Check daily and you'll see red on something close to half of all days. Check once a decade and you'll almost never see it.
Put those together and your checking habit sets how many losses your brain has to process. That, in turn, sets how risky the thing feels to you. Shlomo Benartzi and Richard Thaler named the combination myopic loss aversion.
Their test ran the logic backwards. If investors have prospect theory preferences, how often would they need to evaluate a portfolio for stocks and bonds to look equally attractive? Drawing from CRSP monthly returns for 1926 to 1990, they found the answer sat “in the neighborhood of 1 year (from 9 to 13 months)”. Take a one-year evaluation period as given, and the allocation that falls out is close to a 50-50 split between stocks and bonds.
That's the uncomfortable part. Ask most people about their horizon and you'll hear thirty years. Watch them behave and you'll see someone with a one-year one. The horizon you act on is the horizon you measure on.
It shows up in markets, not just labs
A fair objection to any lab result is that real markets have prices, arbitrage and people with money on the line. Gneezy, Kapteyn and Potters tested exactly that, building a market experiment where assets traded rather than sitting in an envelope.
They found that more information and more flexibility produced less risk taking, and that market prices of risky assets were significantly higher when feedback frequency and decision flexibility were reduced. Market interaction didn't wash the behaviour out. It got capitalised into prices.
Their paper opens with a real case. In February 1999 Bank Hapoalim, Israel's largest mutual fund manager, said it would send fund performance information quarterly rather than monthly. Clients could still check daily if they wanted. The bank expected that fewer prompts would leave investors more willing to hold the funds.
What checking your portfolio costs through the behaviour it drives
Checking isn't expensive on its own. What it does is raise the odds of acting, and acting is where your money goes.
Barber and Odean studied 66,465 households at a large discount broker over 1991 to 1996. The households that traded most earned an annual return of 11.4% while the market returned 17.9%. The average household earned 16.4% net. Gross of costs it earned 18.7%, slightly ahead of the index — so the shortfall was not stock picking. It was the friction of transacting, on turnover that exceeded 250% a year for the most active group.
Say you had been one of those busy households. Your stock picking would have been about as good as anybody's in the sample. You'd still have finished those six years on 11.4% a year while the market did 17.9%. The trading did that, not the choosing.
That's the same shape as the fund-level gap between what investments returned and what investors in them earned, which we looked at in the piece on what bad timing costs the average dollar. Different data, different decade, same direction.
Checking your portfolio is not the same as acting on it
Two habits usually travel together, and they're worth pulling apart.
Opening an app and reading a number costs nothing. It changes nothing about your portfolio, your fees or your expected return. If that were all anyone did, none of this research would matter.
What the evidence describes is a chain. Frequent feedback makes the holding feel riskier, feeling riskier raises the odds you'll trim, and trimming is where the cost lands. Barber and Odean's most active households turned over more than 250% of their portfolios a year. That's not a group calmly observing. That's a group reacting.
The Bank Hapoalim decision is instructive precisely because it didn't restrict anything. Clients could still check daily. The bank only changed what arrived unprompted. That's the lever most of us actually have: not willpower, but how many times a week something puts the number in front of you.
Notifications are the modern version of that monthly statement. Who chose their frequency? A push alert on a 2% move is a design decision made by somebody whose interests aren't identical to yours.
The objections on how often to check investments, and they have force
This is where a lot of behavioural writing stops. It shouldn't.
Myopic loss aversion has a mixed replication record, and the failures aren't obscure. Alexander Klos reviewed them in Judgment and Decision Making. Moher and Koehler reproduced the original Gneezy and Potters design successfully. But they found no myopic loss aversion when subjects played similar gambles rather than the same gamble repeatedly, and none when the task was a choice between two gambles rather than an allocation.
More pointedly, Glätzle-Rützler and co-authors found no evidence of the effect in a sample of 755 adolescents aged 11 to 18, using the original design. Beshears and co-authors found no effect from aggregating return information despite a similar presentation to earlier work.
Klos proposes two candidate explanations. Task structure matters: certainty-equivalent framings appear to remove the effect that binary choices produce. And subjects severely misestimate long-term returns, which may be doing work that gets attributed to myopia.
So the honest position is narrower than the popular version. Feedback frequency changes risk taking in some designs and not others. It's a tendency with boundary conditions, not a law of nature. Anyone who tells you the science is settled here hasn't read the replication literature.
What survives the criticism
Two things, and they're enough to think with.
The observation-window arithmetic isn't a psychological claim at all. Short windows show more losses than long ones because that's what volatility around a positive drift does. That holds whether or not any experiment replicates. It's the same reason a drawdown feels different from a standard deviation, which we pulled apart in the work on the risk you actually feel.
And the Barber and Odean result is field data on real accounts, not a lab. Whatever the mechanism, high activity went with low returns in 66,465 households.
Where to check any of this yourself
Everything above rests on four published papers, and none of them is hidden. The lab result is Gneezy and Potters in the Quarterly Journal of Economics. The evaluation-period estimate is Benartzi and Thaler, built on CRSP monthly returns for 1926 to 1990. The market version is Gneezy, Kapteyn and Potters. The account-level evidence is Barber and Odean's 66,465 households. The replication failures are catalogued by Klos. Every one of them is linked at the foot of this page, and the figures quoted here come from the tables rather than the abstracts.
What would change the conclusion
Three things, and one is already partly true.
If the replication failures extend to the field rather than staying in the lab, the mechanism weakens considerably. Klos's evidence is real and it should temper how confidently anyone states this.
If trading costs approach zero, the Barber and Odean channel narrows. Commissions have collapsed since the 1990s, and the friction that explained most of their gap is smaller now. Turnover may still cost you through spreads, taxes and worse entry points, but the size of the effect from that era shouldn't be assumed to carry over intact.
And if someone genuinely checks without acting, the whole chain breaks. Looking is only expensive because it leads somewhere. Imagine a person who reviews a portfolio quarterly out of interest, then rebalances once a year on a rule they wrote down in advance. That isn't the investor these studies describe. It's closer to the discipline we set out in the piece on how often rebalancing actually needs to happen.
Turning it into something you can hold
The research doesn't prescribe a review cycle, and anyone quoting a specific number as though it were a finding is inventing it. What the work does support is a shape.
A review interval long enough that most reviews show a gain. A change tied to a rule set in advance rather than to what the number looks like on the day. Benartzi and Thaler's result puts the implied evaluation period of a typical investor somewhere near a year. If your review cycle is much shorter than your horizon, you've quietly volunteered for the high-frequency group.
The rule part matters as much as the interval. A quarterly review that ends in a trade whenever something looks bad is a daily habit with extra steps. A quarterly review that only acts when a weight breaches a band is a different thing entirely, and the trade it produces is usually the opposite of the one panic would suggest.
Two caveats. Someone drawing income from a portfolio has to look more often, because the cash has to come from somewhere. And anyone who has just made a large change is right to watch it — not to react, but because unmonitored mistakes stay mistakes.
A reasonable reading
The strong claim, that checking daily makes you measurably poorer, is more than the evidence will carry. The weaker claim is well supported and more useful. The frequency at which you evaluate a portfolio quietly sets the horizon you're actually investing over, and that horizon is usually much shorter than the one you'd say out loud.
So back to the two people who bought the same tracker on the same Monday. If your stated horizon is twenty years and your review cycle is every morning, one of those is wrong. Which of them is cheaper to change?