Two groups of people played the same gamble, with the same odds, for the same money. One group saw the result after every round. The other saw results only after every third round. The second group put 33% more of their money at risk.
Nothing about the bet changed. Only how often they looked.
The experiment
Uri Gneezy and Jan Potters ran the study and published it in the Quarterly Journal of Economics in 1997. Subjects played nine rounds. In each round they chose how much of a 200-cent endowment to stake on a lottery paying 2.5 times the stake with probability one third, and nothing with probability two thirds.
That bet is worth taking. Stake one unit and you gain 2.5 units a third of the time and lose one unit two thirds of the time, so the expected value is (1/3 × 2.5) − (2/3 × 1), or a positive sixth of a unit per round. How much you stake should depend on your risk aversion. It shouldn't depend on how often somebody shows you a scoreboard.
The paper spells out why the schedule matters. Weighing losses more heavily than gains at a rate of λ, a single lottery is attractive only if λ is below 1.25. Evaluate three of them together and it stays attractive up to λ of 1.56, because the probability of seeing a loss at all falls from 0.67 on one round to a much lower figure across three.
It did. Averaged across all nine rounds, the high-frequency group staked 50.5% of their endowment. The low-frequency group staked 67.4%. The gap held in every block of three rounds: 52.0 against 66.7 in rounds one to three, 44.8 against 63.7 in rounds four to six, and 54.7 against 71.9 in rounds seven to nine.
A Mann-Whitney test put the one-tailed significance at p = 0.002 across all rounds. In a second part of the experiment the pattern repeated, with 39.0% against 48.9%.
Why looking more often makes losses louder
The mechanism has two parts, and neither is exotic.
The first is loss aversion. Losses register more heavily than equivalent gains. The second is the arithmetic of observation windows. A volatile asset with a positive drift shows a loss surprisingly often over short windows and rarely over long ones. Check daily and you'll see red on something close to half of all days. Check once a decade and you'll almost never see it.
Put those together and the frequency of checking sets how many losses your brain has to process, which sets how risky the asset feels. Shlomo Benartzi and Richard Thaler named the combination myopic loss aversion.
Their test ran the logic backwards. If investors have prospect theory preferences, how often would they need to evaluate a portfolio for stocks and bonds to look equally attractive? Drawing from CRSP monthly returns for 1926 to 1990, they found the answer sat “in the neighborhood of 1 year (from 9 to 13 months)”. Take a one-year evaluation period as given, and the allocation that falls out is close to a 50-50 split between stocks and bonds.
That's the uncomfortable part. Most people describe a thirty-year horizon and then behave like someone with a one-year one. The horizon they act on is the horizon they measure on.
It shows up in markets, not just labs
A fair objection to any lab result is that real markets have prices, arbitrage and people with money on the line. Gneezy, Kapteyn and Potters tested exactly that, building a market experiment where assets traded rather than sitting in an envelope.
They found that more information and more flexibility produced less risk taking, and that market prices of risky assets were significantly higher when feedback frequency and decision flexibility were reduced. Market interaction didn't wash the behaviour out. It got capitalised into prices.
Their paper opens with a real case. In February 1999 Bank Hapoalim, Israel's largest mutual fund manager, said it would send fund performance information quarterly rather than monthly. Clients could still check daily if they wanted. The bank expected that fewer prompts would leave investors more willing to hold the funds.
The cost of the behaviour it drives
Checking isn't expensive on its own. What it does is raise the odds of acting, and acting is where the money goes.
Barber and Odean studied 66,465 households at a large discount broker over 1991 to 1996. The households that traded most earned an annual return of 11.4% while the market returned 17.9%. The average household earned 16.4% net. Gross of costs it earned 18.7%, slightly ahead of the index — so the shortfall was not stock picking. It was the friction of transacting, on turnover that exceeded 250% a year for the most active group.
That's the same shape as the fund-level gap between what investments returned and what investors in them earned, which we looked at in the piece on what bad timing costs the average dollar. Different data, different decade, same direction.
Checking is not the same as acting
It's worth separating two habits that usually travel together.
Opening an app and reading a number costs nothing. It changes nothing about the portfolio, the fees or the expected return. If that were all anyone did, none of this research would matter.
What the evidence describes is a chain. Frequent feedback makes the asset feel riskier, feeling riskier raises the odds of trimming, and trimming is where the cost lands. Barber and Odean's most active households turned over more than 250% of their portfolios a year. That's not a group calmly observing. That's a group reacting.
The Bank Hapoalim decision is instructive precisely because it didn't restrict anything. Clients could still check daily. The bank only changed what arrived unprompted. That's the lever most people actually have: not willpower, but how many times a week something puts the number in front of you.
Notifications are the modern version of that monthly statement, and they're set to a frequency nobody chose deliberately. A push alert on a 2% move is a design decision made by somebody whose interests aren't identical to yours.
The objections, and they have force
This is where a lot of behavioural writing stops. It shouldn't.
Myopic loss aversion has a mixed replication record, and the failures aren't obscure. Alexander Klos, reviewing them in Judgment and Decision Making, notes that Moher and Koehler reproduced the original Gneezy and Potters design successfully, but found no myopic loss aversion when subjects played similar gambles rather than the same gamble repeatedly, and none when the task was a choice between two gambles rather than an allocation.
More pointedly, Glätzle-Rützler and co-authors found no evidence of the effect in a sample of 755 adolescents aged 11 to 18, using the original design. Beshears and co-authors found no effect from aggregating return information despite a similar presentation to earlier work.
Klos proposes two candidate explanations. Task structure matters: certainty-equivalent framings appear to remove the effect that binary choices produce. And subjects severely misestimate long-term returns, which may be doing work that gets attributed to myopia.
So the honest position is narrower than the popular version. Feedback frequency changes risk taking in some designs and not others. It is a tendency with boundary conditions, not a law of nature. Anyone who tells you the science is settled here hasn't read the replication literature.
What survives the criticism
Two things, and they're enough to act on.
The observation-window arithmetic isn't a psychological claim at all. Short windows show more losses than long ones because that's what volatility around a positive drift does. That holds whether or not any experiment replicates. It's the same reason a drawdown feels different from a standard deviation, which we pulled apart in the work on the risk you actually feel.
And the Barber and Odean result is field data on real accounts, not a lab. Whatever the mechanism, high activity went with low returns in 66,465 households.
What would change the conclusion
Three things, and one is already partly true.
If the replication failures extend to the field rather than staying in the lab, the mechanism weakens considerably. Klos's evidence is real and it should temper how confidently anyone states this.
If trading costs approach zero, the Barber and Odean channel narrows. Commissions have collapsed since the 1990s, and the friction that explained most of their gap is smaller now. The turnover itself may still cost through spreads, taxes and worse entry points, but the size of the effect from that era shouldn't be assumed to carry over intact.
And if someone genuinely checks without acting, the whole chain breaks. Looking is only expensive because it leads somewhere. A person who reviews a portfolio quarterly out of interest and rebalances once a year on a rule is not the investor these studies describe — that's closer to the discipline we set out in the piece on how often rebalancing actually needs to happen.
Turning it into something you can act on
The research doesn't prescribe a review cycle, and anyone quoting a specific number as though it were a finding is inventing it. What the work does support is a shape.
Pick a review interval long enough that most reviews show a gain, and tie any change to a rule set in advance rather than to what the number looks like on the day. Benartzi and Thaler's result puts the implied evaluation period of a typical investor somewhere near a year. If your review cycle is much shorter than your horizon, you're volunteering for a version of the experiment where you're in the high-frequency group.
The rule part matters as much as the interval. A quarterly review that ends in a trade whenever something looks bad is a daily habit with extra steps. A quarterly review that only acts when a weight breaches a band is a different thing entirely, and the trade it produces is usually the opposite of the one panic would suggest.
Two caveats. Someone drawing income from a portfolio has to look more often, because the cash has to come from somewhere. And anyone who has just made a large change should watch it, not to react, but because unmonitored mistakes stay mistakes.
A reasonable reading
The strong claim, that checking daily makes you measurably poorer, is more than the evidence will carry. The weaker claim is well supported and more useful: the frequency at which you evaluate a portfolio quietly sets the horizon you're actually investing over, and that horizon is usually much shorter than the one you'd say out loud.
If your stated horizon is twenty years and your review cycle is every morning, one of those is wrong. The cheapest thing to change is the review cycle.