When I was fresh out of a school I worked in a call center for a while. One of the most important metrics (and everything you did was measured) was the average length of each call. So it was common practice, when you got cases that basically were going to be long calls or ones you probably couldn't actually resolve (e.g., "I want the service back on but cannot pay"), to deliberately upset the customers so that they would demand a supervisor and you could transfer them.
I worked in a call center, as well, but they didn’t prioritize call length as a metric/KPI. Their number one KPI was “first call resolution”, meaning even if it took an hour as long as the problem was resolved the call was considered a success. I feel like it was a much healthier way to measure call center performance.
One support team I used to work with prioritized "first touch resolution" for email cases as well, so the agents would answer every case with a 16-page DYI flowchart listing every single possible problem and its resolution. KPIs are hard...
My experience is most customers can't read past the first line of a response. The idea that they can get through a 16-page flowchart seems unlikely to me. And I do sell to a mostly technical audience.
It sounds pretty good to me if it inspires them to write a manual that answers the questions support gets. Most products come with a "manual" that's useless for troubleshooting.
We rotated through all of them and our pay was tied substantially to them, which I guess also created an incentive to make it hard to get good metrics all the time on the part of someone else in the organization.
I moved from a call center where length was a KPI to one where it is not a KPI, and it makes a world of difference in the pressure/health of the office environment (and thus better outcomes for customers)
It is far better when call length is not a KPI. I would never go back.
IMO, if call length is a KPI, it's a key indicator that the company isn't willing to hire enough agents and is trying to apply pressure to shorten calls for purposes of coverage.
I'd also expect a bunch of "I'm going to transfer you... oops, pressed a wrong button and it hang up!" As a customer, I've had a fair measure of mysteriously hang up calls to support, I wonder how many of those are metrics-driven...
Yeah, eventually they started measuring transfers as a metric, which indeed encouraged that behavior.
Of course some people would just say the issue was solved, falsely (which, you know, it was a little easier to get caught so it was higher-risk, in addition to being probably less moral), and there was also the X factor of how much you liked the customer, considering the extreme amount of verbal abuse the job entailed.
I mean, in general, a hostile tactic like this is likely to get the agent to do as little as possible as he can to help you while staying within the rules. Plus in many cases the front-line agent doesn't have the permissions to do what you need anyway and has to transfer you to a "supervisor" (usually not actually one) or a retention guy to resolve your issue anyway.
The call center I used to work for expected us to give one warning for such hostile behavior, and then simply hang up if the customer kept being abusive and/or threatening.
In general its, simply a bad idea to be rude to the person you expect to provide assistance.
As a corollary, this is along the lines of one of my favorite bits of wisdom from Poor Charlie's Almanac.
If I recall correctly, Munger considers incentives to be the single most important concept to properly understand in order to drive successful business (and arguably life) outcomes.
Not surprisingly, incentives are also chronically underestimated or outright ignored, even in situations where there's a strong, profit-driven incentive (#meta) to get them right.
> From all business, my favorite case on incentives is Federal Express. The heart and soul of their system – which creates the integrity of the product – is having all their airplanes come to one place in the middle of the night and shift all the packages from plane to plane. If there are delays, the whole operation can’t deliver a product full of integrity to Federal Express customers. And it was always screwed up. They could never get it done on time. They tried everything – moral suasion, threats, you name it. And nothing worked. Finally, somebody got the idea to pay all these people not so much an hour, but so much a shift and when it’s all done, they can all go home. Well, their problems cleared up overnight.
For a different view, see "Punished by Rewards: The Trouble with Gold Stars, Incentive Plans, A's, Praise, and Other Bribes".[1]
It's about extrinsic vs intrinsic motivation. Extrinsic motivation comes from outside of the person and is what's typically used by organizations to try to motivate people: salary, praise, bonuses, etc. Intrinsic motivation comes from within.
Luckily in most businesses incentives are used by the people designing the incentives to
1) cost as little as possible to the business
2) "optimize" their own pay/career
(There was a study saying that over 80% of sales people would deny their employer a million-dollar contract if it netted them $500 personally, so compared to that this is nothing)
Which makes sure that the incentives are not aligned with the business goals.
Now I guarantee that the best way to destroy intrinsic motivation, bar none, is to provide extrinsic motivation for things that aren't aligned with the business' goals.
I think this is frequently short-sighted. Surgeons are intrinsically motivated to be surgeons and help people, and they are trained to think clinically (pun somewhat intended) about risk-reward. They realize that trying to help high-risk patients will ultimately result in them being able to help fewer patients.
It also kind of incentivizes the alternative as well: if you aren’t an amazing surgeon, you could work to ensure you exclusively work on “hard” cases you can claim all failures are due to patient difficulty rather than lack of competence (at least competence to deal with that difficulty). Then you can dismiss any claims you’re incompetent with “I have a high mortality rate because I’m willing to risk my reputation to save people who have been abandoned”...
I've listened to a physician make this argument (for a different metric, not mortalities). He sounded as though he expected everyone in the audience to disbelieve him, enough so that I wondered why he even tried.
The problem is that in many cases, including this, you cannot directly measure the intended thing in any meaningful way. Your example is just another indirection which is still certainly somewhat playable and in addition involves estimation from insufficient input data (also known as guessing ;)).
"These doctors are vile dogs. If I was a physician I would take all the hard patients, heal them all, and be a hero of the people."
Yeah, if you were given one operation where the patient has a 100% chance of dying without it, but 80% of you killing them if you operate + 20% chance of them living after, maybe you'd do this once and get lucky. But try doing this twice, you have a a 96% chance of killing a patient, three times makes a 99.2% chance of killing at least a single patient. This probability quickly approaches 100%. If you kill a patient you get dragged into court, branded as a murderer by the opposing lawyer, and your career which you have spent your youth, your 20s and a half a million dollars training for is down the drain. Taking risky cases, if that's not your niche, is is guaranteed to have you out of a job and in debt from legal fees.
To suggest that "bad hearts" are involved in this necessary calculation, is incredibly naive.
People die in surgery all the time, it's not nearly as big a deal for the surgeon as you're suggesting. ~5% chance of death within 30 days per major hart surgery means the average surgeon is losing several people in an average year.
I am pretty sure that the parent was referring to the fact that this article is about heart surgeries. Bad heart is used in the same manner as a bad knee or a bad tooth. It refers to the patient not the doctor.
But aside from that, I completely disagree with your comment. The approach you are defending, even if legal, is completely unethical and immoral. It should be something we chastise, not condone.
Well I think you're an idiot. There's only one viable "approach", determined by the laws of probability I described above. You can try to shame people into doing actions that would ultimately get them fired and sued, but you're only going to look like a self-righteous asshole and nothing will happen.
In another sense, 'Goodhart' seems like a splendid cosmic coincidence. Instead of obsessing over targets, try to perceive the whole system. And who can do that? Good-hearted individuals.
The version I've heard might be a corollary: if you measure it, it will improve.
The subtext is that the system won't improve; rather, something in the system will be sacrificed to improve the metric. The end result (usually quality) won't improve.
Some guy wrote a taxonomy of Goodhart's Law cases. This would be the "Adversarial Goodhart" case: "When you optimize for a proxy, you provide an incentive for adversaries to correlate their goal with your proxy, thus destroying the correlation with your goal."
The first example on that page, "Regressional Goodhart", is totally wrong. If U measures V plus some noise X, assuming V and X form a bivariate normal distribution, then the conditional expectation of V is maximized by selecting the greatest U. In fact, the conditional expectation is linear in U.
Not sure how much stock I can put into the rest of the article given this. Goodhart's Law can seemingly only apply if the variables are highly non-normal, or are negatively correlated (e.g. worse doctors are more likely to refuse the difficult operations than good doctors).
> The first example on that page, "Regressional Goodhart", is totally wrong.
I don't think your statements contradict anything that it says. For reference, here are all the sentences in the "Quick Reference" part, which I'm assuming is the part you read.
Regressional Goodhart - When selecting for a proxy measure, you select not only for the true goal, but also for the difference between the proxy and the goal.
Model: When U is equal to V+X, where X is some noise, a point with a large U value will likely have a large V value, but also a large X value.
Thus, when U is large, you can expect V to be predictably smaller than U.
Example: height is correlated with basketball ability, and does actually directly help, but the best player is only 6'3", and a random 7' person in their 20s would probably not be as good
Is any particular one of these sentences false?
> If U measures V plus some noise X, assuming V and X form a bivariate normal distribution, then the conditional expectation of V is maximized by selecting the greatest U.
The word "maximized" carries some assumptions. If the only thing you can do is select based on U, then that statement is correct. However, if you had some means of selecting directly on V, then it's extremely likely that this would do better than selecting on U.
Suppose you're selecting the top 10 people. Suppose X ranges 0-10 chosen by a die roll, and V ranges 0-10, and it happens there are ten people with V=10, a hundred people with V=9, and a thousand with V=8 (and a lot more with lower Vs). On average, you'll have ten V=9s who score U=19, ten V=9s and a hundred V=8s who score U=18, and so on; in order for selecting on U to perform as well as selecting on V, every one of the V=10s would have to roll X=9 or X=10, which is exceedingly unlikely. It is true that taking people with high U scores yields people with higher Vs than taking people at random, and it is further true that taking the U=19s will give you better results than taking the U=18s or U=12s. It is simultaneously true that, when you take the U=19s, you'll be getting people whose X was 9 or 10, much higher than if you selected people at random or if you selected directly for high V. The first two sentences from the text state exactly this. (One consequence of this observation is, e.g., if you're doing admissions based on some test score, and you're considering raising the required score by ∆U, you should know the effect will be to raise average V and to raise average X, with ∆V=∆U-∆X, and if ∆X is large, you may be disappointed in the results. This is simple regression to the mean.)
I imagine you know all these concepts; I think you're interpreting the text as a stronger statement than it is.
On a less dire area of work, I’m starting to turn away from risky jobs/startups. It seems like there’s more and more newcomers in tech these days, and they’ll reject you for not being at gigs for longer. If I get the impression a future manager will be inexperienced, or that they’d failed to build a team before, I hesitate now. It’s a little frustrating, but that’s just how it is.
I prefer to think that this happens when you pick a measure that ignores key aspects of the results. For example, would the results be the same if the mortality rate was complexity-weighted?
Goodhart law deals with proxies - i.e. measures not of direct target, which is hard to measure, but of something that is easy to measure, and you consider it related to the actual target. If you need wins and measure wins, wins are not a proxy and thus Goodhart law does not apply. But if, say, it were hard for you to count wins in soccer for some reason, and you measured time spent with the ball per player instead, thinking that if players own the ball all the time, they surely going to win eventually - you'd get a team that is engaged in pointless passing around instead of actually scoring goals.
For surgeons, the direct target is for the patient to get better - or at least live longer and with higher quality of life. Proxy is the measure of a mortality of a specific surgeon. It is not a direct target - as the surgeon can achieve zero mortality by not doing any surgery at all, but his patients would die of suffer from the lack of treatment.
Pro tip for those targeting a high win percentage. If you win your first game (beginners luck?), then stop playing and never play again.
That's what the article is describing, surgeons won't even play the game, they won't even attempt to help certain people because it might hurt their "win percentage".
Wins are likely from systems with fixed rules that don't reward innovation, so winning by the rules of a system doesn't necessarily mean a positive outcome. An example might be the system that lead to the housing crisis. Certainly a few people who gamed things the right way "won" big, but only at the expense of millions of homeowners.
Mortality rates are one measure of many that can tell a story, but they're like incarceration rates in that they've become a perverse incentive that exacerbates an issue. If we say that putting more people in prison is a good thing, that's predicated on an assumption that only those that are being put in prisons are people worthy of being there, whose freedom is more costly to society than their imprisonment. When a district is rewarded for putting people in prisons, what types of things might you expect to happen to the legal systems and populations in those places? Would prosecutors then be incentivized to trump up charges and force people into serving prison time unnecessarily to make themselves look like bigger "winners"? Might police make more frivolous arrests or even stir up trouble in communities to paint a picture of rampant criminality so they can look to be "tough on crime"? Wouldn't you expect to see an all-causes decrease in long term crime?
The problem isn't just cherry picking, but if people can game a system built around limited and gameable measures, then it's going to encourage min-maxing for profit.
the problem here is that mortality rate is not correlated/normalized with the patient risk factor. the problem lies in the performance indicator itself, not strictly people trying to optimizing for what they are asked for. it is likely inhumane to refuse patients that have slightly above average risk - even if there are some considerations to be had, like having replacement organs go to waste, but that is already covered by the donor lists queues - but given people can find themselves out of work for having bad mortality rates, optimizing for the performance indicator it's the only logical conclusion.
also, normalizing for the patient risk will likely just lead to people overstating patient risk, because once you reduce performance to an index, people will focus on what's measured before on what the intended goal for the measurement is.
Physicians already have good reasons to overstate the patient's severity. It's part of the payment formula used by CMS. But there are plenty of incentives not to lie such as jail time.
I'm not an expert in cardiatric medicine or anything, but couldn't patient risk be determined by objective factors like patient age and history of previous heart issues?
I don't see how the analogy works. The entire point is you have a medical record with objective characteristics that can be tied to the risk of any given procedure.
My entire point is that isn't true at all - you have a medical record with some objective characteristics, but it is not possible to calculate the risk of a procedure based on that information. Just like in software, we have a project requirements document with many objective characteristics and we are unable to calculate the time required based on those characteristics.
So what is the risk calculation based on? Presumably the doctors refusing to do procedures they deem high-risk are not determining completely at random. The risk calculation doesn't have to be perfectly accurate, and it seems to me as though heart procedures are better defined than software projects.
(And besides that, although software estimation is notoriously inaccurate, it is still done and written into contracts all the time, because it serves a useful purpose that outweighs its inaccuracy)
Another way of looking at it is that the patients who are refused treatment count as losses. So the current system is not actually maximizing wins, it's maximizing win percentage for doctors.
Ya, I went down this line of thinking too. It does seem like this might too be a poor incentive because now doctors would operate on everyone to try and claim wins, even if the likelihood is exceedingly low. There’s probably some measure like “healthy six months from now” that incentivizes them to try in situations with high mortality where operating is the only chance to improve things but keep them from operating if the risk is too high.
Maybe it's a good metric for hospitals. There are other pressures discouraging hospitals from rejecting patients altogether, or such is my impression, anyway.
You say that as if you had a choice. The point of the article is that you don't. That yahoo doctor is the only one willing to operate on you. The rest -- presumably the smart ones -- are playing the game to win, so they'll sit out this turn, because you're a bad bet.
The point is that if you are already dying then you would prefer the surgery instead of doing nothing. The doctors seem to have no problem with actually doing the work but don't like their stats ruined, which is rather disconcerting given the consequences.
Your comment also shows that patients are just as much of an issue by focusing on the metric more than the cases involved.
You are committing a common fallacy - when evaluating risk, you are considering only one of the alternatives and forgetting the risk of the other. Yes, nobody wants to be operated on by a yahoo. But mortality is not a good measure of yahoo-ness, because in some cases not operating carries larger risk of death than operating. And if a surgeon decides to operate in this case, the decision is right even if resulting risk is high - because the alternative is even higher risk. Considering only mortality rate of operations ignores this. It is a variation of survivorship bias, only on the opposite side.
You presumably want the metric (operation, mortality rate); its useless to consider mortality across the board, because that assumes all activities are equally risky. Otherwise, you end up like this, just entirely removing the option of high-risk activities, regardless of the potential benefit.
Essentially, the doctor has decided that he doesn't want to take the risk of killing you, even if you're willing to accept it. Why should that be the doctor's choice to make, and why is he being given this additional risk, when the benefit/loss should clearly be yours?
A doctor might have a 5% mortality track record for a surgery that generally has a 20% mortality risk. I feel like you are pushing it to an extreme with this yahoo doctor example.
It becomes a bad measure because once it's a target people start to manipulate it in order to hit the target. In this case if a surgeon refuses to take anything but easy surgeries their stats will look amazing but it says very little about their actual abilities.
but he'll become good at accurately measuring risk, so if accepts to operate on you, you can be certain that your case is simple and your odds of dying on the table are slim.
Some doctors have high mortality records because they specialize in treating deadly diseases. And you'd rather be treated by a doctor who has only ever treated non-fatal ones? Well, that'll work if you never have anything worse than a common cold.
"When a measure becomes a target, it ceases to be a good measure."[a]
[a] https://en.wikipedia.org/wiki/Goodhart%27s_law